Skip to content

feat(conversation): ask a person before a tool call runs, on its ai_tool_call step - #725

Merged
martinmitrevski merged 3 commits into
acceleratefrom
athena/tool-approvals
Oct 2, 2026
Merged

martinmitrevski merged 3 commits into
acceleratefrom
athena/tool-approvals

Conversation

@martinmitrevski

@martinmitrevski martinmitrevski commented Oct 2, 2026 •

Copy link
Copy Markdown

What

A caller can declare that a person must allow each call of a tool before it runs. The question and the answer travel on the call's existing ai_tool_call attachment (the reply steps from #706), so clients read them the same way they read the rest of the step.

{ "type": "ai_tool_call", "id": "toolu_…", "name": "athena_device_location",
  "status": "awaiting_approval", "executor": "client",
  "target_user_id": "emp_…", "target_client_id": "ios-…",
  "approval": { "title": "Share your location?", "message": "Only your city, region, country and time zone are shared.",
                "reason": "to check the local weather", "allow_title": "Share location", "decline_title": "Don't share" } }
  • Declaration: SessionTool.approval holds title, message, reason_argument (the string argument in which the model says why), allow_title and decline_title.
  • The step: a call to such a tool opens as awaiting_approval. It is addressed to the person whose command it answers, plus their install for a client tool, and carries the question.
    • It is a step whatever visible_tools says, as a client tool's is, because somebody is waiting on it.
    • A client tool with no install to address asks nobody and runs as before.
  • The answer: the caller reports it on the events socket with tool_approval {tool_call_id, command_id, turn_id, allowed, summary}. It is bound to the command and turn like tool_result.
    • Allowed: approval.decision: "allowed", and the call goes on (awaiting_client for a client tool, running otherwise).
    • Declined: cancelled with decision: "declined" and the caller's summary, so older clients still read it as finished.
    • The call is still answered with tool_result either way. An unanswered one shows the caller's reason as its summary.
  • Limits: an open question is never trimmed or dropped; an answered one gives up its message and reason before summaries go.

The schemas are in api/legacy.yaml. generated.go, openapi.yaml and the Go client are regenerated. The JS, Python plugin and Swift clients are left as they were. All three had drifted from the spec before this change: on accelerate the js job already fails at npm run types -- --check, and regenerating breaks the JS SDK's own code, which still uses fields the spec dropped. So this PR's js job fails the same way it does on accelerate.

Athena's side (the API holding the call until the employee answers, and the iOS/web clients) is GetStream/athena-ai#18. The iOS views are in stream-chat-swift-ai 0.10.0.

Tests

  • New tests for the step, the decision, a server tool's question, Stream's limits, the conversation path (DisplaySuite), the command/turn binding (SessionSuite) and the declaration reaching the spec.
  • internal/conversation, internal/session, internal/harness, internal/agent and internal/api pass. go vet ./... is clean, and CI's generated-code check passes locally.
  • Timing flakes: in the full go test ./..., a different unrelated subtest failed in each run, and both pass alone and on rerun. Locally that was TestAgentSuite/TestASecondVoiceAtTheSameMicrophoneIsNotTakenForTheCaller and TestHarnessSuite/TestTheModelAskingAgainDoesNotReplaceTheCallersImages. CI's first go run failed TestAgentSuite/TestAToolThatKeepsFailingIsNotAnsweredForever, which passes 8 of 8 locally. None of them touch tools or approvals. A full local run on clean accelerate passed.
  • On a local Athena stack (this runtime, DeepSeek V4 Flash), a Stream client playing the iPhone went awaiting_approval → allowed → awaiting_client → completed. It also checked the declined path. The unanswered path (failed after the caller's 120 s wait) was checked on an earlier build of the same change.

🤖 Generated with Claude Code

…ool_call step

A caller can declare that a person must allow each call of a tool
(SessionTool.approval: a title, message, the argument holding the model's
reason, and button labels). In a persistent conversation the call's
ai_tool_call attachment then opens as awaiting_approval, addressed to the
person whose command it answers (and, for a client tool, their install), with
an approval object for their client to ask:

  "approval": {"title": "Share your location?", "message": "...",
               "reason": "to check the local weather",
               "allow_title": "Share location", "decline_title": "Don't share"}

Such a call is a step whatever visible_tools says, as a client tool's is:
somebody is waiting on it.

The caller collects the answer and reports it on the events socket with
tool_approval {tool_call_id, command_id, turn_id, allowed, summary}, bound to
the command and turn like tool_result. Allowed, the call goes on as it would
have (awaiting_client for a client tool, running otherwise) with
approval.decision "allowed". Declined, it is cancelled with decision
"declined" and the caller's summary, so older clients still read it as over.
The call is still answered with tool_result either way; an unanswered one
shows the caller's reason as its summary.

An open question is never trimmed or dropped to fit Stream's limits; an
answered one gives up its message and reason before summaries go.

The schemas are in api/legacy.yaml; generated.go, openapi.yaml and the Go and
JS clients are regenerated. The JS types also catch up with spec changes they
had not picked up before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added the config label Oct 2, 2026
martinmitrevski and others added 2 commits October 2, 2026 16:25
The committed types were already out of step with the spec on accelerate
(the js job fails at `npm run types -- --check` there), and regenerating them
breaks the SDK's own code, which still uses fields the spec dropped. That is
work of its own; this change does not need the JS client.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rovals

# Conflicts:
#	acceleration/internal/api/generated.go
@martinmitrevski
martinmitrevski merged commit c8c0d4c into accelerate Oct 2, 2026
10 of 12 checks passed
@martinmitrevski
martinmitrevski deleted the athena/tool-approvals branch October 2, 2026 15:57
kanat pushed a commit that referenced this pull request Oct 5, 2026
…ool_call step (#725)

* feat(conversation): ask a person before a tool call runs, on its ai_tool_call step

A caller can declare that a person must allow each call of a tool
(SessionTool.approval: a title, message, the argument holding the model's
reason, and button labels). In a persistent conversation the call's
ai_tool_call attachment then opens as awaiting_approval, addressed to the
person whose command it answers (and, for a client tool, their install), with
an approval object for their client to ask:

  "approval": {"title": "Share your location?", "message": "...",
               "reason": "to check the local weather",
               "allow_title": "Share location", "decline_title": "Don't share"}

Such a call is a step whatever visible_tools says, as a client tool's is:
somebody is waiting on it.

The caller collects the answer and reports it on the events socket with
tool_approval {tool_call_id, command_id, turn_id, allowed, summary}, bound to
the command and turn like tool_result. Allowed, the call goes on as it would
have (awaiting_client for a client tool, running otherwise) with
approval.decision "allowed". Declined, it is cancelled with decision
"declined" and the caller's summary, so older clients still read it as over.
The call is still answered with tool_result either way; an unanswered one
shows the caller's reason as its summary.

An open question is never trimmed or dropped to fit Stream's limits; an
answered one gives up its message and reason before summaries go.

The schemas are in api/legacy.yaml; generated.go, openapi.yaml and the Go and
JS clients are regenerated. The JS types also catch up with spec changes they
had not picked up before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore(sdks/js): leave the JS types as accelerate has them

The committed types were already out of step with the spec on accelerate
(the js job fails at `npm run types -- --check` there), and regenerating them
breaks the SDK's own code, which still uses fields the spec dropped. That is
work of its own; this change does not need the JS client.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant