Skip to content

feat: add optional jev thinking-mode voice router - #2344

Merged
halajohn merged 7 commits into
TEN-framework:mainfrom
JianYan11:feat/jev-thinking-router
Oct 3, 2026
Merged

halajohn merged 7 commits into
TEN-framework:mainfrom
JianYan11:feat/jev-thinking-router

Conversation

@JianYan11

@JianYan11 JianYan11 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add a TypeSafe Jev decision extension with choice, score, and noul support, cancellation, and stable error responses.
  • Add an opt-in voice assistant graph that routes DeepSeek Flash between non-thinking and high-effort thinking modes without tools.
  • Show the Jev route and completed DeepSeek reasoning separately in chat, while keeping reasoning out of TTS output.
  • Add placeholder environment settings and regression tests.

Rollback state

  • Restored the initial PR commit 70e4d8d.
  • The later OpenAI migration and review fixes are not included. The two P1 findings about the default graph in the existing review comment apply again.

Verification

  • The initial revision previously passed 40 related Python tests, formatting, and a frontend type check.
  • Live voice end-to-end playback remains unverified.

@github-actions

Copy link
Copy Markdown

Found two regressions affecting the default voice_assistant graph:

  • [P1] Restore the default LLM API key (ai_agents/agents/examples/voice-assistant/tenapp/property.json:45). This edit removes ${env:OPENAI_API_KEY} from the auto-start graph. OpenAILLM2Config.api_key defaults to empty and openai_llm2_python/extension.py:on_start returns without creating a client when it is empty, so users with a configured OpenAI key now get no LLM responses. Restore the property and assert it in the graph regression test.

  • [P1] Restrict the DeepSeek thinking parameter to the routed graph (ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_python/agent/llm_exec.py:194). Even when model_routing.enabled is false, every request includes extra_body.thinking={"type":"disabled"}. The shared OpenAI-compatible extension forwards that field to the default OpenAI endpoint (and other existing providers), where thinking is not a Chat Completions parameter; a valid key alone will not restore those conversations. Keep the original {"temperature": 0.7} request for routing-disabled graphs, and test the parameters sent in that mode.

ASR checklist (new speech_final consumer, no ASR plugin diff): lifecycle N/A; connection state N/A; buffering N/A; finalize N/A; reconnect N/A; result shape pass (reads metadata.asr_info.speech_final, leaving ASR output unchanged); metrics N/A; tests pass for simulated turn events, live voice/ASR-to-LLM behavior remains unverified. No ASR-plugin MUST rule was changed or violated by this diff.

JianYan11 added a commit to JianYan11/ten-framework that referenced this pull request Sep 28, 2026
@github-actions

Copy link
Copy Markdown

[P3] Log the reasoning effort actually sent to OpenAI (ai_agents/agents/ten_packages/extension/openai_llm2_python/openai.py:240). The routed graph passes reasoning_effort=none or high as a top-level request parameter, but this new log reads only extra_body.thinking. Both routes therefore report thinking=default, hiding the selected mode when diagnosing real requests. Log req.get("reasoning_effort") for OpenAI requests (and retain the extra_body check for providers that use it). OpenAI Docs confirms GPT-5.4 mini supports these reasoning efforts.

ASR design review

  • Lifecycle: N/A (no ASR plugin changes).
  • Connection state: N/A.
  • Buffering: N/A for ASR audio; the new consumer accumulates finalized text until a turn boundary.
  • Finalize: N/A (no finalize protocol changes).
  • Reconnect: N/A.
  • Result shape: pass; the new graph reads existing metadata.asr_info.speech_final from Deepgram without changing emitted ASR results.
  • Metrics: N/A.
  • Tests: pass for simulated turn events; live ASR-to-LLM playback remains unverified.

No ASR-plugin MUST rule is changed by this diff.

JianYan11 added a commit to JianYan11/ten-framework that referenced this pull request Sep 28, 2026
@github-actions

Copy link
Copy Markdown

[P2] Fall back to the fast route for malformed Jev confidence (ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_python/routing.py:74-88). The response validator turns a missing, boolean, non-finite, or out-of-range confidence into None, but then passes the original unvalidated decision to select_dest(), which selects llm_deep for any such response. A deep choice is even reported as selected_deep despite the invalid confidence. The README promises that invalid responses fall back to no thinking, and the new test at tests/test_jev_routing.py:61 codifies the opposite behavior for boolean confidence. Return invalid_response/fast_dest when confidence is malformed, and reserve the deep route for a valid fast confidence below the threshold or a valid deep decision; cover missing and out-of-range confidence too.

ASR checklist

  • Lifecycle: N/A (ASR extension unchanged).
  • Connection state: N/A.
  • Buffering: N/A for ASR audio; the new consumer accumulates finalized transcript segments until speech_final and resets on session change.
  • Finalize: N/A (ASR finalize protocol unchanged).
  • Reconnect: N/A.
  • Result shape: pass; reads existing Deepgram metadata.asr_info.speech_final without changing emitted ASR results.
  • Metrics: N/A.
  • Tests: pass for simulated turn boundaries; live ASR-to-LLM playback has not been rerun.

No ASR-plugin MUST rule is changed by this PR.

@github-actions

Copy link
Copy Markdown

[P2] Avoid retransmitting the growing utterance on every ASR update (ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_python/extension.py:125-140). In speech_final mode, every interim/final segment copies and joins all earlier final segments, then sends the entire growing string as a new non-final transcript. message_collector2 queues every snapshot and sends only one ~1 KB chunk per 40 ms. During sustained speech without a speech_final pause, this creates quadratic work/traffic and an unbounded collector backlog; the final transcript and visible answer can be stuck behind stale partials. Bound the pending text and throttle/coalesce partial updates (or send only changed segments), then cover a long continuous utterance and final-message ordering in a test.

ASR checklist

  • Lifecycle: N/A (no ASR plugin change).
  • Connection state: N/A.
  • Buffering: fail in the new downstream transcript consumer (unbounded snapshots); ASR audio buffering unchanged.
  • Finalize: N/A (plugin finalize unchanged).
  • Reconnect: N/A.
  • Result shape: pass (reads existing metadata.asr_info.speech_final; does not change ASR output).
  • Metrics: N/A.
  • Tests: fail for sustained/long turns; short simulated turns are covered, live playback is unverified.

No ASR-plugin MUST contract is changed by this diff; the finding concerns its new downstream consumer.

@JianYan11
JianYan11 force-pushed the feat/jev-thinking-router branch from 58a917a to 70e4d8d Compare September 30, 2026 07:59
@github-actions

Copy link
Copy Markdown

The restored PR head (70e4d8d) still has four actionable issues reported in earlier comments:

  • [P1] Restore the OpenAI key in the default graph (ai_agents/agents/examples/voice-assistant/tenapp/property.json:45). The auto-start voice_assistant graph no longer supplies api_key to openai_llm2_python; its empty default prevents client startup even when OPENAI_API_KEY is configured. Please assert the key in the graph test.
  • [P1] Do not send DeepSeek thinking options from routing-disabled graphs (ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_python/agent/llm_exec.py:194). The default graph now forwards extra_body.thinking={"type":"disabled"} to the configured OpenAI-compatible provider. Keep the original temperature-only parameters when routing is disabled and test that request.
  • [P2] Reject malformed confidence before selecting a route (ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_python/routing.py:74). Boolean, missing, non-finite, or out-of-range confidence becomes None, but select_dest() still chooses llm_deep; invalid responses should fall back to llm_fast. The boolean-confidence test currently expects the faulty deep route.
  • [P2] Bound/coalesce transcript snapshots during sustained speech (ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_python/extension.py:125). Each ASR update re-joins all accumulated final segments and queues the growing utterance with message_collector2, which sends only one chunk every 40 ms. Long turns can accumulate stale partials ahead of the final transcript and answer. Test long-turn final-message ordering.

ASR checklist

Lifecycle: N/A; connection state: N/A; buffering: fail (new downstream transcript consumer, not ASR audio buffering); finalize: N/A; reconnect: N/A; result shape: pass (uses existing metadata.asr_info.speech_final); metrics: N/A; tests: fail for long continuous turns (short simulated turns covered; live voice unverified). No ASR plugin is modified, so no plugin-specific MUST contract change was identified.

@github-actions

Copy link
Copy Markdown

Actionable findings:

  • [P2] Reject malformed Jev confidence before selecting a route — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/routing.py:74-88 normalizes a missing, boolean, non-finite, or out-of-range confidence to None, but then passes the original decision to select_dest(). That method returns llm_deep for the invalid value, so a malformed provider response is reported as below_threshold/selected_deep and enables thinking. The README says invalid Jev responses fall back to Flash without thinking, and the boolean-confidence test currently codifies the faulty deep destination. Validate the confidence and return fast_dest with an invalid_response status before calling select_dest(); add cases for missing and out-of-range values.

  • [P2] Bound/coalesce ASR transcript snapshots — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/extension.py:132-147 joins every finalized segment with the current interim text and sends the whole growing utterance as a new non-final message on every ASR update. message_collector2 gives each snapshot a new ID and puts all chunks on an unbounded queue, so a long uninterrupted utterance creates quadratic join/serialization work and a backlog of stale snapshots ahead of the final transcript and answer. Throttle/coalesce updates or send only the changed portion, and cover a long-turn ordering case.

  • [P2] Keep the disabled-routing target present — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/routing.py:45-46 returns the hard-coded destination llm when routing is disabled, but voice_assistant_jev_router defines only llm_fast and llm_deep. Any property override that disables routing (the config is public and defaults to disabled) sends chat_completion to an extension that is not in this graph, so the turn gets no response. Route disabled mode to an existing node (or add an llm node) and test the actual graph configuration.

  • [P2] Honor non-streaming LLM requests — ai_agents/agents/ten_packages/extension/deepseek_llm2_python/client.py:213-263 forwards request_input.streaming=False but always iterates the result with async for. The OpenAI SDK returns a ChatCompletion object for stream=False, which is not async-iterable, so the standard chat_completion interface fails for non-streaming callers. Add a non-streaming response path and a regression test; the current test suite only exercises streaming=True.

  • [P2] Run the new example regression tests in CI — the Jev tests were added under ai_agents/agents/examples/voice-assistant/tests/, while .github/workflows/ai_agents.yaml runs task test, whose test-agent-extensions target only discovers agents/ten_packages/extension/*/tests. Consequently the graph, routing, reasoning-display, and LLM-exec tests in this PR are not executed by the required CI job. Wire that directory into the task/workflow or move the tests to a discovered target.

ASR design review

  • Lifecycle: N/A — no ASR extension lifecycle changed.
  • Connection state: N/A — Deepgram connection handling is unchanged.
  • Buffering: FAIL — the new downstream consumer queues unbounded, repeated text snapshots during a turn.
  • Finalize: N/A — no ASR finalize protocol changed.
  • Reconnect: N/A — no reconnect path changed.
  • Result shape: PASS — the graph reads the existing metadata.asr_info.speech_final field without changing emitted ASR results.
  • Metrics: N/A — no ASR metrics changed.
  • Tests: PARTIAL — short synthetic turn cases exist, but the long-turn/backlog case and live ASR-to-LLM playback are unverified; the example tests are also not currently discovered by CI.

No ASR-plugin MUST rule is changed by this PR; the buffering finding is in the new downstream ASR consumer.

@github-actions

Copy link
Copy Markdown

Findings:

  • [P2] Reject malformed Jev confidence before selecting a destination — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/routing.py:74-88. Missing, boolean, non-finite, or out-of-range confidence is normalized to None, but the original response is then passed to select_dest(), which sends the turn to llm_deep and reports below_threshold/selected_deep. The README promises that invalid Jev responses fall back to Flash without thinking, so a malformed provider response currently increases latency and exposes reasoning. Validate the confidence before selecting a route and return fast_dest with an invalid_response (or equivalent fallback) status; add cases for missing and out-of-range values alongside the existing boolean case.

  • [P2] Coalesce speech-final transcript snapshots — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/extension.py:132-147. Every ASR interim/final update re-joins all accumulated segments and queues the entire growing utterance as a new message. message_collector2 emits only one queued chunk every 40 ms, so sustained speech creates quadratic join work and an ever-growing queue of stale partials; the final user transcript and answer can sit behind old snapshots. Send only changed/latest partial state (or throttle/coalesce it) and add a long-turn test that verifies final-message ordering.

  • [P2] Bound the routed conversation context — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/agent/llm_exec.py:161-187. _queue_context() appends every user and assistant turn, and every request copies the full list into messages; there is no turn/token limit or use of the configured max_memory_length convention. A long-lived voice session will eventually exceed DeepSeek's context window and fail subsequent turns while latency and request size grow without bound. Add a configurable bounded history/token budget (preserving the current turn) and a multi-turn regression test.

ASR design review

  • Lifecycle: N/A (no ASR extension lifecycle code changed).
  • Connection state: N/A (the existing Deepgram connection is unchanged).
  • Buffering: FAIL for the downstream transcript consumer; the speech-final path queues unbounded snapshots (finding above).
  • Finalize: N/A (no ASR finalize protocol changed).
  • Reconnect: N/A (no ASR reconnect code changed).
  • Result shape: PASS; the controller reads existing metadata.asr_info.speech_final and does not rewrite ASR results.
  • Metrics: N/A (no ASR metrics code changed).
  • Tests: FAIL for sustained/long turns and final-message ordering; current tests cover only short simulated events.

Validation: repository CI reports the voice-assistant job, lint, and build checks passing; live provider voice playback remains unverified.

@github-actions

Copy link
Copy Markdown

Additional confirmed findings:

  • [P2] Keep disabled routing on an extension that exists in the graph — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/routing.py:45-46. The disabled branch returns the hard-coded destination llm, but voice_assistant_jev_router defines only llm_fast and llm_deep. If the exposed model_routing.enabled property is overridden to false, _send_to_llm() sends chat_completion to a nonexistent extension and the turn receives no answer. Route disabled mode to an actual node (or add the llm node) and test the graph configuration.

  • [P2] Handle non-streaming chat completions — ai_agents/agents/ten_packages/extension/deepseek_llm2_python/client.py:213-263. The request forwards request_input.streaming=False, but the response is always consumed with async for. The OpenAI SDK returns a ChatCompletion object for stream=False, which is not async-iterable, so callers using the standard non-streaming LLM request path get a TypeError wrapped as CreateChatCompletion failed. Add the single-response conversion and a streaming=False regression test; the current tests cover only streaming.

  • [P2] Ensure the new example tests run in CI — ai_agents/Taskfile.yml:68-83 discovers only agents/ten_packages/extension/**/tests, while this PR puts the Jev tests under agents/examples/voice-assistant/tests. The ci (voice-assistant) workflow invokes task test, so these graph/routing/reasoning tests are skipped despite the green CI result. Add the example test directory to the task/workflow or move the tests under a discovered target.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Findings on the current head (e9a8078):

  • [P1] Keep disabled routing on an extension that exists in the graph — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/routing.py:45-46 and agent/llm_exec.py:228-230 return the hard-coded destination llm when model_routing.enabled is false. The voice_assistant_jev_router graph defines only llm_fast and llm_deep, so a normal property override to disable routing (and the config's default) sends chat_completion to a nonexistent destination and produces no answer. Return config.fast_dest (or add an llm node) and exercise the actual graph with routing disabled.

  • [P2] Fall back before selecting a route when Jev confidence is malformed — routing.py:74-88 normalizes missing, boolean, non-finite, and out-of-range confidence to None, but then passes the original response to select_dest(). That method sends every such response to llm_deep, so an invalid fast response is reported as below_threshold and an invalid deep response is reported as selected_deep, enabling reasoning on malformed provider data. Return fast_dest with an invalid_response status whenever confidence is invalid; add missing and out-of-range cases to the routing tests.

  • [P2] Coalesce speech-final transcript snapshots — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/extension.py:215-225 joins all finalized segments with the current interim text and enqueues the complete growing utterance for every ASR update. message_collector2 assigns each snapshot a new queue item and drains only one chunk every 40 ms, so a sustained utterance causes quadratic joining/serialization work and an unbounded backlog of stale partials ahead of the final transcript and answer. Throttle/coalesce to the latest partial (or send only the changed segment), and test a long turn's final-message ordering.

  • [P2] Bound the routed conversation history — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/agent/llm_exec.py:151-187 appends every user and assistant turn to self.contexts and copies the full list into each request. There is no turn/token limit or use of the existing max_memory_length convention, so a long-lived session eventually exceeds the DeepSeek context window while request size and latency grow without bound. Add a configurable bounded history/token budget that preserves the current turn and cover multi-turn eviction.

  • [P2] Handle non-streaming LLM requests — ai_agents/agents/ten_packages/extension/deepseek_llm2_python/client.py:213-263 forwards request_input.streaming=False but always consumes the SDK result with async for. The OpenAI SDK returns a ChatCompletion object for stream=False, which is not async-iterable, so callers using the standard non-streaming LLM interface receive CreateChatCompletion failed. Add the single-response conversion and a streaming=False regression test.

  • [P2] Run the new example regression tests in CI — ai_agents/Taskfile.yml:68-80 discovers only agents/ten_packages/extension/**/tests, while this PR places the graph, routing, reasoning-display, and DeepSeek tests under agents/examples/voice-assistant/tests. The workflow's task test therefore reports success without executing those tests. Add the example directory to the test target/workflow or move the tests under a discovered extension target.

ASR design review

  • Lifecycle: N/A — no ASR extension lifecycle code changed.
  • Connection state: N/A.
  • Buffering: FAIL — the new downstream speech-final consumer queues unbounded repeated snapshots during a sustained turn.
  • Finalize: N/A — the vendor ASR finalize protocol is unchanged.
  • Reconnect: N/A.
  • Result shape: PASS — the controller reads the existing metadata.asr_info.speech_final field without changing emitted ASR results.
  • Metrics: N/A.
  • Tests: PARTIAL/FAIL — short synthetic turn cases exist, but long-turn/final ordering is not covered, the example tests are skipped by CI, and live ASR-to-LLM playback is unverified.

No ASR-plugin MUST lifecycle or wire-contract rule is changed by this PR; the buffering failure is in the new downstream consumer.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Actionable findings:

  • [P1] Package main_jev_python in the voice-assistant image — ai_agents/agents/examples/voice-assistant/tenapp/property.json:1784 adds a graph whose controller is the new main_jev_python addon, but ai_agents/agents/examples/voice-assistant/Dockerfile still copies only ten_packages/extension/main_python. The builder therefore has no source for the referenced addon, and the released image cannot instantiate voice_assistant_jev_router even though the Docker build passes. Copy this extension in the Dockerfile (or move it to a packaged shared extension) and add a build/runtime check for the opt-in graph.

  • [P2] Treat malformed Jev confidence as an invalid response — ai_agents/agents/examples/voice-assistant/tenapp/ten_packages/extension/main_jev_python/routing.py:74-83 converts missing, boolean, non-finite, and out-of-range confidence to None, then passes the original response to select_dest(). That method sends the request to llm_deep, so malformed provider output enables thinking and is reported as below_threshold/selected_deep; the boolean case at tests/test_jev_routing.py:61 currently codifies this. Return fast_dest with an invalid-response fallback before route selection, and cover missing and out-of-range confidence too.

  • [P2] Keep routing-disabled mode on an extension that exists — main_jev_python/routing.py:45-46 returns the hard-coded destination llm when model_routing.enabled is false, while voice_assistant_jev_router defines only llm_fast and llm_deep. Any per-session property override that disables routing sends chat_completion to a nonexistent extension and silently loses the turn. Use an existing destination (the fast/no-thinking node) or add an llm node, and test the actual graph configuration.

  • [P2] Handle non-streaming LLM requests — ai_agents/agents/ten_packages/extension/deepseek_llm2_python/client.py:213-263 forwards request_input.streaming=False but always consumes the OpenAI response with async for. The SDK returns a single ChatCompletion object for stream=False, so standard non-streaming callers get a TypeError wrapped as CreateChatCompletion failed. Add a non-streaming conversion path and a regression test; the current test only exercises streaming.

  • [P2] Coalesce ASR transcript snapshots — main_jev_python/extension.py:215-225 joins the complete accumulated utterance and sends a new message on every accepted ASR update. message_collector2 puts every snapshot on an unbounded queue and emits one chunk every 40 ms; sustained speech therefore creates quadratic join/serialization work and stale messages that can sit ahead of the final transcript and answer. Throttle/coalesce partials (and flush stale snapshots on interruption) and add a long-turn ordering test.

  • [P2] Bound the routed conversation history — main_jev_python/agent/llm_exec.py:161,186-187 appends every user/assistant message and copies the full list into each request without a turn or token limit. Long-lived sessions will eventually exceed DeepSeek's context window while request size and latency grow without bound. Add a configurable history/token budget and a multi-turn regression test.

  • [P2] Run the new example tests in CI — ai_agents/Taskfile.yml:75 discovers only agents/ten_packages/extension/**/tests, while this PR places the graph, routing, ASR buffering, reasoning-display, and executor tests under agents/examples/voice-assistant/tests. The green task test job therefore skips the new behavior entirely. Add that directory to the test task/workflow or move the tests under a discovered extension target.

ASR design review

  • Lifecycle: N/A — the Deepgram ASR extension itself is unchanged.
  • Connection state: N/A — no ASR connection code changed.
  • Buffering: FAIL — the new downstream consumer queues unbounded, repeated transcript snapshots during a turn (finding above).
  • Finalize: N/A — no ASR finalize protocol changed.
  • Reconnect: N/A — no ASR reconnect path changed.
  • Result shape: PASS — the controller reads the existing metadata.session_id and metadata.asr_info.speech_final fields without changing emitted ASR results.
  • Metrics: N/A — no ASR metrics changed.
  • Tests: PARTIAL — short synthetic turn cases exist, but long-turn queue ordering, live ASR-to-LLM playback, and CI collection of these example tests remain unverified.

No ASR-plugin MUST rule is changed by this PR; the buffering failure is in the new downstream consumer.

@halajohn
halajohn merged commit ad20a17 into TEN-framework:main Oct 3, 2026
20 of 21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants