Skip to content

feat: add ModelBest VoxCPM TTS plugin - #658

Open
lottshin wants to merge 3 commits into
GetStream:mainfrom
lottshin:feat/modelbest-voxcpm-tts
Open

lottshin wants to merge 3 commits into
GetStream:mainfrom
lottshin:feat/modelbest-voxcpm-tts

Conversation

@lottshin

@lottshin lottshin commented Sep 20, 2026 •

Copy link
Copy Markdown

What this adds

Adds a VoxCPM TTS plugin backed by ModelBest's hosted Audio Speech API.

The plugin:

  • streams ModelBest's SSE/WAV response into Vision Agents PcmData chunks
  • supports reference-audio voice cloning
  • supports delivery conditioning with prompt_audio and prompt_text
  • closes an in-flight request when synthesis is interrupted
  • works without local VoxCPM weights, PyTorch, or CUDA
  • includes a small smoke-test example for checking credentials and model access

Verification

  • 20 non-integration tests pass
  • Ruff, mypy, extra validation, and uv lock --check pass
  • Python 3.13 workspace dependency resolution passes
  • The plugin wheel and source distribution build successfully
  • Tested against the ModelBest endpoint with normal synthesis and reference-audio cloning
  • Verified the generated output at 48 kHz
  • Confirmed that the example files are not included in the plugin wheel

The credential-gated integration test is included but was not run in CI or as part of this PR.

@coderabbitai

coderabbitai Bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

Understand this PR’s impact

Explore downstream dependencies and potential security impact with Blast Radius.

View blast radius →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: GetStream/Vision-Agents/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 900337e4-9504-414e-8a78-6efa72f4c992

📥 Commits

Reviewing files that changed from the base of the PR and between 8d4cc85 and e823307.

📒 Files selected for processing (1)
  • plugins/voxcpm/example/voxcpm_smoke.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • plugins/voxcpm/example/voxcpm_smoke.py

Included review availability: Your plan provides up to 8 included reviews per hour; 6 remain after this review.


📝 Walkthrough

Walkthrough

This change adds a VoxCPM TTS plugin backed by ModelBest’s Audio Speech API. The client validates WAV and cloning inputs, sends authenticated streaming requests, parses SSE responses, yields PCM audio, and supports cancellation and cleanup. The package is registered in the workspace and optional dependencies. Documentation, examples, public exports, type metadata, and tests are included.

Priority: ➖ Normal


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: GetStream/Vision-Agents/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 79d299ae-4479-41f3-befc-ec0b50adc9bc

📥 Commits

Reviewing files that changed from the base of the PR and between ba2fbf1 and 638bf6c.

⛔ Files ignored due to path filters (1)
  • uv.lock is excluded by !**/*.lock
📒 Files selected for processing (9)
  • README.md
  • agents-core/pyproject.toml
  • plugins/voxcpm/README.md
  • plugins/voxcpm/py.typed
  • plugins/voxcpm/pyproject.toml
  • plugins/voxcpm/tests/test_tts.py
  • plugins/voxcpm/vision_agents/plugins/voxcpm/__init__.py
  • plugins/voxcpm/vision_agents/plugins/voxcpm/tts.py
  • pyproject.toml

Included review availability: Your plan provides up to 8 included reviews per hour; 7 remain after this review.

Comment thread plugins/voxcpm/py.typed
@@ -0,0 +1 @@

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Move py.typed into the import package.

The wheel only includes the vision_agents tree. This root-level marker is not installed with vision_agents.plugins.voxcpm.

Move it to plugins/voxcpm/vision_agents/plugins/voxcpm/py.typed.

data_lines = []
continue
if line.startswith("data:"):
data_lines.append(line.removeprefix("data:").lstrip())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🛡️ Analyzed with Security Review | 🟠 Major | ⚡ Quick win

Denial of Service

Reachability: External
Exploitability: Difficult
CWE: CWE-400 — Uncontrolled Resource Consumption

Bound the accumulated SSE event size.

A server can send unlimited newline-terminated data: lines without a blank delimiter. data_lines then grows until the process exhausts memory. The socket read timeout only limits idle periods.

Track the accumulated event size and reject events above a fixed limit, such as 1 MiB.

raise ValueError("prompt_audio and prompt_text must be provided together")

self.voice = voice
self._base_url = base_url.rstrip("/")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🛡️ Analyzed with Security Review | 🟠 Major | ⚡ Quick win

Sensitive Data Exposure

Reachability: Internal
Exploitability: Moderate
CWE: CWE-319 — Cleartext Transmission of Sensitive Information

Reject plaintext credential destinations.

base_url accepts non-HTTPS URLs. _request() then sends Authorization: Bearer ... to that URL. A network observer can recover the API key when a caller configures an HTTP endpoint.

Require HTTPS for non-loopback hosts. The HTTPS default does not protect a custom base_url.

async def _stream() -> AsyncIterator[PcmData]:
async with self._lock:
self._stop_event.clear()
response = await self._request(text)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

🔎 Supported by static analysis

🏁 Script executed:

sed -n '215,335p' plugins/voxcpm/vision_agents/plugins/voxcpm/tts.py
sed -n '200,250p' plugins/voxcpm/tests/test_tts.py

Repository: GetStream/Vision-Agents

Length of output: 6961


🏁 Script executed:

sed -n '1,235p' plugins/voxcpm/vision_agents/plugins/voxcpm/tts.py
printf '\n--- focused tests and fixtures ---\n'
rg -n -C 8 "stop_audio|release|modelbest_server|stream_audio" plugins/voxcpm/tests/test_tts.py plugins/voxcpm/tests
printf '\n--- base lifecycle binding ---\n'
rg -n -C 6 "class .*TTS|async def stop_audio|async def close|def send_iter" vision_agents plugins/voxcpm/vision_agents | head -n 240

Repository: GetStream/Vision-Agents

Length of output: 36977


Cancel the pending request when stopping audio.

If stop_audio() runs while _request(text) is awaiting response headers, _response is still None, so the request continues. When it returns, the code assigns the response and enters _iter_sse_events() without checking _stop_event. The stream can then wait for the first SSE event before it observes the stop.

Track and cancel the pending request task. Treat its cancellation as a normal stop. Also check _stop_event immediately after _request() returns, close the response, clear _response, and return before iterating.

@lottshin lottshin closed this Sep 20, 2026
@lottshin lottshin reopened this Sep 21, 2026
@lottshin
lottshin marked this pull request as ready for review September 21, 2026 02:14

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
plugins/voxcpm/example/voxcpm_smoke.py (1)

14-14: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add explicit type annotations.

Annotate OUTPUT_PATH, first_chunk_at, and pcm_chunks. This keeps the example compliant with the repository typing rule.

As per coding guidelines: "Use type annotations everywhere."

Also applies to: 32-33

Source: Coding guidelines


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: GetStream/Vision-Agents/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: a5c3d30e-9198-46ce-961d-f459e3625505

📥 Commits

Reviewing files that changed from the base of the PR and between 638bf6c and 8d4cc85.

📒 Files selected for processing (6)
  • plugins/voxcpm/README.md
  • plugins/voxcpm/example/.env.example
  • plugins/voxcpm/example/README.md
  • plugins/voxcpm/example/__init__.py
  • plugins/voxcpm/example/pyproject.toml
  • plugins/voxcpm/example/voxcpm_smoke.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • plugins/voxcpm/README.md

Included review availability: Your plan provides up to 8 included reviews per hour; 7 remain after this review.

@lottshin

lottshin commented Oct 1, 2026

Copy link
Copy Markdown
Author

Hi, just following up on this PR. I've tested the plugin against ModelBest's hosted API for standard synthesis and reference-audio voice cloning, and included tests that run without credentials plus a runnable smoke example.

Is there anything else I should add or adjust to help with review? Happy to make changes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant