Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions .coderabbit.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,3 +12,23 @@ reviews:
drafts: true
base_branches:
- ".*"
path_instructions:
- path: "**/*.{cu,cuh,cpp,c,h,hpp}"
instructions: |
Prioritize correctness, performance and portability:
- Correctness: races, uninitialized or stale state, edge cases (empty, masked,
boundary sizes), numerical stability, and fallback paths that differ from the fast
path.
- Performance: regressions on hot paths, redundant work, unnecessary allocations or
synchronization, and caches or reuse that do not actually take effect.
- Portability: assumptions tied to one GPU, architecture, backend or platform; hardware
limits must be queried or guarded, with a working fallback.
- path: "ggml-patches/**"
instructions: |
Patches to the pinned ggml submodule. Apply the same correctness, performance and
portability checks; also check that op preconditions match what the kernels support
and that the patch series stays consistent.
- path: "{BENCHMARK.md,README.md,docs/**,app/bench*}"
instructions: |
Check that measurements are sound and that documented numbers, defaults and flags match
the code.
162 changes: 162 additions & 0 deletions BENCHMARK.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,162 @@
# Benchmarks

## Speech recognition (Nemotron Speech Streaming)

Nemotron Speech Streaming EN 0.6B, cache-aware streaming, LibriSpeech
test-clean (2,620 utterances, 5.4 h).

### GeForce RTX 4090

| Engine | Chunk | Compute per chunk (ms)<br>avg · p99 | Throughput (RTFX) | WER |
|---|:---:|:---:|:---:|:---:|
| NeMo-Speech.cpp (Q8_0) | 1.12 s | **3.4** · 4.6 | **238.4×** | 2.51% |
| NeMo (FP32) | 1.12 s | 21.2 · 23.1 | 49.3× | 2.32% |
| NeMo-Speech.cpp (Q8_0) | 160 ms | **2.3** · 2.8 | **66.6×** | 2.66% |
| NeMo (FP32) | 160 ms | 20.7 · 22.7 | 7.6× | 2.69% |

### DGX Spark (GB10)

| Engine | Chunk | Compute per chunk (ms)<br>avg · p99 | Throughput (RTFX) | WER |
|---|:---:|:---:|:---:|:---:|
| NeMo-Speech.cpp (Q8_0) | 1.12 s | **6.6** · 8.6 | **120.4×** | 2.50% |
| NeMo (FP32) | 1.12 s | 20.4 · 21.7 | 50.9× | 2.32% |
| NeMo-Speech.cpp (Q8_0) | 160 ms | **4.7** · 5.4 | **32.2×** | 2.64% |
| NeMo (FP32) | 160 ms | 18.9 · 19.7 | 8.4× | 2.69% |

### CPU

Intel Core i7-11700K, 8 threads for both engines, on a 100-utterance subset of test-clean (13.3 min).

| Engine | Chunk | Compute per chunk (ms)<br>avg · p99 | Throughput (RTFX) | WER |
|---|:---:|:---:|:---:|:---:|
| NeMo-Speech.cpp (Q8_0) | 1.12 s | **42.7** · 56.4 | **19.2×** | 3.10% |
| NeMo (FP32) | 1.12 s | 187.9 · 211.6 | 5.6× | 2.96% |
| NeMo-Speech.cpp (Q8_0) | 160 ms | **27.0** · 28.9 | **5.6×** | 3.00% |
| NeMo (FP32) | 160 ms | 103.5 · 112.5 | 1.5× | 3.14% |

## Speech synthesis (MagpieTTS)

MagpieTTS Multilingual 357M v2607, streaming.

### GeForce RTX 4090

| Engine | Time to first audio (ms)<br>avg · p99 | Inter-chunk latency (ms)<br>avg · p99 | Throughput (RTFX) |
|---|:---:|:---:|:---:|
| NeMo-Speech.cpp (Q8_0) | **8.6** · 9.4 | **2.9** · 3.1 | **59.8×** |
| NeMo (FP32) | 129.3 · 137.0 | 123.5 · 131.7 | 1.5× |

By input length:

| Input | Time to first audio (ms)<br>avg · p99 | Throughput (RTFX) | NeMo (FP32): time to first audio (ms)<br>avg · p99 | NeMo (FP32): throughput (RTFX) |
|---|:---:|:---:|:---:|:---:|
| Short (8 words) | **8.4** · 9.8 | **56.5×** | 126.0 · 127.9 | 1.5× |
| Medium (55 words) | **9.0** · 10.3 | **54.3×** | 127.1 · 131.1 | 1.5× |
| Long (258 words) | **10.1** · 11.0 | **56.0×** | 127.1 · 128.2 | 1.5× |

### DGX Spark (GB10)

| Engine | Time to first audio (ms)<br>avg · p99 | Inter-chunk latency (ms)<br>avg · p99 | Throughput (RTFX) |
|---|:---:|:---:|:---:|
| NeMo-Speech.cpp (Q8_0) | **16.8** · 18.1 | **5.8** · 25.5 | **30.0×** |
| NeMo (FP32) | 87.3 · 95.5 | 86.7 · 92.8 | 2.1× |

By input length:

| Input | Time to first audio (ms)<br>avg · p99 | Throughput (RTFX) | NeMo (FP32): time to first audio (ms)<br>avg · p99 | NeMo (FP32): throughput (RTFX) |
|---|:---:|:---:|:---:|:---:|
| Short (8 words) | **18.6** · 23.2 | **27.9×** | 86.6 · 87.3 | 2.2× |
| Medium (55 words) | **16.3** · 17.8 | **27.5×** | 90.4 · 91.8 | 2.0× |
| Long (258 words) | **18.4** · 19.2 | **28.3×** | 90.1 · 91.2 | 2.0× |

### CPU

Intel Core i7-11700K, 8 threads for both engines.

| Engine | Time to first audio (ms)<br>avg · p99 | Inter-chunk latency (ms)<br>avg · p99 | Throughput (RTFX) |
|---|:---:|:---:|:---:|
| NeMo-Speech.cpp (Q8_0) | **203.4** · 238.2 | **63.8** · 86.8 | **2.72×** |
| NeMo (FP32) | 659.8 · 679.1 | 665.0 · 761.6 | 0.28× |

By input length:

| Input | Time to first audio (ms)<br>avg · p99 | Throughput (RTFX) | NeMo (FP32): time to first audio (ms)<br>avg · p99 | NeMo (FP32): throughput (RTFX) |
|---|:---:|:---:|:---:|:---:|
| Short (8 words) | **191.0** · 208.2 | **2.58×** | 659.7 · 663.8 | 0.28× |
| Medium (55 words) | **199.2** · 233.1 | **2.60×** | 663.3 · 666.5 | 0.27× |
| Long (258 words) | **216.4** · 229.9 | **2.61×** | 686.5 · 697.9 | 0.27× |

## Devices

### GeForce RTX 4090

| | |
|---|---|
| System | GeForce RTX 4090 (128 SMs, 24 GB), Intel Core i7-11700K (16 threads), 128 GB RAM |
| Software | Ubuntu 24.04, NVIDIA driver 595.84, CUDA 13.2 |
| Build | Release, `-DCMAKE_CUDA_ARCHITECTURES=89` |
| NeMo | NeMo 3.1 (`main`), PyTorch 2.12.1 (CUDA 13.2), TF32 matmuls |

### CPU

Intel Core i7-11700K, the RTX 4090 host's CPU, with the same build and software, run with `--device cpu`.

| | |
|---|---|
| System | Intel Core i7-11700K (8 cores, 16 threads, AVX-512), 128 GB RAM |
| Threads | 8 for both engines: `--asr.backend.threads 8` / `--tts.threads 8`; NeMo `torch.set_num_threads(8)` |

### DGX Spark (GB10)

| | |
|---|---|
| System | GB10 GPU (48 SMs), 20-core Arm CPU, 128 GB unified memory |
| Software | Ubuntu 24.04, NVIDIA driver 580.126.09, CUDA 13.0 |
| Build | Release, `-DCMAKE_CUDA_ARCHITECTURES=121` |
| NeMo | NeMo 3.1, PyTorch 2.11 (CUDA 13.0), TF32 matmuls |

## Methodology

Precision is shown per engine. NeMo-Speech.cpp runs Q8_0 GGUF weights; NeMo runs
FP32, which was faster than BF16 here.

### Speech recognition

| | |
|---|---|
| Model | Nemotron Speech Streaming EN 0.6B |
| Chunks | 1.12 s (`--asr.streaming.rnnt_right_context 13`), 160 ms (`1`) |
| Metrics | Compute per chunk is the time to process one chunk; NeMo's excludes feature extraction, which its streaming example runs per utterance. Throughput is audio duration over wall time; NeMo's likewise excludes feature extraction. WER uses the Whisper English normalizer. |
| CPU subset | Every 26th utterance of test-clean (100 utterances, 40 speakers, 13.3 min) |
| Trials | RTX 4090 and CPU: NeMo-Speech.cpp average of three, NeMo one |

```bash
OUT_DIR=datasets scripts/asr/prepare_librispeech.sh test-clean

nemo-speech bench asr datasets/librispeech-test-clean -r --mode stream -c 1 -n 1 \
--model nemotron-speech-streaming-en-0.6b.q8_0.gguf \
--asr.streaming.rnnt_right_context 13 --save hyp/
```

Score `hyp/` against `datasets/librispeech-test-clean/transcripts.json` with the
Whisper English normalizer.

### Speech synthesis

| | |
|---|---|
| Models | MagpieTTS Multilingual 357M v2607, NeMo NanoCodec 22 kHz (F16) |
| Synthesis | `en-US`, default voice, seed 1, 22.05 kHz audio in 186 ms chunks (4 codec frames) |
| Inputs | The 10 LJSpeech sentences of the Riva TTS performance reports ([`ljs_audio_text_test_filelist_small.txt`](test_files/tts/ljs_audio_text_test_filelist_small.txt), 20 requests); by length, [`test_files/tts/bench`](test_files/tts/bench) (5 requests per input) |
| Metrics | Latencies at the client from the streaming audio callback; throughput is audio duration over wall time |
| Trials | NeMo-Speech.cpp: average of three; NeMo: one |

NeMo audio is decoded in 4-frame chunks as codes are produced.

```bash
nemo-speech bench tts --text-file test_files/tts/ljs_audio_text_test_filelist_small.txt \
--per-stream 20 --magpie-model magpie.q8_0.gguf --codec-model nanocodec.gguf \
--tokenizer-dir tokenizer/

nemo-speech bench tts test_files/tts/bench -n 5 \
--magpie-model magpie.q8_0.gguf --codec-model nanocodec.gguf --tokenizer-dir tokenizer/
```
4 changes: 4 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -259,6 +259,10 @@ if(GGML_METAL)
list(APPEND GGML_DEPENDENCIES ggml-metal)
endif()

# Tiled CPU GEMM (llamafile tinyBLAS) for F16/F32/Q8_0 matmuls. Stock ggml leaves it off; without
# it CPU matmuls run one dot product per output.
set(GGML_LLAMAFILE ON CACHE BOOL "ggml: use llamafile SGEMM")

add_subdirectory(ggml EXCLUDE_FROM_ALL)

# ggml is an implementation dependency, so install only the runtime libraries
Expand Down
30 changes: 27 additions & 3 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,22 @@

We welcome external contributions to NeMo-Speech.cpp.

## AI usage

You may use AI tools as assistants for code, but the contribution must be yours. Pull request descriptions, issues and review replies must be written by you.

- **Disclose it.** If AI was used, mention in the pull request in what capacity it was used.
- **Write it yourself.** Issues, pull request descriptions and review replies
must be your own words, not AI output.
- **Review it.** Review every line of code before submitting. You should understand and be able to
explain the design and any line of code when a reviewer asks.
- **Verify it.** Build and run the relevant tests and benchmarks yourself, and
report the commands and results you actually ran.
- **Keep it focused.** Don't include unrelated refactors, reformatting or
speculative changes that a tool added along the way.

Pull requests that don't follow these guidelines may be closed without review.

## Development checks

Follow the [source-build guide](docs/build.md) for prerequisites and submodules.
Expand Down Expand Up @@ -40,9 +56,13 @@ may be held for provenance and license review before acceptance.

## Signing off your work

Every commit must be signed off. The sign-off certifies that you have the right
to submit the contribution under the license indicated in the file. Commits
without a `Signed-off-by` line will not be accepted.
Every commit must be signed off. By adding a `Signed-off-by` line to a commit,
you agree to the [Developer Certificate of Origin (DCO)
1.1](#developer-certificate-of-origin) reproduced below: you certify that you
wrote the contribution or otherwise have the right to submit it under the
project's open source license, and you acknowledge that the contribution and
your sign-off are public and kept permanently. Commits without a
`Signed-off-by` line will not be accepted.

Use Git's `--signoff` (or `-s`) option:

Expand All @@ -56,6 +76,10 @@ This appends:
Signed-off-by: Your Name <your@email.com>
```

### Developer Certificate of Origin

Signing off certifies that at least one of (a), (b), or (c) below applies to
your contribution, and that you agree to (d).
The full, unmodified [Developer Certificate of Origin
1.1](https://developercertificate.org/) follows:

Expand Down
38 changes: 34 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,36 @@
| Full-duplex voicechat | [Nemotron Labs VoiceChat](https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B), including realtime audio, transcripts, and tool calling |
| Speech processing | [Silero VAD](https://github.com/snakers4/silero-vad), punctuation and capitalization, endpointing, text normalization, and subtitles |

## Performance

NeMo-Speech.cpp is blazing fast and built for real-time streaming. Speech recognition
transcribes audio in 160 ms chunks up to 67× faster than real time, and speech synthesis
generates speech up to 60× faster than real time with its first audio in under 10 ms. Both stay
faster than real time even on a CPU.

Streaming speech recognition with Nemotron Speech Streaming 0.6B (Q8_0), 160 ms chunks:

| Device | Latency per chunk | Throughput | Speedup over NeMo (FP32) |
|---|:---:|:---:|:---:|
| GeForce RTX 4090 | **2.3 ms** | **67× real time** | **8.8×** |
| CPU | **27 ms** | **6× real time** | **3.7×** |

Streaming speech synthesis with MagpieTTS Multilingual (Q8_0), 186 ms audio chunks:

| Device | Time to first audio | Inter-chunk latency | Throughput | Speedup over NeMo (FP32) |
|---|:---:|:---:|:---:|:---:|
| GeForce RTX 4090 | **9 ms** | **3 ms** | **60× real time** | **40×** |
| CPU | **203 ms** | **64 ms** | **2.7× real time** | **9.7×** |

See [BENCHMARK.md](BENCHMARK.md) for the methodology and more results.

## Installation

> [!IMPORTANT]
> **For the best performance and the latest features, build natively from source.** A native
> build is compiled for your machine, and release tags can be out of sync with the
> `main` branch. See [Build from source](#build-from-source).

Install the `nemo-speech` CLI for the detected platform and backend:

On Linux or macOS, run:
Expand All @@ -45,10 +73,12 @@ irm https://github.com/NVIDIA/NeMo-Speech.cpp/raw/main/scripts/install.ps1 | iex
Open a new PowerShell window after installation so the updated user `PATH`
takes effect.

The installer prefers a verified native release and falls back to a source
build when an artifact is unavailable. A source build requires Git, CMake 3.26
or newer, Ninja, a C++17 compiler, SentencePiece development files, and the
toolchain required by the selected backend, if any. See
The installer downloads the prebuilt archive for the latest release, checks it
against the SHA-256 checksum published with the release, and builds from source
when no archive is available for your platform. **Pass `--source` (`-Source` on
Windows) to always build from the `main` branch.** A source build requires Git,
CMake 3.26 or newer, Ninja, a C++17 compiler, SentencePiece development files,
Comment thread
pskrunner14 marked this conversation as resolved.
and the toolchain required by the selected backend, if any. See
[Installation](docs/install.md) for platform-specific prerequisites and
options.

Expand Down
16 changes: 8 additions & 8 deletions app/CMakeLists.txt
Original file line number Diff line number Diff line change
@@ -1,8 +1,10 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0

find_package(Threads REQUIRED)
add_executable(nemo_speech_cli
main.cpp
bench.cpp
cli_util.cpp
doctor.cpp
model.cpp
Expand All @@ -15,21 +17,19 @@ target_link_libraries(nemo_speech_cli PRIVATE
nemo_speech_engine_registry
ggml
ggml-base
Threads::Threads
)
target_include_directories(nemo_speech_cli PRIVATE ${CMAKE_SOURCE_DIR}/include)
if(WIN32)
target_link_libraries(nemo_speech_cli PRIVATE shell32)
endif()

if(NEMO_SPEECH_BUILD_ASR)
find_package(Threads REQUIRED)
target_sources(nemo_speech_cli PRIVATE
bench.cpp
bench_asr.cpp
transcribe.cpp)
target_compile_definitions(nemo_speech_cli PRIVATE NEMO_SPEECH_CLI_ASR=1)
target_link_libraries(nemo_speech_cli PRIVATE
nemo_speech_asr
Threads::Threads)
target_link_libraries(nemo_speech_cli PRIVATE nemo_speech_asr)
if(NEMO_SPEECH_BUILD_MIC_CAPTURE)
if(NOT EXISTS "${CMAKE_SOURCE_DIR}/llama.cpp/vendor/miniaudio/miniaudio.h")
message(FATAL_ERROR
Expand Down Expand Up @@ -85,12 +85,12 @@ if(NEMO_SPEECH_WITH_NORM)
endif()
if(NEMO_SPEECH_BUILD_DIAR)
target_compile_definitions(nemo_speech_cli PRIVATE NEMO_SPEECH_CLI_DIAR=1)
target_sources(nemo_speech_cli PRIVATE diarize.cpp)
target_sources(nemo_speech_cli PRIVATE bench_diarize.cpp diarize.cpp)
target_link_libraries(nemo_speech_cli PRIVATE nemo_speech_asr)
endif()
if(NEMO_SPEECH_BUILD_TTS)
target_compile_definitions(nemo_speech_cli PRIVATE NEMO_SPEECH_CLI_TTS=1)
target_sources(nemo_speech_cli PRIVATE synthesize.cpp)
target_sources(nemo_speech_cli PRIVATE bench_tts.cpp synthesize.cpp)
target_link_libraries(nemo_speech_cli PRIVATE nemo_speech_tts_magpietts)
if(NEMO_SPEECH_TTS_WITH_JA)
target_compile_definitions(nemo_speech_cli PRIVATE NEMO_SPEECH_CLI_TTS_JA=1)
Expand All @@ -101,7 +101,7 @@ if(NEMO_SPEECH_BUILD_TTS)
endif()
if(NEMO_SPEECH_BUILD_NMT)
target_compile_definitions(nemo_speech_cli PRIVATE NEMO_SPEECH_CLI_NMT=1)
target_sources(nemo_speech_cli PRIVATE translate.cpp)
target_sources(nemo_speech_cli PRIVATE bench_translate.cpp translate.cpp)
target_link_libraries(nemo_speech_cli PRIVATE nemo_speech_nmt)
endif()

Expand Down
Loading
Loading