Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

@audio/neural

The opt-in ML lane: capture (dry/wet → formula) + pretrained weights.

Package What Status
@audio/neural-capture staged identifier: dry/wet pair → cheapest audiojs formula (IR → Hammerstein → W–H → W–H+tail → TCN) + null report; probe signal + IR-onset alignment rungs 1–4 shipped, tested vs sox + real NAM amp
@audio/neural-amp NAM .nam WaveNet playback, dependency-free shipped, real-capture verified
@audio/neural-synth sound matching: configure any knob-synth to a target sound (CMA-ES + mel-spectral loss) shipped, patch recovery verified
@audio/neural-timbre timbre recognition, gpu-font's recipe: one note → constant-Q features → 1.85M-parameter encoder → cosine against a catalog of preset vectors bound to the encoder's SHA-256; Yamaha DX7 first (synth-dx7 renders the labels), the named preset seeds neural-synth's match trained; 2,925 voices from held-out cartridges: 75.0% top-1, 92.6% top-5 clean, 56.4% / 79.0% augmented; ONNX = PyTorch within 2e-6 in Node and on WebGPU; int8 encoder (1.9 MB) and the dx7-factory and dx7-all catalogs (4-bit vectors, a point of top-1 or less) bundled, 4.4 MB with gzip: createMatcher() runs offline, tests too (ROM1A fixture: 29/32 top-1, 32/32 top-5 against dx7-factory); voice sources state no terms for this use, catalogs hold names and vectors only
@audio/neural-runtime one inference adapter: ONNX Runtime (onnxruntime-node / onnxruntime-web wasm + webgpu), worklet-ready, cached model fetch (~/.cache/audiojs/neural, $AUDIO_NEURAL_CACHE) shipped, 12/12 tests
@audio/neural-asr Whisper speech-to-text (transformers.js / ONNX Runtime), segment + word timestamps, cues bridge to @audio/subtitle shipped, real-inference verified (whisper-tiny, MIT weights)
@audio/neural-align forced alignment: pure-JS CTC trellis (torchaudio forced_align semantics) + wav2vec2-base-960h adapter (Apache-2.0) → word timestamps, enhanced LRC shipped, core tested without a model, adapter verified live
@audio/neural-separate stems: Open-Unmix-class spectrogram masking + multichannel Wiener EM (norbert port), Hybrid Transformer Demucs (STFT outside the graph), waveform path, through neural-runtime umxhq (MIT weights) and htdemucs exported and verified against the Python originals (same SDR on MUSDB18 previews); weights not bundled: scripts/export-*.py write them to the neural cache; Demucs weights research-only
@audio/neural-diarize who spoke when: @audio/vad regions → WavLM speaker embeddings (MIT) → agglomerative clustering → speaker segments, VTT <v> voice tags shipped, clustering tested without a model, adapter verified live
@audio/neural-tts text to speech: SpeechT5 (MIT) via transformers.js, sentence chunking, any output rate shipped, real-synthesis verified
@audio/neural-pitch monophonic pitch posterior and voicing: a 6,066-parameter transposition-equivariant network (PESTO's design, our own MIT weights: audiojs synth renders with exact f0, VocalSet CC BY 4.0), plain JS, no runtime; a stage 1 for @audio/pitch-pyin's HMM and note model (candidates option) trained; JS = PyTorch within 6e-5; 12 KB weights bundled; through pYIN's HMM beats YIN at 10 and 0 dB SNR and in a measured room (Vocadito 0 dB raw pitch 0.918 vs 0.499), slightly behind it on clean audio
@audio/neural-transcribe polyphonic notes with pitch bends: Basic Pitch (Spotify, ICASSP 2022; Apache-2.0 code and weights), note creation ported from the Python, frame posteriors exposed, through neural-runtime notes identical to Python basic-pitch on the same 22.05 kHz input; from 44.1 kHz all 969 of its notes within one frame, 968 identical (resample-sinc 1.2); bends in cents from the note's pitch, upstream: true for Basic Pitch's +33.3-cent offset (its #87); 230 KB model fetched to the neural cache on first use; package Apache-2.0 with NOTICE
@audio/neural-denoise speech enhancement: RNNoise ported to JS (BSD-3-Clause, weights bundled, 3.5 MB; int8 products in a 419-byte WebAssembly SIMD kernel), AudioWorklet with a constant 29.3 ms delay; DeepFilterNet3 through neural-runtime, 10 s chunks RNNoise bit-exact to upstream's portable C; DeepFilterNet3 128–134 dB SNR from Python, same scores; VoiceBank+DEMAND PESQ 3.16 (DeepFilterNet3 unlimited; 2.67 at its default 12 dB limit, which keeps room tone), 2.46 (RNNoise at its default 20 dB limit), 2.19 (best classical, wiener); limit: 0 gives upstream's output; DeepFilterNet3 weight terms unconfirmed upstream, fetched from the repo, never bundled

See research.md for theory (Boyd–Chua feasibility boundary, device-class ladder) and todo.md for the plan.

Policy (keeps the classical stance honest): classical tools never require this lane; weights are hosted separately and licensed-audited before any promise (many audio models are research-only — the freemium "premium ML weights" conflict in the site todo resolves here); deterministic pipelines stay classical. MIR's deferred ML tier (genre/mood/tags/separate) lands here when it lands.

About

Neural lane — umbrella for @audio/neural-* atoms (runtime adapter, denoise, amp, separation)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages