Add native Coqui XTTS v2 voice cloning - #520
DrewThomasson wants to merge 11 commits into
Conversation
Native Q8_0 synthesis demoGenerated end-to-end by this PR's xtts-v2-audiocpp-demo.mp4Independent Whisper-base.en transcription: “Hello? This is a native voice clone.” |
|
Same as usual if this passes your requirements this also needs to be merged https://huggingface.co/audio-cpp/audio.cpp-gguf/discussions/9 👍 |
Synthetic voice-cloning demo — alternate referenceThis is an explicitly labeled synthetic test using the user-supplied Spoken text: “This is a synthetic voice clone generated locally by audio dot C P P. It is not a real recording of David Attenborough.” xtts-v2-attenborough-demo.mp4 |
|
Hm that output is good but... perhaps quantization the model has lowered the voice cloning ability |
|
That model is VERY old ...so you can not expect would work as good as nowadays models. |
9bdd1d9 to
4d88768
Compare
|
No.... We use XttsV2 all the time in Ebook2audiobook and the voice cloning as well as pronunciation quality is MUCH better when using full precision some quality reduction has defiantly occurred when turning this into a gguf, perhaps I need to shoot for a higher than f16 |
|
I mean this was a output of out xttsV2 raw which sounds a lot more like David shortStory.mp4Hmmmm |
|
Did you ask your AI agent to compare audio quality during making a port? I think the easiest way i star from beginning and comparing quality on each step. |
Summary
Model artifacts
Hugging Face PR: https://huggingface.co/audio-cpp/audio.cpp-gguf/discussions/9
xtts-v2-q8_0.gguf: 616,802,944 bytes, SHA-256e07fd16d7365575fc9fe3946a70e9510f73f904983459d8eef060e91e62cfd4axtts-v2-f16.gguf: 976,290,816 bytes, SHA-2561db573bd06510f9fb8513913cc4a20b3b4340d23f918dbaf7dde695dda1ce055Both packages embed the tokenizer, config, CPML license, and this model spec, and both pass the conditioning → GPT → HiFiGAN native probe.
Validation
0.999999999993, RMSE1.66e-70.99999999983, RMSE6.31e-70.9999105, identical argmax token808Hello? This is a native voice clone.for inputHello, this is a native voice clone./varvs/private/varalias failure infun_asr_nano_assets_testThe XTTS model and generated output are governed by the Coqui Public Model License 1.0.0 and are non-commercial.