Claude:"On an RTX A6000 (Windows Server 2025, portable Python 3.10, cu121, TTS_BF16=on), Turbo's T3 loop ran at ~22 tok/s, below the 25 tok/s needed for realtime. The loop is CPU-bound: each token issues several hundred small GPU ops from Python, and torch.compile isn't practical in the portable Windows install.
I [Claude] wrote an opt-out add-on (t3_graph.py + an 11-line hook in chatterbox/models/t3/t3.py). It reimplements the GPT-2 decode step with a preallocated KV cache and records one token step, including sampling (temperature/top-k/top-p/repetition penalty, same order as stock), as a CUDA graph that is replayed per token. No Triton or MSVC needed. Result: ~385 tok/s, ~9–10× realtime end to end, with no audible quality change. Logits match the HF model within bf16 noise (max relative diff 0.0076 over 33 positions).
Safety: Turbo only, batch 1, CUDA only; everything else goes to the unmodified stock loop. Each new decoder is self-checked against the HF model before use, and a mismatch, exception, too-long prompt, or low VRAM falls back to stock, freeing its memory. It can be disabled with T3_CUDA_GRAPH=off. The RNG state is preserved, so seeds are reproducible (though they don't reproduce stock's exact audio).
Caveats: ~200 MB extra VRAM; T3 is serialized across concurrent requests; it patches a file inside the chatterbox-tts package, so it would need to live in the server as a monkeypatch or go upstream to Resemble. Not yet tested on Linux, other GPUs, or under concurrent load."
DiBianoR: This is SO much faster than stock that I felt like I had to bring it to your attention - literally an order of magnitude. I've attached the code so you can look over / vet it if you want to.
Happy to open a PR if you're interested.
in .\python_embedded\Lib\site-packages\chatterbox:
Modified: tts.py
Added: tts_turbo.py
Claude:"On an RTX A6000 (Windows Server 2025, portable Python 3.10, cu121, TTS_BF16=on), Turbo's T3 loop ran at ~22 tok/s, below the 25 tok/s needed for realtime. The loop is CPU-bound: each token issues several hundred small GPU ops from Python, and torch.compile isn't practical in the portable Windows install.
I [Claude] wrote an opt-out add-on (t3_graph.py + an 11-line hook in chatterbox/models/t3/t3.py). It reimplements the GPT-2 decode step with a preallocated KV cache and records one token step, including sampling (temperature/top-k/top-p/repetition penalty, same order as stock), as a CUDA graph that is replayed per token. No Triton or MSVC needed. Result: ~385 tok/s, ~9–10× realtime end to end, with no audible quality change. Logits match the HF model within bf16 noise (max relative diff 0.0076 over 33 positions).
Safety: Turbo only, batch 1, CUDA only; everything else goes to the unmodified stock loop. Each new decoder is self-checked against the HF model before use, and a mismatch, exception, too-long prompt, or low VRAM falls back to stock, freeing its memory. It can be disabled with T3_CUDA_GRAPH=off. The RNG state is preserved, so seeds are reproducible (though they don't reproduce stock's exact audio).
Caveats: ~200 MB extra VRAM; T3 is serialized across concurrent requests; it patches a file inside the chatterbox-tts package, so it would need to live in the server as a monkeypatch or go upstream to Resemble. Not yet tested on Linux, other GPUs, or under concurrent load."
DiBianoR: This is SO much faster than stock that I felt like I had to bring it to your attention - literally an order of magnitude. I've attached the code so you can look over / vet it if you want to.
Happy to open a PR if you're interested.
in .\python_embedded\Lib\site-packages\chatterbox:
Modified: tts.py
Added: tts_turbo.py