Local voice cloning and speech generation with Pocket TTS Raven.
voice-lab is the standalone home of the Pocket TTS Raven voice pipeline that
was previously embedded in another project. It deals only with cloning a voice
from a consented reference recording and generating speech audio with the
managed local Raven runtime.
A Podcast style Demo rendered by generate_manuscript.py from examples/welcome-call-manuscript.json
git submodule update --init --recursive
python3 tools/setup_raven.py --accept-model-terms
python3 tools/setup_raven.py --check-installationThe installer downloads hash-pinned model inputs only after the user accepts
the model terms, rebuilds the ONNX graph locally, builds Raven from source in
isolation, and publishes a schema-v2 installation under
~/Library/Caches/voice-lab/pocket-tts-raven/ only after a successful server
health check.
Supported platforms: macOS arm64 and Linux x86_64.
The two built-in reference voices live in voices/. For batch work, use the
interactive helper — pick a voice, then type sentences one at a time and get
out/1-out.mp3, out/2-out.mp3, ... (numbering continues across runs; the
intermediate PCM16 WAV is kept next to each MP3):
python3 make_sentences.pyThe prompt shows the active voice ([deja-thoris] sentence>). Type a sentence
to synthesize it with that voice, or enter a voice number (e.g. 2) to switch
speakers mid-discussion; an empty line quits.
For a single line directly:
python3 generate_voice.py \
--voice ./voices/speaker.wav \
--text "Hello from voice-lab." \
--output ./out/speaker-hello.wavOptions: --precision int8|fp32, --temperature, --lsd-steps, --threads,
--keep-commas, --voice-id. The voice reference must be an existing WAV or
MP3 file with an explicit lawful consent basis; see
docs/pocket-tts-raven-distribution-review.md for model attribution, consent
requirements, and the prebuilt-release gate.
For scripted, multi-voice work, define the voices and lines in one JSON
manuscript and render it in one run. The JSON format is documented in the
manuscript-voice skill (.agents/skills/manuscript-voice/SKILL.md), and
examples/welcome-call-manuscript.json shows a working example.
python3 generate_manuscript.py manuscript.json --dry-run # validate + plan
python3 generate_manuscript.py manuscript.json --concat out/manuscript.mp3This writes one numbered MP3 per line (001-narrator.mp3, ...) and can join
all lines into a single file with --concat.
voice_providers/pocket_tts_raven.py— persistent loopback HTTP provider that starts, health-checks, and drives the local Raven runtime; validates manifests by path, size, and SHA-256; stages and conditions voice references.voice_providers/__init__.py—VoiceRequest/VoiceResult/VoiceProvidercontracts.audio.py— WAV inspection, PCM16 conditioning/normalization, text normalization, and synthesis fingerprints.generate_voice.py— one-shot CLI: reference WAV/MP3 + text → PCM16 WAV.make_sentences.py— interactive batch helper: pick a voice, prompt for sentences, write numbered MP3 takes via ffmpeg.generate_manuscript.py— batch CLI: JSON manuscript (named voices + ordered lines) → one MP3 per line, optional--concatjoin..agents/skills/manuscript-voice/SKILL.md— format specification for manuscript JSON files, discoverable as a project skill.examples/welcome-call-manuscript.json— example two-voice manuscript.voices/— built-in in-house reference voices (Reginald Ashworth, Deja Thoris); seevoices/README.md.tools/setup_raven.py— hash-pinned local builder/installer for Raven and its ONNX models (--accept-model-termsrequired).tools/compare_raven_profiles.py,tools/benchmark_raven_adapters.py— profile comparison and adapter benchmarking utilities.tests/test_pocket_tts_raven.py— provider, installer-validation, and adapter test suite.vendor/pocket-tts-raven— Raven source, pinned as a Git submodule.docs/— distribution review and adapter benchmark documentation.
python3 tools/benchmark_raven_adapters.py \
--voice ./voices/speaker.wav \
--ffi-library /path/to/libpocket_tts.dylib
python3 tools/compare_raven_profiles.py \
--voice speaker=./voices/speaker.wavThe persistent loopback HTTP adapter is the production default; direct ctypes
FFI remains an opt-in benchmark path.
python3 -m unittest discover -s testsProject code is MPL 2.0 (LICENSE). The Pocket TTS model is CC BY 4.0
(kyutai/pocket-tts, ONNX export KevinAHM/pocket-tts-onnx); third-party
dependency notices are in THIRD_PARTY_NOTICES.md and the Raven submodule.
Model files, transformed ONNX artifacts, prebuilt binaries, and voice
references are not distributed by this repository.