Convert articles and documents to high-quality audio using Qwen3-TTS.
Turn your markdown notes, articles, and text files into podcast-style audio you can listen to anywhere — powered by local AI inference on your GPU.
- 10 languages — Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian
- 9 premium voices — Male and female speakers across languages, dialects, and age ranges
- Multi-format support —
.md,.markdown,.txt,.rst,.text - Intelligent text cleaning — Strips markdown syntax, code blocks, links, and front-matter before synthesis
- Batch processing — Convert multiple files in a single command
- Chunked synthesis — Splits long text at sentence boundaries for consistent quality
- Rich CLI output — Progress bars, tables, and styled panels via Rich
- GPU accelerated — Runs on CUDA with auto-detection fallback to CPU
- Python 3.10+
- NVIDIA GPU with CUDA support (recommended) or CPU
- uv (recommended) or pip
git clone https://github.com/luis-codex/qwen-reader.git
cd qwen-reader
# Create virtual environment and install
uv venv
uv pip install -e .
# Install PyTorch with CUDA support (adjust cu128 to your CUDA version)
uv pip install torch torchaudio --index-url https://download.pytorch.org/whl/cu128
# (Optional, Linux only) Install FlashAttention 2 for ~2x faster inference
pip install -U flash-attn --no-build-isolationNote
The first run downloads the model (~3.5 GB) from HuggingFace. Subsequent runs load it from cache in ~30s.
Tip
FlashAttention 2 significantly reduces GPU memory usage and speeds up inference, but is only available on Linux with Ampere+ GPUs (RTX 30xx/40xx). Windows users can safely ignore the flash-attn warning — the manual PyTorch attention path works correctly, just slower.
# Install with dev dependencies (pytest, pytest-cov)
uv pip install -e ".[dev]"
# Run the test suite
python -m pytestAdd the virtual environment's Scripts (Windows) or bin (Linux/macOS) directory to your system PATH:
# Windows (PowerShell) — run once
$scriptsPath = "$PWD\.venv\Scripts"
[Environment]::SetEnvironmentVariable("Path", "$([Environment]::GetEnvironmentVariable('Path', 'User'));$scriptsPath", "User")# Linux / macOS — add to ~/.bashrc or ~/.zshrc
export PATH="/path/to/qwen-reader/.venv/bin:$PATH"Then open a new terminal and use qwen-reader from anywhere.
Usage: qwen-reader [OPTIONS] COMMAND [ARGS]
Commands:
read Convert one or more files to audio
speak Convert inline text to audio
speakers List available TTS voices
list List previously generated audio files
# Single file
qwen-reader read article.md
# Multiple files at once
qwen-reader read notes.txt report.md spec.rst
# Choose a voice and language
qwen-reader read article.md --speaker Ryan --lang English
# Custom output directory
qwen-reader read article.md --output-dir ./my-audio
# Custom output filename
qwen-reader read article.md --name my-podcastqwen-reader speak "Hello world, this is a test."
qwen-reader speak "Hola mundo, esto es una prueba." --lang Spanish --speaker VivianPre-generated audio samples across all 10 supported languages are available in the demos/ folder.
Listen to them to hear the quality before setting up the tool yourself!
| Sample | Language | Speaker | Voice |
|---|---|---|---|
demo_english.wav |
🇬🇧 English | Ryan | Dynamic male, strong rhythmic drive |
demo_spanish.wav |
🇪🇸 Spanish | Vivian | Bright, edgy young female |
demo_chinese.wav |
🇨🇳 Chinese | Serena | Warm, gentle young female |
demo_japanese.wav |
🇯🇵 Japanese | Ono_Anna | Playful female, light nimble timbre |
demo_korean.wav |
🇰🇷 Korean | Sohee | Warm female, rich emotion |
demo_french.wav |
🇫🇷 French | Aiden | Sunny American male |
demo_german.wav |
🇩🇪 German | Aiden | Sunny American male |
demo_italian.wav |
🇮🇹 Italian | Vivian | Bright, edgy young female |
demo_portuguese.wav |
🇧🇷 Portuguese | Ryan | Dynamic male |
demo_russian.wav |
🇷🇺 Russian | Aiden | Sunny American male |
To regenerate all demos:
pwsh scripts/generate_demos.ps1
qwen-reader speakers| Speaker | Voice Description | Native Language |
|---|---|---|
| Vivian | Bright, slightly edgy young female | Chinese |
| Serena | Warm, gentle young female | Chinese |
| Uncle_Fu | Seasoned male, low mellow timbre | Chinese |
| Dylan | Youthful Beijing male, clear natural | Chinese (Beijing Dialect) |
| Eric | Lively Chengdu male, husky brightness | Chinese (Sichuan Dialect) |
| Ryan | Dynamic male, strong rhythmic drive | English |
| Aiden | Sunny American male, clear midrange | English |
| Ono_Anna | Playful Japanese female, light nimble | Japanese |
| Sohee | Warm Korean female, rich emotion | Korean |
Tip
Each speaker can speak any of the 10 supported languages, but sounds best in their native language.
qwen-reader list 📂 ~/qwen-reader-audio
┌──────────────────────┬────────┬──────────────────┐
│ File │ Size │ Modified │
├──────────────────────┼────────┼──────────────────┤
│ 🔊 article.wav │ 4.2 MB │ 2026-04-25 22:10 │
│ 🔊 spoken_text.wav │ 0.1 MB │ 2026-04-25 21:52 │
├──────────────────────┼────────┼──────────────────┤
│ 2 files │ 4.3 MB │ │
└──────────────────────┴────────┴──────────────────┘
| Option | Short | Default | Description |
|---|---|---|---|
--speaker |
-s |
Aiden |
TTS voice to use |
--lang |
-l |
Auto |
Language (Auto, English, Chinese, Spanish, etc.) |
--instruct |
-i |
conversational | Style instruction for the TTS engine |
--output-dir |
-o |
~/qwen-reader-audio |
Output directory |
--name |
-n |
filename stem | Custom output filename (without extension) |
--device |
-d |
auto-detected | Compute device (cuda:0, cpu) — auto-detects CUDA |
--version |
-v |
— | Show version |
--help |
-h |
— | Show help |
| Extension | Processing |
|---|---|
.md, .markdown |
Strips YAML front-matter, code blocks, links, images, emphasis, headers |
.rst |
Strips directives, section underlines, inline markup |
.txt, .text |
Passed through as-is |
This project follows Clean Architecture with strict layer boundaries and a unidirectional dependency rule.
┌───────────────────────────────────────────────────────────┐
│ Interface Layer cli/ │
│ app.py · commands.py · options.py · rendering.py │
│ click + rich · args, output, exit codes │
├───────────────────────────────────────────────────────────┤
│ Use-Case Layer core/synthesis.py │
│ Orchestration · chunking → TTS → WAV assembly │
├──────────────────────┬────────────────────────────────────┤
│ Domain Layer │ Infrastructure Layer │
│ core/text.py │ core/model.py │
│ core/storage.py │ Model lifecycle, GPU management │
│ Pure transforms, │ torch, qwen_tts (deferred import) │
│ file listing │ │
│ stdlib only │ │
└──────────────────────┴────────────────────────────────────┘
Interface → Use-Case → Domain
→ Infrastructure → External Systems
No inner layer ever imports an outer layer. Core modules never call print(), sys.exit(), or import click/rich.
qwen_reader/
├── __init__.py # Package version
├── __main__.py # python -m qwen_reader entry
├── cli/
│ ├── __init__.py # Re-exports cli, main
│ ├── app.py # Click group, entry point, Windows UTF-8
│ ├── commands.py # read, speak, speakers, list commands
│ ├── options.py # Shared option decorators
│ └── rendering.py # Rich console, progress bars, result panel
└── core/
├── __init__.py # Docstring only — no re-exports
├── text.py # Domain: text cleaning & chunking
├── storage.py # Domain: audio file listing & output dir
├── model.py # Infrastructure: lazy model singleton
└── synthesis.py # Use-Case: audio generation orchestration
tests/
├── conftest.py # FakeModel stub + shared fixtures
├── test_text.py # Domain layer — no mocks, stdlib only
├── test_storage.py # Domain layer — file listing, no mocks
├── test_synthesis.py # Use-Case layer — mocked infrastructure
└── test_cli.py # Interface layer — click CliRunner
| Layer | Module(s) | Responsibility | Allowed deps | Forbidden |
|---|---|---|---|---|
| Interface | cli/app.py, cli/commands.py, cli/options.py, cli/rendering.py |
Parse args, render output, map exit codes | click, rich, Use-Case | torch, numpy, direct I/O |
| Use-Case | core/synthesis.py |
Orchestrate domain + infra into workflows | Domain, Infrastructure, numpy, soundfile | click, rich, print() |
| Domain | core/text.py, core/storage.py |
Pure text transforms, file listing | stdlib only (re, os, time) |
Any third-party package |
| Infrastructure | core/model.py |
External system lifecycle (model load, GPU) | torch, qwen_tts, stdlib | click, rich, domain logic |
| Mechanism | Example | Purpose |
|---|---|---|
@dataclass(frozen=True) |
ModelConfig, SynthesisResult |
Immutable snapshots passed between layers |
@dataclass (mutable) |
SynthesisConfig |
Aggregates user inputs before passing down |
| Callbacks | on_chunk(current, total, preview) |
Interface layer decides how to display progress |
Each architectural layer has its own test file with a tailored testing approach:
| File | Layer | Tests | Mocking | Speed |
|---|---|---|---|---|
test_text.py |
Domain | 37 | None — pure functions, stdlib only | < 1ms per test |
test_storage.py |
Domain | 9 | None — real temp files | < 1ms per test |
test_synthesis.py |
Use-Case | 17 | FakeModel stubs infrastructure |
< 100ms per test |
test_cli.py |
Interface | 19 | patch_model + CliRunner |
< 500ms per test |
# Quick run
python -m pytest
# With coverage report
python -m pytest --cov=qwen_reader --cov-report=term-missing
# Single layer
python -m pytest tests/test_text.py -v| Module | Stmts | Miss | Cover |
|---|---|---|---|
__init__.py |
1 | 0 | 100% |
cli/app.py |
22 | 2 | 91% |
cli/commands.py |
103 | 8 | 92% |
cli/options.py |
12 | 0 | 100% |
cli/rendering.py |
22 | 0 | 100% |
core/text.py |
52 | 0 | 100% |
core/storage.py |
35 | 0 | 100% |
core/synthesis.py |
74 | 3 | 96% |
core/model.py |
33 | 15 | 55% |
| Total | 358 | 30 | 92% |
Note
core/model.py coverage is lower by design — it wraps torch and qwen_tts which are mocked in tests. The remaining uncovered lines are the actual model loading path that requires a GPU.
| Layer | Minimum | Target | Actual |
|---|---|---|---|
| Domain | 90% | 100% | ✅ 100% |
| Use-Case | 80% | 90% | ✅ 96% |
| Interface | 60% | 80% | ✅ 92% |
| Category | Exception | CLI behavior |
|---|---|---|
| File not found | FileNotFoundError |
Print message, continue batch |
| Unsupported format | ValueError |
Print message, continue batch |
| Empty content | ValueError |
Print message, continue batch |
| Model failure | RuntimeError |
Print message, exit 1 |
| Synthesis failure | RuntimeError |
Print message, exit 1 |
| Code | Meaning |
|---|---|
0 |
All operations succeeded |
1 |
One or more operations failed |
2 |
CLI usage error (missing args, bad flags) |
Core modules never call sys.exit() — they raise typed exceptions. Only cli/commands.py converts exceptions to exit codes.
Configuration follows a strict priority order: CLI flags → Environment variables → Dataclass defaults.
| Variable | Default | Description |
|---|---|---|
QWEN_TTS_MODEL |
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice |
HuggingFace model ID |
QWEN_TTS_DEVICE |
auto (cuda:0 if available, else cpu) |
Inference device |
QWEN_TTS_OUTPUT_DIR |
~/qwen-reader-audio |
Default output directory |
Environment variables are read inside default_factory on config dataclasses — never scattered through application logic.
| Dependency | Purpose |
|---|---|
| qwen-tts | Qwen3-TTS model inference |
| torch | Deep learning runtime |
| soundfile | WAV file I/O |
| numpy | Audio array operations |
| click | CLI framework |
| rich | Terminal formatting |
| flash-attn (optional, Linux) | ~2× faster inference, less VRAM |
Dev dependencies (optional):
| Dependency | Purpose |
|---|---|
| pytest | Test framework |
| pytest-cov | Coverage reporting |
Every item must pass before merge to main.
- Interface layer (
cli/) imports no infrastructure/domain heavy deps - Domain layer (
core/text.py,core/storage.py) has zero third-party imports - Core modules never call
print(),sys.exit(), or importclick/rich - All cross-layer data flows via
@dataclassor callbacks - Heavy imports (torch, model libs) are deferred inside functions
-
pyproject.tomlhas[project.scripts]entry -
__main__.pyexists and delegates tocli:main -
__init__.pyexports only__version__ -
pip install -e .+qwen-reader --helpsucceeds
-
-h/--helpavailable on every group and command -
-v/--versionprints version and exits - All options have
show_default=Truewhere applicable - Success output is a structured Rich panel/table
- Error output uses
[red]❌prefix - Exit codes follow contract (0/1/2)
- Empty file / empty text raises
ValueError, not crash - Unsupported extension raises
ValueErrorwith list of valid types - File encoding fallback (UTF-8 → Latin-1) is implemented
- Windows UTF-8 stdout reconfiguration is present
- Batch processing continues on per-file errors
- Every public function has a docstring with Args/Returns/Raises
- Module-level docstring states purpose and dependency contract
- Type annotations on all public function signatures
- No
# type: ignorewithout adjacent comment explaining why - Constants use
UPPER_SNAKE_CASE, classesPascalCase
- Domain layer has unit tests with no mocks (46 tests)
- Use-Case layer has tests that mock infrastructure (17 tests)
- CLI layer has click
CliRunnertests (19 tests)
This project is licensed under the MIT License.