- 2026-08-06: The Hugging Face Refiner repository now includes the
mlx-int4andonnx-int4Refiner variants. Thanks to xcc-zach. - 2026-07-30: Paper, code, AASR-Bench, and Refiner checkpoint are released.
AgenticASR turns speech into clean written text while preserving the speaker's final intent. It removes disfluencies, resolves self-corrections, normalizes written form, and can revise previously emitted text when later speech adds new evidence.
- Bilingual: AgenticASR supports both English and Chinese speech-to-clean-text refinement.
- ASR-agnostic: The Refiner is decoupled from the recognizer and can be attached to any ASR frontend that produces text hypotheses.
- Online + offline: Refine a complete transcript once, or continually replace a bounded active span as speech arrives.
- AASR-Bench: 917 samples and 6,637 atomic rubrics covering Content, Format, Filter, and Rephrase.
| English Demo | 中文演示 |
|---|---|
▶ Play English demo |
▶ 播放中文演示 |
We evaluate the Refiner with Qwen3-ASR and Whisper. The packaged Windows/macOS application is available from the VibeXASR product page. This repository contains the research code and reproducible core implementation.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pippython -m pip install -r requirements.txt
python -m pip install -r system/requirements.txtRefiner training additionally requires LLaMA Factory. Install it separately before running llamafactory-cli train.
The Refiner accepts JSONL output from any ASR frontend. Each record must contain source_record_id and output.raw_text.
python experiments/scripts/postprocess_asr.py \
/path/to/asr_output.jsonl \
/path/to/refined_output.jsonl \
--model /path/to/refiner-checkpointThe checkpoint uses the Refiner system prompt defined in the inference and training code. See experiments/README.md for the input contract and backend options.
The streaming system combines VAD, an online sherpa-onnx ASR frontend, stable text chunking, and a default K=3 sliding-window Refiner. The current local backend uses MLX-LM on macOS.
bash system/download_vad.sh /path/to/models
python -m system.live_asr \
--wav /path/to/example.wav \
--asr-dir /path/to/models/asr \
--refiner /path/to/models/refiner-mlxUse --identity-refiner only for ASR and chunking diagnostics. See system/README.md for model preparation and runtime options.
Start an OpenAI-compatible vLLM service, or configure an OpenRouter-compatible service:
export VLLM_MODEL_NAME=/path/to/gemma-4-31b-it
export VLLM_BASE_URL=http://127.0.0.1:8000/v1
python run_pipeline.pyThe five-stage pipeline generates Oral/Clean pairs, simulates ASR hypotheses, performs semantic quality control, and deduplicates the final records. Outputs are written to data/final/. See pipeline/README.md.
Export the finalized records into LLaMA Factory format:
python pipeline/scripts/export_sft.py \
--inputs /path/to/AgenticASR/data/final/train.jsonl \
--train-output /path/to/llamafactory-data/train_sft.json \
--val-output /path/to/llamafactory-data/val_sft.jsonRegister those files in the LLaMA Factory dataset directory at /path/to/llamafactory-data/dataset_info.json:
{
"refiner_train": {"file_name": "train_sft.json"},
"refiner_val": {"file_name": "val_sft.json"}
}Then edit a copy of refiner.yaml and replace every machine-specific path:
model_name_or_path: /path/to/base-model
dataset_dir: /path/to/llamafactory-data
dataset: refiner_train
eval_dataset: refiner_val
output_dir: /path/to/refiner-outputLaunch training with the edited configuration:
llamafactory-cli train /path/to/refiner.yamlDownload rubric.json from ModelScope or Hugging Face, then run the judge in experiments/scripts/main.py with --rubric /path/to/rubric.json. See experiments/README.md.
pipeline/: synthetic training-data generation and SFT export.experiments/: Refiner inference and AASR-Bench evaluation.system/: streaming VAD, ASR, chunk management, and online refinement.
@misc{jiang2026agenticasrrefiningspeechrecognition,
title={AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach},
author={Zixuan Jiang and Binghao Qiang and Jiaying Chi and Yanqiao Zhu and Kai Yu and Xie Chen},
year={2026},
eprint={2607.28175},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2607.28175},
}We thank the authors and contributors of LLaMA Factory, MiniCPM, X-ASR, Gemma, and Qwen3-ASR for their great work and open-source contributions.
This project is released under the Apache License 2.0.

