Skip to content

About

Official Repo for "AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach"

Topics

Resources

Stars

60 stars

Watchers

2 watching

Forks

Repository files navigation

AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach

Paper Project Page VibeXASR Desktop App
AASR-Bench on Hugging Face AASR-Bench on ModelScope | Refiner on Hugging Face Refiner on ModelScope

News

  • 2026-08-06: The Hugging Face Refiner repository now includes the mlx-int4 and onnx-int4 Refiner variants. Thanks to xcc-zach.
  • 2026-07-30: Paper, code, AASR-Bench, and Refiner checkpoint are released.

Features

AgenticASR turns speech into clean written text while preserving the speaker's final intent. It removes disfluencies, resolves self-corrections, normalizes written form, and can revise previously emitted text when later speech adds new evidence.

  • Bilingual: AgenticASR supports both English and Chinese speech-to-clean-text refinement.
  • ASR-agnostic: The Refiner is decoupled from the recognizer and can be attached to any ASR frontend that produces text hypotheses.
  • Online + offline: Refine a complete transcript once, or continually replace a bounded active span as speech arrives.
  • AASR-Bench: 917 samples and 6,637 atomic rubrics covering Content, Format, Filter, and Rephrase.

Bilingual Demo

English Demo 中文演示
Play the English AgenticASR demo
▶ Play English demo
播放 AgenticASR 中文演示
▶ 播放中文演示

We evaluate the Refiner with Qwen3-ASR and Whisper. The packaged Windows/macOS application is available from the VibeXASR product page. This repository contains the research code and reproducible core implementation.

Installation

Create a separate environment

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip

Install the dependencies

python -m pip install -r requirements.txt
python -m pip install -r system/requirements.txt

Refiner training additionally requires LLaMA Factory. Install it separately before running llamafactory-cli train.

Inference

1. Batch Refiner inference

The Refiner accepts JSONL output from any ASR frontend. Each record must contain source_record_id and output.raw_text.

python experiments/scripts/postprocess_asr.py \
  /path/to/asr_output.jsonl \
  /path/to/refined_output.jsonl \
  --model /path/to/refiner-checkpoint

The checkpoint uses the Refiner system prompt defined in the inference and training code. See experiments/README.md for the input contract and backend options.

2. Streaming AgenticASR

The streaming system combines VAD, an online sherpa-onnx ASR frontend, stable text chunking, and a default K=3 sliding-window Refiner. The current local backend uses MLX-LM on macOS.

bash system/download_vad.sh /path/to/models
python -m system.live_asr \
  --wav /path/to/example.wav \
  --asr-dir /path/to/models/asr \
  --refiner /path/to/models/refiner-mlx

Use --identity-refiner only for ASR and chunking diagnostics. See system/README.md for model preparation and runtime options.

Training

1. Generate Refiner training data

Start an OpenAI-compatible vLLM service, or configure an OpenRouter-compatible service:

export VLLM_MODEL_NAME=/path/to/gemma-4-31b-it
export VLLM_BASE_URL=http://127.0.0.1:8000/v1
python run_pipeline.py

The five-stage pipeline generates Oral/Clean pairs, simulates ASR hypotheses, performs semantic quality control, and deduplicates the final records. Outputs are written to data/final/. See pipeline/README.md.

2. Export SFT data

Export the finalized records into LLaMA Factory format:

python pipeline/scripts/export_sft.py \
  --inputs /path/to/AgenticASR/data/final/train.jsonl \
  --train-output /path/to/llamafactory-data/train_sft.json \
  --val-output /path/to/llamafactory-data/val_sft.json

3. Fine-tune the Refiner

Register those files in the LLaMA Factory dataset directory at /path/to/llamafactory-data/dataset_info.json:

{
  "refiner_train": {"file_name": "train_sft.json"},
  "refiner_val": {"file_name": "val_sft.json"}
}

Then edit a copy of refiner.yaml and replace every machine-specific path:

model_name_or_path: /path/to/base-model
dataset_dir: /path/to/llamafactory-data
dataset: refiner_train
eval_dataset: refiner_val
output_dir: /path/to/refiner-output

Launch training with the edited configuration:

llamafactory-cli train /path/to/refiner.yaml

Evaluation

Download rubric.json from ModelScope or Hugging Face, then run the judge in experiments/scripts/main.py with --rubric /path/to/rubric.json. See experiments/README.md.

Repository Guide

  • pipeline/: synthetic training-data generation and SFT export.
  • experiments/: Refiner inference and AASR-Bench evaluation.
  • system/: streaming VAD, ASR, chunk management, and online refinement.

Citation

@misc{jiang2026agenticasrrefiningspeechrecognition,
      title={AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach},
      author={Zixuan Jiang and Binghao Qiang and Jiaying Chi and Yanqiao Zhu and Kai Yu and Xie Chen},
      year={2026},
      eprint={2607.28175},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2607.28175},
}

Acknowledgements

We thank the authors and contributors of LLaMA Factory, MiniCPM, X-ASR, Gemma, and Qwen3-ASR for their great work and open-source contributions.

License

This project is released under the Apache License 2.0.

About

Official Repo for "AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach"

Topics

Resources

Stars

60 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages