You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Star1 (1)You must be signed in to star a repository
About
An Open Source WhisprFlow Clone Speech-to-Text for macOS; eupports press to speak, toggle dictation, local inference, and configurable enhancement with OpenAI compatible models.
An Open Source WhisprFlow Clone Speech-to-Text for macOS; eupports press to speak, toggle dictation, local inference, and configurable enhancement with OpenAI compatible models.
Installation
pip install whosspr
Quick Start
# Start service and use keyboard shortcuts (below) to speak> whosspr start
11:25:20 [INFO] whosspr.cli: WhOSSpr Flow v0.1.0 starting...
11:25:20 [INFO] whosspr.config: No config file found, using defaults
╭───── Starting ──────╮
│ WhOSSpr Flow v0.1.0 │
│ │
│ Model: base │
│ Language: en │
│ Device: auto │
│ Enhancement: off │
│ │
│ Hold: ctrl+cmd+1 │
│ Toggle: ctrl+cmd+2 │
╰─────────────────────╯
Press Ctrl+C to stop.
11:25:20 [INFO] whosspr.controller: Loading Whisper model...
11:25:20 [INFO] whosspr.transcriber: Loading Whisper model 'base' on mps
11:25:22 [INFO] whosspr.transcriber: Model loaded
11:25:22 [INFO] whosspr.keyboard: Registered shortcut: ctrl+cmd+1 (hold)
11:25:22 [INFO] whosspr.keyboard: Registered shortcut: ctrl+cmd+2 (toggle)
11:25:22 [INFO] whosspr.keyboard: Keyboard listener started
11:25:22 [INFO] whosspr.controller: Dictation service started
Default Shortcuts
Shortcut
Action
Ctrl+Cmd+1 (hold)
Hold to record, release to transcribe
Ctrl+Cmd+2
Toggle dictation on/off
Permissions Setup
WhOSSpr requires two macOS permissions:
Permission
Purpose
How to Grant
Microphone
Record audio
System Preferences → Privacy → Microphone → Enable Terminal
Accessibility
Insert text
System Preferences → Privacy → Accessibility → Add Terminal
Recommendation: Start with base for a balance of speed and accuracy.
Requirements
Requirement
Details
OS
macOS 10.14+ (optimized for Apple Silicon)
Python
3.12+
Permissions
Microphone access, Accessibility access
RAM
2GB+ (more for larger models)
Usage
Starting the Service
whosspr start
Command-line Options
Option
Description
--model
Whisper model size (tiny/base/small/medium/large/turbo)
--language
Language code (e.g., en, es, fr)
--device
Device for inference (auto/cpu/mps/cuda)
--enhancement
Enable LLM text enhancement
--api-key
API key for enhancement
Examples
# Use small model with Spanish
whosspr start --model small --language es
# Use MPS (Apple Silicon GPU)
whosspr start --device mps
# Enable enhancement
whosspr start --enhancement --api-key sk-xxx
Text Enhancement
WhOSSpr can improve transcribed text using an OpenAI-compatible API:
# Using OpenAIexport OPENAI_API_KEY=sk-your-api-key
whosspr start --enhancement
# Using local LLM (Ollama)
whosspr start --enhancement \
--api-key ollama \
--api-base-url http://localhost:11434/v1
Troubleshooting
Problem
Solution
Permission denied
Run whosspr check, grant permissions, restart terminal
No audio input
Check microphone connection and permissions
Text not appearing
Verify Accessibility permission, try different app
Model download fails
Check internet, try --model tiny
High CPU/memory
Use smaller model, try --device mps on Apple Silicon
Development
Install from Source
git clone https://github.com/axsaucedo/WhOSSprFlow.git
cd WhOSSprFlow
pip install -e ".[dev]"
Running Tests
# All automated tests
pytest
# With coverage
pytest --cov=whosspr
# Manual E2E tests (interactive)
WHOSSPR_MANUAL_TESTS=1 pytest tests/test_e2e_manual.py -v -s
License
Apache 2.0
Architecture
This document describes the architecture of WhOSSpr Flow, an open-source speech-to-text application for macOS.
Overview
WhOSSpr Flow captures audio from the microphone, transcribes it using OpenAI Whisper, optionally enhances the text with an LLM, and inserts the result into the active application.
flowchart LR
A[User presses shortcut] --> B[Record audio]
B --> C[Transcribe with Whisper]
C --> D{Enhancement enabled?}
D -->|Yes| E[Enhance with LLM]
D -->|No| F[Insert text]
E --> F
Copy to clipboard, paste with Cmd+V, universal application support
enhancer.py
OpenAI-compatible API, API key resolution, custom prompts, grammar/punctuation improvement
permissions.py
Microphone access check, accessibility access check, pass/fail status
Data Flow
flowchart TD
subgraph Input
KB[Keyboard Shortcuts]
end
subgraph Controller
CTRL[Controller]
end
subgraph Recording
REC[Recorder]
AUDIO[(Audio numpy)]
end
subgraph Processing
TRANS[Transcriber]
ENH[Enhancer]
TEXT[(Text)]
end
subgraph Output
INS[Inserter]
end
KB --> CTRL
CTRL --> REC
REC --> AUDIO
AUDIO --> TRANS
TRANS --> TEXT
TEXT --> ENH
ENH --> INS
TEXT --> INS
Loading
State Machine
stateDiagram-v2
[*] --> IDLE
IDLE --> RECORDING: shortcut pressed
RECORDING --> IDLE: cancelled / too short
RECORDING --> PROCESSING: shortcut released
PROCESSING --> IDLE: complete / error
Loading
Design Principles
Principle
Description
Simple Modules
Each module has single responsibility, none exceeds ~300 lines
Controller imports others; other modules don't import each other (except config)
Direct Initialization
Components created when needed, no lazy patterns or factories
Callbacks for UI
Controller uses on_state/on_text/on_error callbacks to separate UI from logic
Threading Model
Component
Threading
sounddevice
Handles audio callback internally
pynput
Runs keyboard listener in separate thread
Processing
Sequential - no background threads for transcription
This simplifies debugging and reduces race conditions.
Configuration Schema
Section
Field
Type
Default
whisper
model_size
ModelSize
base
whisper
language
str
en
whisper
device
DeviceType
auto
shortcuts
hold_to_dictate
str
ctrl+cmd+1
shortcuts
toggle_dictation
str
ctrl+cmd+2
enhancement
enabled
bool
false
enhancement
api_key
str
""
enhancement
model
str
gpt-4o-mini
audio
sample_rate
int
16000
audio
channels
int
1
Test Structure
Test File
Coverage
test_config.py
Config loading/saving
test_recorder.py
Audio recording
test_transcriber.py
Whisper wrapper
test_keyboard.py
Shortcut parsing/handling
test_controller.py
Orchestration logic
test_enhancer.py
LLM enhancement
test_cli.py
CLI commands
test_e2e_manual.py
Interactive tests (require user)
Dependencies
Package
Purpose
sounddevice
Audio recording (no portaudio headers needed)
openai-whisper
Local speech-to-text
pynput
Global keyboard shortcuts
pyperclip
Clipboard operations
typer + rich
CLI framework
pydantic
Configuration validation
openai
LLM API client
About
An Open Source WhisprFlow Clone Speech-to-Text for macOS; eupports press to speak, toggle dictation, local inference, and configurable enhancement with OpenAI compatible models.