A face-to-face chat GUI for LLMs. Claude (or any local Ollama model) speaks through a real animated avatar with a natural neural voice, sees you via the webcam, listens through your microphone, and remembers you across sessions.
You type or speak; the avatar replies in voice and on screen. A persistent per-persona cognition layer (episodic + semantic + emotional memory + relationship progression + character sheet) is injected into the system prompt every turn — so the same memories steer Anthropic, Ollama, or the demo engine identically. Every panel is a detachable Qt dock; every operation has a CLI + HTTP control surface so Claude Code can drive the GUI from outside.
- Eight characters, one engine-agnostic memory — pick your default Claude (max-capability, ICT 3D face, British female voice) or swap to one of seven authored personas: Iris the neuroscience PhD, Bayard the retired guitarist, Niko the indie dev, Soraya the ER nurse, Theo the bookshop owner, playful cartoon Claude, or Iris on x-ray. Each has a full character sheet (Big Five traits, backstory, catchphrases, goals, preferred voice) and their own persistent memory file.
- Natural neural voice — Kokoro-onnx local TTS runs in real time on
Apple Silicon CPU. 54 voices; each character has a defaulted voice
(
bf_emma,bm_george,af_sky, …). Pyttsx3 fallback if Kokoro isn't installed. - Real STT → LLM — sounddevice mic → silero-vad → faster-whisper → bridged into the chat panel. Mic is auto-muted during TTS so the avatar doesn't echo itself; 🎤 Hold to talk button interrupts the avatar and bypasses the echo gate.
- Test mode: two bots conversing in character — pick a partner
persona and engine (canned / Ollama / Anthropic / demo). Each bot
has its own
Conversation+ character; replies route throughchat_panel.append_external_messageso they don't re-trigger the main client. - Live engine swap — change between Anthropic / Ollama / demo from
the config dialog or CLI without restarting the app. Status pill
reflects what's actually driving conversation (green = Anthropic,
blue = Ollama, grey = demo,
⇄prefix = test-mode override). - Detachable layout — drag any panel out as a floating window, tab two panels together, hide one, save the arrangement with Cmd-Shift-Y, reset to defaults with Cmd-Shift-L. Persists via QSettings.
- Persona editor — Tools → Edit personas… (Cmd-Shift-I). Live edit any character's identity, traits, backstory, catchphrases, goals; save rebinds the running avatar.
- Driveable from Claude Code —
tools/faceview_drive.pylaunches + controls the GUI,tools/faceview_monitor.pyreads state. Both talk to a 127.0.0.1 FastAPI control plane; an MCP server adapter exposes the same surface as native Claude Code tools.
The default avatar is max-capability Claude on the USC ICT-FaceKit
photo-real 3D head with an x-ray glow shader, speaking through the
bf_emma Kokoro voice (British female). Mood pill, status pills, and
the chat history all update in real time.
Left: max-capability Claude (default boot persona — ICT face, bf_emma voice).
Right: Iris, neuroscience PhD student, with a different voice (af_nicole) and
her own conversation memory under .faceview/memory/ict_xray.json.
Playful Claude on the stylised cartoon face, Bayard the retired classical guitarist, and Soraya the ER nurse. Each has a distinct backstory, catchphrases, Big Five trait profile, and Kokoro voice.
Test mode replaces the user-side webcam with a second avatar and routes
both sides through real LLMs (or canned seed prompts). Pick the partner
persona + engine + model from the config dialog; each bot uses its
character's narrate_identity() as its system prompt and grows its own
in-memory Conversation history.
Test mode running two Llama-3 bots: Theo (bookshop owner) in the camera window
chatting with max-capability Claude in the avatar window about James Ellroy. Both
replies are real Ollama output. The LLM pill in the status panel shows
⇄ llama3 in Ollama-blue while test mode is on, with the
status bar reading Test mode: two bots conversing — LLM (ollama:llama3:latest).
The config dialog (Tools → Configuration… / Cmd-,) is tabbed:
- General — Camera, Microphone, Claude voice (TTS), Avatar window, Test mode, Mirror mode (avatar mimics user's expression + head), plus the TTS engine + voice pickers.
- LLM — Live engine swap (Auto / Anthropic / Ollama / Demo), Anthropic model combo, Ollama model combo (with Refresh), API-key status; below the separator, a separate Test-mode bots section with its own engine + model + Partner-persona combos. Changing the partner persona restarts test mode automatically so the new camera-side avatar takes effect.
- Avatar — Persona combo with shortcut to the full picker (41 bundled appearance presets), head-nod cascade mode, body-rig weighting mode.
Open with Tools → Edit personas… (Cmd-Shift-I). Edit any character's identity inline; saving rebinds the running cognition store so the live avatar picks up the new traits on the next reply.
Each persona has its own JSON file under .faceview/memory/<persona>.json
holding three memory layers + a relationship score:
- Episodic —
{ts, type, text, significance, emotion, recalled}rows. Recall is scored by recency × significance × emotion × context × rehearsal. Consolidates down to 400 entries by retention when the list exceeds 500. - Semantic — facts/beliefs keyed by subject (
player,history,self) with confidence values. No decay. - Emotional — current emotions with exponential decay (~6h half-life).
- Relationship score — accumulates from each significant turn; brackets into character-defined levels (Acquaintance → Companion). Each level "unlocks" deeper conversational latitude in the prompt.
$ python tools/faceview_monitor.py memory
╭─ cognition · ict_xray_young (Claude) ──
│ path: /Users/george/claude_test/faceView/.faceview/memory/ict_xray_young.json
│ first_seen: 2026-05-13 session #4
│ user_name: George
│ relationship Lv 2 · Familiar (score 38)
│ mood joy (24%)
│ episodic 27 entries
│ semantic subjects: player, history
│
│ known about player:
│ name George
│ pref_1778685920 I love dark roast coffee
│
│ recent episodic:
│ sig=8 joy User: Hi! My name is George … — You: Nice to meet you, George!
│ sig=4 neutral User: What's new today? — You: Same Claude, fresher context …
╰──────────────────────────────────────────At inference, CognitionStore.narrate_for_prompt() builds a system-prompt
prefix from the character's identity, the relationship level, the current
mood, the semantic facts, and the top recalled memories. The same narrative
is injected via Conversation.effective_system() regardless of whether the
backing engine is Anthropic, Ollama, or demo — so the avatar stays itself
no matter what LLM is driving it.
Kokoro-onnx neural TTS runs locally on Apple Silicon CPU at real-time speed.
First-time setup downloads the model + voices (~340 MB) into
.faceview/tts/:
python -m faceview.speech.tts_kokoro --download
python -m faceview.speech.tts_kokoro --say "Hello — testing the voice." --voice af_sarah54 voices: af_* American female, am_* American male, bf_* British
female, bm_* British male. Each character has a per-persona default
(see assets/config/characters.json) — persona swap also swaps the
voice. You can override per-session in Tools → Configuration… → General
tab → Voice combo, or with FACEVIEW_TTS_VOICE=bf_lily.
If kokoro isn't installed or the model isn't on disk, TtsWorker
transparently falls back to pyttsx3 (macOS NSSpeechSynthesizer).
The avatar's own voice playing through speakers used to leak back into the mic and trigger another LLM call. Two layers prevent this:
- Audio mute at source —
AudioCapture.muted = TrueonTTS_STARTED, released 250 ms afterTTS_FINISHED. VAD never sees the avatar's voice, so the transcript panel + LLM bridge never see it. - STT-to-chat bridge gate — defence in depth; drops any
TRANSCRIPT_FINALthat lands within 2.5 s ofTTS_FINISHED(covers faster-whisper's async transcribe lag).
To talk over the avatar, hold the 🎤 Hold to talk button in the
chat panel. Press kills the current Kokoro utterance (terminates the
tracked afplay subprocess), un-mutes the mic, and overrides the gate
so your voice routes straight into chat. Release returns to normal
mute-during-TTS behaviour.
Two scripts let Claude Code (or a human) drive faceView from the shell.
# Read-only — status / chat / events / memory / watch / screenshot
python tools/faceview_monitor.py
python tools/faceview_monitor.py chat -n 20
python tools/faceview_monitor.py memory
python tools/faceview_monitor.py watch # live snapshot loop
# Write — launch / stop / chat / say / persona / emotion / engine / test / lifecycle / memory
python tools/faceview_drive.py launch # pulls Anthropic key from Keychain
python tools/faceview_drive.py persona ict_xray
python tools/faceview_drive.py engine ollama --model llama3:latest
python tools/faceview_drive.py test ollama --model llama3:latest
python tools/faceview_drive.py lifecycle test_mode --on
python tools/faceview_drive.py chat "What did we talk about yesterday?"
python tools/faceview_drive.py say "Hello!"
python tools/faceview_drive.py memory show
python tools/faceview_drive.py memory clear
python tools/faceview_drive.py stopBehind the scenes both talk to a FastAPI server on 127.0.0.1:8765
that the GUI starts at boot — full endpoint list lives in
src/faceview/server/api.py. An MCP server adapter exposes the same
operations as native Claude Code tools.
Conda-based, Python 3.11, Apple Silicon (M1/M2/M3/M4). Other macOS should work; Linux/Windows untested.
conda create -n faceview python=3.11
conda activate faceview
pip install -e ".[dev,speech,vision]" # add identity,emotion,mcp as needed
pip install kokoro-onnx soundfile # natural voice (optional but recommended)
# One-time voice asset download (~340 MB):
python -m faceview.speech.tts_kokoro --download
# Optional: USC ICT-FaceKit photo-real head (~23 MB after compile):
git clone --depth 1 https://github.com/USC-ICT/ICT-FaceKit /tmp/ICT-FaceKit
python -m tools.build_ict_blendshapes /tmp/ICT-FaceKitfaceview # GUI
python -m faceview # equivalent
ANTHROPIC_API_KEY=sk-ant-... faceview # with real Claude
FACEVIEW_HEADLESS=1 faceview # offscreen smoke
FACEVIEW_TEST_MODE=1 \
FACEVIEW_TEST_ENGINE=ollama \
FACEVIEW_TEST_MODEL=llama3:latest \
faceview # boot straight into two-bot test mode
pytest # 158 testsIf you store your API key in macOS Keychain, the recommended pattern:
security add-generic-password -a "$USER" -s "ANTHROPIC_API_KEY" -w # one-time
alias faceview-run='ANTHROPIC_API_KEY="$(security find-generic-password -a "$USER" -s ANTHROPIC_API_KEY -w)" /opt/anaconda3/envs/faceview/bin/faceview'See docs/TROUBLESHOOTING.md for the
most common issues and fixes — camera/mic permissions, broken VLMs,
demo-mode fallback, port conflicts, persona swap freezes, etc.
The HTTP server speaks OpenAI's /v1/chat/completions and /v1/models
wire format, so any OpenAI-SDK-compatible tool (Cursor, langchain,
the raw openai Python client, …) can point at faceView and get back
replies from whatever engine you've selected — plus faceView's live
perception block (what the camera sees, who's recognised, etc.)
prepended to the system prompt.
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:8765/v1",
api_key="not-used", # local-only, no auth
)
reply = client.chat.completions.create(
model="faceview",
messages=[{"role": "user", "content": "What can you see?"}],
)
print(reply.choices[0].message.content)Streaming SSE is not yet supported (planned). Set stream=False.
See INTERFACE.md for the full module map. Short
version:
mic ─► AudioCapture ──► VAD ──► STT ──► (echo gate) ──► CHAT_USER_MESSAGE
(muted during TTS) │
▼
chat input ─► ChatPanel ─────────────────► CHAT_USER_MESSAGE ── ClaudeClient
(memory narration
prepended to system)
│
┌───────────────────────────────────────────────┤
▼ ▼ ▼
ChatPanel TtsWorker SimCameraWorker
(history + CHAT_LOG) (kokoro/pyttsx3 (avatar.say →
→ afplay) lip-sync + mood)
cam ─► Camera ─► Presence/Identity/Emotion/Mouth/HeadPose ─► StatusPanel
│
▼
(mirror mode) SimCameraWorker
HTTP / MCP ─► Service ─► _GuiBridge slots ─► MainWindow handlers
- PySide6 GUI with one
QThreadper heavy stage (audio, video, ML inference, LLM, server) and an in-process pub/sub bus on Qt signals — thread-safe by construction viaQt.QueuedConnection. - Vision pipeline: webcam → MediaPipe presence + 478-point landmarks → InsightFace ArcFace owner-vs-stranger → DeepFace emotion → mouth-activity / viseme / head-pose detection. All ML deps are lazy-imported, so the GUI shell, tests, and CI screenshot capture run with the minimum install.
- Speech pipeline:
sounddevicemic → silero-vad → faster-whisper STT → LLM → Kokoro / pyttsx3 TTS. Same lazy-import policy. - Detachable layout — every panel is a
QDockWidget.LayoutManagersnapshots a default state at build time and persists user choices viaQSettings. - Live + headless screenshot via
widget.grab().save(), working underQT_QPA_PLATFORM=offscreenso CI can produce real PNGs.
$ pytest -q
158 passed in 65.7s
GitHub Actions runs the full suite + the headless smoke screenshot on every push, archiving the PNG as a build artefact. ML libs are lazy-loaded so CI doesn't need them installed.
This is a personal-use project. It's stable enough to use daily as a conversation interface and to drive from Claude Code. PRs and issues welcome; expect rough edges around platform-specific bits (macOS arm64 is the tested path).
- USC ICT-FaceKit — the photo-real avatar mesh + blendshape model.
- Kokoro-onnx — local neural TTS.
- faster-whisper, silero-vad, InsightFace, MediaPipe, DeepFace.
- Cognition architecture adapted from the
autonomous_worldNPC memory system and thetable_gamesLiving-AI design.




