Governed AI pipelines where the LLM can't fire actions — one Go binary, self-hosted, human-in-the-loop.
AI suggests. Deterministic code decides. The operator signs off.
Draftcat runs YAML-defined pipelines that triage email, qualify leads, draft replies, extract data from PDFs, and govern self-hosted voice AI. Every outbound action passes an operator approval gate, every LLM call is budget-checked, and every fetched item is deduped against a SQLite state store. One business per instance, self-hosted, auditable.
A normal approval server sees how every person voted. Draftcat's experimental FHE encrypted tally lets three or more reviewers turn approve or reject into unreadable ciphertext on their own machines. A collector combines those files without opening them; only the key owner can reveal the final count and learn whether quorum was met.
reviewers encrypt votes → collector adds unreadable ballots → key owner opens one total
collector never sees yes or no
# Once per vote: create the private key and the public key reviewers receive.
./draftcat fhe-vote keygen
# Each reviewer encrypts locally. The readable vote is never sent.
./draftcat fhe-vote encrypt --public fhe-public.json --context invoice-4821 \
--ballot <unique-random-invite> --vote approve --out reviewer.vote.json
# The collector combines 3+ encrypted ballots; the owner alone opens the result.
./draftcat fhe-vote tally --public fhe-public.json --context invoice-4821 \
--out tally.json alice.vote.json bob.vote.json carol.vote.json
./draftcat fhe-vote decrypt --secret fhe-secret.json --context invoice-4821 \
--expected 3 --quorum 2 tally.jsonUse it when separate teams, companies, or committee members need a shared approval but the tally host must not know individual votes. Keep the collector separate from the key owner and give the key owner only the final tally. Skip it when the same trusted Draftcat owner may see the votes, fewer than three people vote, or you need a public audit receipt—the zero-knowledge feature below is for that. Ciphertext files are still sent; the plaintext votes are not. Read the encrypted vote walkthrough and threat model before evaluating it.
Sometimes a customer, auditor, or partner needs evidence that a human approved an AI action — but should not receive the message, the reviewer's identity, or your internal workflow. Draftcat can turn a signed approval row into a zero-knowledge proof:
| The verifier learns | What stays private |
|---|---|
| A direct human approval was recorded | Customer message and payload hash |
| The required reviewer quorum was met | Reviewer identity and exact vote counts |
| The proof came from the Draftcat instance key they pinned | Pipeline, step, time, nonce, and instance secret |
# Operator: publish this commitment once through a trusted channel.
./draftcat zk-receipt key-id
# Operator: create a shareable proof for the latest human approval.
./draftcat zk-receipt prove --out approval.proof.json invoice-due-diligence
# Customer or auditor: verify it without DRAFTCAT_APPROVAL_SECRET or database access.
./draftcat zk-receipt verify --expect-key <pinned-key-commitment> approval.proof.jsonThis is an experimental cryptographic preview, not a production compliance claim. It uses an embedded BN254/Groth16 circuit and a development single-party setup; the circuit has not received an independent audit. Use it to evaluate the disclosure model, then replace the setup through a ceremony before relying on it in production. See zero-knowledge approval proofs for the trust model, exact statement, and limitations.
New in v0.9.0: daily model usage persists in SQLite, parallel pipelines keep separate budget totals, and every engine model call passes one admission gate. Approval records commit before work is released. Unconsumed tool actions can be revoked, consumed actions accept durable caller-reported outcomes, and live or recovered permits share expiration and policy checks. Exact scalar output validation and bounded provider reads complete the release. See the budget and lifecycle guide.
New in v0.8.0: webhook retries can carry a durable
Idempotency-Key, signed replay identities are claimed atomically, and tool permits recheck their policy before execution. Requests reject oversized or ambiguous data and retain exact numbers. Older state stores upgrade safely; completed and failed runs carry exact approval identities. Audit commands read without modifying the database, anddraftcat receipts verifychecks exported JSONL offline. See the upgrade and reliability guide.New in v0.7.0: execution decisions now carry their proof. Every tool-gate route is authenticated, each request has a stable action identity and exact policy binding, and an allowed decision becomes an atomic consume-once permit before the side effect runs. Webhook acceptance is durable before HTTP 202 and can be polled after handoff. Versioned receipts bind action, payload, policy, and expiry, with
draftcat receipts list|show|exportfor verification-ready JSONL. Orderedmodel_policyrules can deny or send matching model input/output to a human, while/healthzand/readyzgive orchestrators a safe listener contract.In v0.6.0: the gate holds under load. The tool-call gate answers asynchronously (
mode: async,wait:) so a harness with a short HTTP timeout never loses a decision, and a tool call waiting on a human is durable across a restart. Rules constrain arguments (args:- glob, regex,one_of,min/max) and never widen on a mismatch. A repeat guard stops an agent that loops on one call from paging you, the operator hears about denials the gate made on its own,/pendinganddraftcat pendinglist every open gate,/statusshows spend against caps, cost caps enforce the provider's real charge, rate limits back off instead of failing the run - and one Telegram update pump fixes taps that were silently lost while two gates were open at once.
| Draftcat | n8n | LangChain agents | Agent harnesses (Flue, Claude Code) | |
|---|---|---|---|---|
| AI execution model | Deterministic boundary; AI cannot fire actions | Bolt-on LLM nodes in visual workflows | Agent decides next action freely | Agent acts autonomously in a sandbox |
| Human-in-the-loop | Required on every outbound step | Optional manual nodes | Optional; not the default | Optional (dispatch a message mid-run) |
| Token budgets | Per-step / pipeline / day, enforced | None | None | App-managed, not built in |
| Prompt-injection defense | Input sanitization + output schema validation | None | None | Sandbox isolation; app-managed |
| State & dedup | SQLite-backed; items processed at most once | DB-backed | In-memory | Session store / Durable Objects |
| Runtime | Single Go binary | Node.js + Postgres | Python + dependency tree | TypeScript, runtime-agnostic |
Use n8n for drag-drop integrations across 400+ services. Use LangChain for research and open-ended exploration. Use an agent harness like Flue when you want an agent to roam a sandbox and choose its own steps. Use Draftcat when a wrong LLM choice means a real customer gets emailed.
However your agent runs, draftcat sits between it and your customer systems as a mandatory approval gate — not a tool the model can route around. The same gate holds in both setups:
You only talk to your agent. You don't control its runtime, so it hands work to draftcat over a webhook — but it can only start a gated pipeline, never fire a customer-facing action itself.
You control the harness. Your runtime (n8n, your own agent loop, Dograh) does the roaming and integrations, then routes every outbound action through draftcat — the one gate it can't bypass — and gets the result plus an audit trail back.
- Token budgets — per-step / pipeline / day; any breach halts the run immediately.
- Cost budgets —
per_day_cost/per_pipeline_costcap spend in money. On OpenRouter the caps are enforced on the charge the provider reports for each call (cached and reasoning tokens included); elsewhere on your configured per-1k rates. The approval prompt shows what the run has spent, and/statusshows the day against every cap. - Human-in-the-loop — every outbound action requires an explicit operator decision, made live or declared in advance.
- Any operator channel — the
hitl/v0protocol keeps draftcat as the gate and lets an untrusted relay own presentation. Teams runs through a Power Automate flow in your own tenant: no bot, no Azure app registration, no admin consent. Check yours withdraftcat hitl verify <relay-url>. - Consume-once tool permits -
POST /gate/tool-callputs an agent's MCP or SDK call through the same gate as a pipeline step. Bearer authentication covers ask, poll, and consume. A stableaction_idmakes retries idempotent; the binding covers the exact arguments, policy, and expiry; only the first successfulPOST /gate/tool-call/<id>/consumecarriespermit: execute. Seedocs/tool-gate.md. - Repeat guard — inside
repeat_windowan identical tool call (same agent, tool, arguments) gets the gate's remembered answer instead of a new prompt: a denied call stays denied, an in-flight call joins the open prompt, andmax_repeatsstops a looping agent from paging you. - Denial notices — a refusal the gate makes on its own (unlisted tool, argument outside a rule, repeat guard) is reported to the operator channel, one notice per agent, tool and reason per window, so nothing is refused silently.
- Open gates —
/pendingon the channel anddraftcat pendingon the host list every approval waiting on a human, pipeline steps and tool calls alike, with how long each has waited and how long it has left. - Risk tiers — steps declare
risk: low | normal | high, andapproval_policycan pre-approve a declared class. High risk never qualifies, and each exemption is audited aspolicy_approvewith the rule that fired. - Escalation —
escalate_afterre-notifies before a gate times out;escalate_towidens who is told, never who may decide. - Durable, run-correlated gates — every gate is written to SQLite before the draft goes out, so an approval in flight survives a restart, and each decision records the run it released.
- Approver scoping —
approvers:on a step narrows who may decide it to a subset ofallowed_users. Quorum says how many; this says which ones. It can only narrow, never widen. - Model I/O policy - ordered
model_policyregex rules check exact input before it reaches the provider and output before it leaves Draftcat. A match can deny or enter the existing human approval gate, and the decision is written as a versioned receipt. - Input sanitization — operator input is scrubbed for prompt-injection patterns before the LLM.
- Output validation — AI output is checked against the skill's
output_schema(field types, numericmin/max,enummembership) and rejected if it doesn't conform. - Checked action receipts - v2 receipts bind immutable action ID, payload hash, policy digest, validity window, run, and decision. List, inspect, or stream JSONL from SQLite with
draftcat receipts; seedocs/action-receipts.md. - Durable webhook admission - Draftcat writes a body-hash-only admission row before returning HTTP 202. The response includes
admission_idand an authenticated poll URL; unfinished admissions becomeinterruptedafter restart. - Health contract -
GET /healthzreports process liveness andGET /readyzsucceeds only while the SQLite decision store is available. - Private approval proofs — share proof that a direct human approval met quorum without sharing the action, approver, or counts; see
docs/zk-approval-proofs.md. - Encrypted approval tally — combine three or more encrypted votes without letting the collector read any individual vote; see
docs/fhe-vote-tally.md. - Rate limiting — per-user, per-minute caps on operator interactions.
- Channel security — allowed-user lists + input-length limits enforced at startup; the engine refuses to start without them.
- Config validated on boot — the engine runs the same checks as
draftcat validateat startup and refuses to start on errors, so problems surface at boot rather than mid-run.DRAFTCAT_SKIP_VALIDATE=1overrides. - Observability — opt-in structured JSON spans, one per pipeline and step (duration, status, tokens, cost). Off by default;
observability.spans: trueorDRAFTCAT_TRACE=1.
Install the native binary through npm (Node.js 18 or newer):
npm install -g draftcat
draftcat --helpThe installer downloads the matching Linux, macOS, or Windows binary and verifies it against the checksums attached to the GitHub release. No Go toolchain is required.
With a Go toolchain, install the CLI from the module:
go install github.com/renezander030/draftcat@latestOr build from source:
git clone https://github.com/renezander030/draftcat.git && cd draftcat
cp secrets.yaml.example secrets.yaml # operator IDs + API keys
go build -o draftcat . && ./draftcatOr with Docker (no Go toolchain needed):
git clone https://github.com/renezander030/draftcat.git && cd draftcat
cp secrets.yaml.example secrets.yaml
docker compose upPipelines live in config.yaml, prompts in skills/. Run draftcat doctor to check credentials, operators, the state store and schedules before the first start. A SQLite store opens at ./state.db on first boot. To add the EU-resident voice AI plugin: go build -tags voice -o draftcat . — the lean binary is unchanged when the tag is off.
draftcat is a service you self-host, not a plugin an agent loads. It runs as a long-lived process and pings you on Telegram to approve each action. Pick the path that fits.
Click, sign in, and paste three values — your Telegram bot token, an OpenRouter key, and your Telegram user ID. Render runs it always-on with a persistent disk: no VPS, no shell, no TLS to configure. (It deploys as a background worker, so it has no public URL — ideal for the "watch my inbox, approve on Telegram" job. For inbound agent webhooks, use a host you control, below.)
curl -fsSL https://raw.githubusercontent.com/renezander030/draftcat/master/install.sh | shPulls the image, scaffolds ~/draftcat/.env + a compose file with a state volume, and prints the two steps left (fill the .env, then docker compose up -d). Config and skills are baked into the image.
docker run -d --restart unless-stopped \
-v draftcat-state:/data -e DRAFTCAT_STATE_PATH=/data/state.db \
-e DRAFTCAT_TG_TOKEN -e OPENROUTER_API_KEY \
-e DRAFTCAT_TG_ALLOWED_USERS=<your-telegram-id> \
ghcr.io/renezander030/draftcatOr docker compose up from a clone — builds the same image and mounts your local config.yaml/skills/ so you can edit pipelines.
The setups above run the always-on operator loop. To let an external agent or harness trigger pipelines, enable the webhook in config.yaml (schedule: webhook on the pipeline), publish the port, and put Caddy/nginx in front for TLS:
webhook:
enabled: true
addr: 0.0.0.0:8088
secret_env: DRAFTCAT_WEBHOOK_SECRETcurl -X POST https://draftcat.yourco.eu/hooks/<pipeline> \
-H "Authorization: Bearer $DRAFTCAT_WEBHOOK_SECRET" \
-d '{ "lead": "..." }'The POST only starts a gated pipeline — the approval step still runs, so inbound can never make the LLM fire a customer-facing action.
Draftcat writes the admission to SQLite before returning 202:
{"admission_id":"wh_...","status":"accepted","poll":"/hooks/status/wh_..."}Poll that path with the same bearer token. GET /healthz is a liveness check;
GET /readyz verifies that the decision store is reachable.
Each pipeline is a fixed sequence of typed steps. The LLM never chooses the next action — it produces structured output, the engine validates it against a schema, and an operator approves before anything reaches a customer.
| Step type | What it does |
|---|---|
deterministic |
Plain Go — fetch emails, parse PDFs, dedup, route, notify |
ai |
LLM inference with a skill template, budget-checked, schema-validated |
approval |
Operator reviews via Telegram: approve / edit / reject |
pipelines:
- name: invoice-due-diligence
schedule: 1h
steps:
- {name: parse-pdf, type: deterministic, action: pdf_extract, vars: {path: /inbox/invoice.pdf}}
- {name: extract, type: ai, skill: extract-line-items}
- {name: verify, type: deterministic, action: pdf_verify_cite, vars: {fail_on_unresolved: "true"}}
- {name: review, type: approval, mode: hitl, channel: telegram}| Action | What it does |
|---|---|
gmail_unread |
Fetch unread Gmail messages (deduped per pipeline) |
whatsapp_intake |
Normalize inbound WhatsApp JSON into governed pipeline input |
ghl_new_contacts |
Fetch recent GoHighLevel contacts (deduped) |
ghl_stale_opportunities |
Fetch stalled GHL opportunities |
ghl_unread_conversations |
Fetch unread GHL conversations |
pdf_extract |
Parse a PDF into text + per-fragment bounding boxes (pure-Go) |
pdf_verify_cite |
Resolve <cite> tags in AI output against the parsed PDF |
notify |
Send AI output to the operator channel |
voice_* / dograh_* |
Voice plugin actions (-tags voice) |
Add an action by appending a case to the deterministic switch in main.go and registering its name in internal/validate/. See internal/ghl/ and internal/dograh/ for connector patterns.
For WhatsApp, run a small whatsmeow receiver as the session owner and POST its
normalized message JSON into a schedule: webhook pipeline that starts with
whatsapp_intake; see docs/whatsapp.md.
provider:
type: openrouter
api_key_env: OPENROUTER_API_KEY
models:
haiku: {model: anthropic/claude-haiku-5.5, max_tokens: 4096}
budgets:
per_step_tokens: 2048
per_pipeline_tokens: 10000
per_day_tokens: 100000
per_day_cost: 5.00 # money cap, same unit as your model rates (0 = off)
per_pipeline_cost: 0.50
alert_at: [0.5, 0.8] # tell the operator at 50% and 80% of a daily cap
observability: {spans: false} # or DRAFTCAT_TRACE=1
state: {path: ./state.db}An approval step can narrow who may decide it:
- name: release-payment
type: approval
channel: telegram
quorum: 2 # how many must approve
approvers: [111111, 222222] # which ones (subset of allowed_users)Cost caps are checked between calls: a call is refused once spend has reached the cap. Pair them with per_step_tokens to bound the size of any single call. Explicit rate-limit rejections (HTTP 429 without declared usage) retry with backoff and honour Retry-After; provider.max_retries sets the attempt limit. Transport errors, timeouts, server errors, and responses with uncertain billing halt the call and require usage reconciliation before another dispatch. See the budget recovery guide.
The tool-call gate is configured the same way, per tool:
tool_gate:
enabled: true # served on the webhook listener
repeat_window: 10m # identical call → same answer, no second prompt
tools:
- name: read_calendar # listed = allowed, audited
- name: send_email
risk: high
require_approval: true
args:
to: {glob: "*@example.com"} # inside the rule: ask as usual
on_mismatch: deny # outside it: refuse without askingEvery gate request uses the webhook bearer token. Send a stable action_id,
then consume an allowed binding exactly once before running the side effect:
curl -X POST http://127.0.0.1:8088/gate/tool-call \
-H "Authorization: Bearer $DRAFTCAT_WEBHOOK_SECRET" \
-H 'Content-Type: application/json' \
-d '{"action_id":"send-invoice-4821","tool":"send_email","args":{"to":"[email protected]"}}'
curl -X POST http://127.0.0.1:8088/gate/tool-call/send-invoice-4821/consume \
-H "Authorization: Bearer $DRAFTCAT_WEBHOOK_SECRET" \
-H 'Content-Type: application/json' \
-d '{"binding_hash":"sha256:..."}'Model input and output policy is ordered and deterministic. deny fails the
LLM call closed; review pauses at the configured operator channel:
model_policy:
max_preview_chars: 800
rules:
- id: credentials-in-input
phase: input
pattern: '(?i)(api[_ -]?key|password)'
action: review
reason: Credentials require an explicit operator decision.
- id: unsupported-claim
phase: output
roles: [drafter]
pattern: '(?i)guaranteed results'
action: deny
reason: Do not send unsupported guarantees.Skills are YAML prompt templates in skills/ with an output_schema the engine enforces. With -tags voice, a voice: block configures the webhook receivers, Dograh endpoints, and pre-call lookup — see docs/voice.md.
draftcat # run the engine (validates config first; refuses to start on errors)
draftcat validate [--strict] # lint config + skills
draftcat doctor [--json] # read-only preflight: credentials, operators, state, ports, schedules
kill -HUP <pid> # reload pipelines, budgets, policies and skills without a restart (or /reload)
draftcat test <pipeline> # dry-run against fixtures/<pipeline>/ (never touches real APIs)
draftcat runs [pipeline] # recent runs + the approval decisions in each (--json to archive)
draftcat pending # approval gates waiting on a human right now (--json)
draftcat receipts list # approval receipts and verification status (--json)
draftcat receipts show <id> # one versioned receipt
draftcat receipts export # JSONL to stdout (--out path writes mode 0600)
draftcat audit-verify # verify signed approval receipts
draftcat hitl verify <url> # run the hitl/v0 conformance suite against a relaydraftcat runs reads the governance record back out of SQLite: what ran, when, and who decided what.
2026-07-26T05:37:31Z invoices ok 60.0s
release-payment adjust by 111 (0/2)
release-payment approve by 222 (2/2) [signed]
Per-step timings and token counts live in the observability spans (observability.spans, OTLP/Prometheus). This is the durable record of decisions.
Pre-commit hooks (lefthook) run gofmt, go vet, go build, go test -short, and golangci-lint on new code; pre-push runs draftcat validate.
State persists to SQLite (./state.db by default): fetched item IDs are deduped per (pipeline, scope) so items process at most once, every run is recorded (started_at / ended_at / status), and writes use WAL mode for crash safety without per-write fsync.
Approval rows can be made tamper-evident with signed receipts, so a later audit
can verify which operator approved which payload hash. See
docs/action-receipts.md.
A pipeline's schedule decides when it runs — an interval (1h), a cron expression in the pipeline's time zone ("0 8 * * 1-5" with timezone: Europe/Berlin), @daily-style shortcuts, manual (operator /run only), or webhook. Schedules continue from the recorded run history across restarts, pause_after_failures stops a pipeline that keeps failing, and SIGTERM lets running pipelines finish within timeouts.shutdown_grace. See running Draftcat unattended.
The webhook server is opt-in and opens no port unless enabled:
webhook: {enabled: true, addr: 127.0.0.1:8088, secret_env: DRAFTCAT_WEBHOOK_SECRET}curl -X POST http://127.0.0.1:8088/hooks/invoice-due-diligence \
-H "Authorization: Bearer $DRAFTCAT_WEBHOOK_SECRET" -d '{"path": "/inbox/invoice.pdf"}'The body reaches the pipeline as {{webhook_body}} / {{input}}; bearer auth is constant-time, and a second trigger while the pipeline is running gets 409. Before 202, Draftcat stores an admission ID, pipeline, body hash, and status in SQLite. GET /hooks/status/<admission_id> returns the authenticated status without retaining the request body. A webhook only starts a pipeline - the approval gate still runs, so an inbound request can never make the LLM fire an outbound action.
Signed requests. Bind each trigger to its exact body and a timestamp with an HMAC receipt, on top of the bearer token:
webhook:
enabled: true
secret_env: DRAFTCAT_WEBHOOK_SECRET
require_signature: true
max_skew_seconds: 300 # defaultX-Draftcat-Signature: t=<unix>,v1=<hex hmac-sha256(t + "." + body)>
Requests outside the skew window are refused, and each signature is spent once (recorded in the dedup table), so a captured request cannot be re-fired. A signature header is always verified when present, even with require_signature: false.
Hook up your Dograh to your draftcat instance!
Built with -tags voice, Draftcat becomes the EU-resident writeback + governance layer for self-hosted voice agents (Dograh, Pipecat, or any orchestrator that posts JSON webhooks): 5 lifecycle webhook receivers, sub-300ms pre-call context lookup, a 7-step Learning-Item review pipeline before any prompt/KB change ships, Dograh REST admin actions, and per-day call/minute budgets with bearer-auth webhooks. Full wiring recipe and runnable DACH fixtures in docs/voice.md.
The deterministic-boundary architecture is documented in the Production AI Automation Notes gist series, each mapping to draftcat code:
- #1 Agent Approval Gates — proposed actions, schema validation, audit log
- #2 Token Budgets — per-step / pipeline / day enforcement
- #5 SQLite Dedup + Crash Safety — WAL mode,
seen_items, run audit - #6 Prompt-Injection Defense — input sanitization + output schema validation
- #7 PDF Cite Verification — auditable LLM extraction with per-fragment bounding boxes
- #11 Pipeline Fixture Testing — dry-run pipelines from JSON fixtures; zero API calls in CI
- #12 LLM Skills as YAML — prompt + output_schema + role in versioned YAML, validated by a linter
- #13 Inbound Agent Webhook Auth — constant-time bearer token, fail-closed on empty secret, async 202 dispatch
- #14 Self-Improving Voice Agent — harvest Learning-Items, group, propose a minimal workflow diff, two approval gates, git commit + auto-versioned Dograh publish
- #15 AI Action Audit Trail — versioned action-bound
action_approvalsreceipts: who approved which payload, when; action-level receipt binding, expiry and fail-closed audit persistence
- capcut-cli — edit CapCut / JianYing video drafts from the CLI. Same DNA: single binary, no API, structured JSON boundary between agent and tool.
MIT. See LICENSE.





