Runnable recipes for Orca Agent Engine. Each one is a small, complete program that solves a real task and can be read in a sitting.
Recipes are TypeScript, built on the Orca TypeScript SDK, @runorca/orca-sdk. Every one reads its
configuration from the environment, creates what it needs, and cleans up after itself, so you can
run any of them repeatedly without leaving residue behind.
Documentation for the engine and its API lives at docs.runorca.ai.
- Node.js 22 or later, with pnpm.
corepack enableprovides the pnpm version pinned inpackage.json. Recent Node.js releases no longer include corepack; if the command isn't found, runnpm install -g corepackfirst. - A reachable Orca Agent Engine deployment. To run one yourself, follow the engine's quick start. The recipes run real agents, so the engine needs a model provider configured, as its quick start describes. A hosted workspace works too.
- A workspace API key for that deployment. For a self-hosted engine, the engine docs explain how to create one; a hosted workspace issues keys from its settings.
git clone https://github.com/orca-ae/orca-cookbooks.git
cd orca-cookbooks
cp .env.example .env # fill in ORCA_BASE_URL, ORCA_API_KEY and ORCA_MODEL
pnpm installORCA_MODEL is deliberately not defaulted, because the ids available to you depend on the model
provider your deployment uses:
-
Self-hosted engine: use a model id that the provider configured in your engine accepts.
-
Hosted workspace: list the ids with the
orkCLI.brew install orca-ae/tap/ork ork agent providers list
pnpm recipe --list
pnpm recipe fix-failing-testsStart with fix-failing-tests. It is the smallest recipe that still uses the whole loop, and the
others assume it makes sense.
| Recipe | What it shows |
|---|---|
| Fix failing tests | An agent copies mounted fixtures into a writable output workspace, then iterates on four entangled faults until its suite is green, verified from the final bash tool result against the original read-only tests. |
| Recipe | What it shows |
|---|---|
| Gate an agent behind a human decision | Custom tools as a round trip. The agent recommends; the host process holds the ledger and refuses an over-limit approval at the boundary. |
| Take an issue to a merged pull request | Three turns on one session - fix, CI failure, review comment - against a mock CLI rigged to fail the first check run. |
| Recipe | What it shows |
|---|---|
| Explore an unfamiliar codebase | Documentation that is confidently wrong, and an agent that has to ground itself in the source. Adds a resource to a session already running. |
| Turn a CSV into a report you can open | A sandbox built with pandas and plotly, a year of sales data, and one collapse buried in it. The generated HTML is downloaded and saved. |
| Recipe | What it shows |
|---|---|
| Remember preferences between sessions | Two sessions, one memory store. The second never heard the first conversation and has to recover what matters from what was written down. |
| Recipe | What it shows |
|---|---|
| Coordinate a team whose members have different tools | Three specialists with three tool surfaces, including one with no tools at all - so the coordinator cannot shortcut its own team. |
| Plan big, execute small | A coordinator with no read tool reviews six documents by handing each to a worker, spending its own context on the comparison instead. |
| Watch subagents work, live | A coordinator delegates to three specialists, and a follower attaches to each thread as it appears so the work is visible while it happens. |
| Recipe | What it shows |
|---|---|
| Triage an incident with a skill, and gate the fix | Runbooks as an uploaded skill, telemetry whose alert name is misleading, and a remediation tool that refuses the one thing the runbook forbids. |
| Recipe | What it shows |
|---|---|
| Wire an agent to an MCP server | The three pieces that must line up - server, toolset reference, and a vault credential bound by URL - plus answering always_ask confirmations. |
| Recipe | What it shows |
|---|---|
| Version a prompt, catch a regression, roll back | A v2 that reads like an improvement and scores worse. Measures both against labelled data, then pins new sessions to the version that won. |
| Recipe | What it shows |
|---|---|
| A chat bot whose backend is the session | Thread id in session metadata as the only record of the mapping, recovered by search after a restart. Two interleaved threads, no crosstalk. |
| Give an agent a database without giving it the database | Custom tools bridge to a store the sandbox cannot reach, and the query tool validates a model-composed filter before executing it. |
All fixture data is fictional.
- data-analyst builds its sandbox with
pippackages. It needs a deployment that allows environment package installation, and not every deployment does. - watch-subagents-live attaches to each subagent thread when its
session.thread_createdevent arrives. Some deployments don't project that event onto the session's primary stream; there the recipe finds threads by polling instead, so its followers attach a little later. - operate-in-production needs an external MCP server that speaks Streamable HTTP, so it isn't
self-contained. It explains what to set and exits when
ORCA_MCP_SERVER_URLis unset.
Some capabilities don't exist in the engine, so no recipe pretends they do:
| Capability | Status |
|---|---|
| Session spend budgets | POST /v1/sessions rejects a budget field |
| Advisor roster entries | The roster accepts agents and self only |
| Pinning inference to a geography | model.inference_geo is accepted, then dropped |
| Outcome grading that re-drives an agent | Outcomes are advisory |
| Skills discovered from a mounted repository | Mounted checkouts are not scanned |
| Scheduled deployments | Not implemented; triggers, in hosted workspaces, are the closest thing |
| Path | What it is |
|---|---|
recipes/<name>/ |
One recipe: README.md, main.ts, and its fixtures |
lib/ |
Shared helpers: client, bootstrap, streaming, custom tools, cleanup |
scripts/ |
The recipe runner and the checks |
registry.yaml |
Recipe metadata; the table above is generated from it |
authors.yaml |
Recipe authors, keyed by GitHub handle |
pnpm check # naming, license headers, registry, unit tests, typecheck
pnpm index # regenerate the recipe tablepnpm check runs offline and needs no credentials. Running a recipe against a live deployment
consumes real session time, so that is never automatic.
Three behaviours are worth knowing about, because they are easy to get wrong and are handled once here rather than in each recipe.
Resume cursors are transport cursors. The SDK exposes JSON event UUIDs but discards the SSE
id: sequence required by from_cursor. The helper streams a first turn live, then polls persisted
history after the last UUID for later turns so nothing is replayed or skipped.
Idle is not immediate. The session.status_idle event can arrive before the server has
committed the status change, so deleting straight after the event fails. waitForIdle polls the
session record to close that window.
A session has to stop before it can go. The server won't delete a running session, and a
recipe that fails or is interrupted mid-turn leaves one running. trackSession registers a release
that interrupts every thread still working, waits for idle, then deletes, ahead of the files and
agents the session uses. withCleanup runs it on Ctrl-C too, and the streaming helpers stop the
recipe at its next turn boundary so it doesn't start new work while cleanup runs.
New recipes and fixes are welcome. CONTRIBUTING.md explains how to add a recipe
and what pnpm check expects. If you use an AI assistant, read the AI policy first.
Everyone taking part follows the code of conduct.
- Issues for broken recipes and recipe requests.
- Discussions for questions and ideas.
- Security problems go through SECURITY.md, never a public issue.
Orca Cookbooks is licensed under the Apache License 2.0. See NOTICE.