typed

Migrate to typed

For Pro / Max customers evaluating typed as a fallback for an existing AI coding CLI subscription (or a replacement, if it fits).


1. The config change

The fastest path is the typed launcher. One curl installs it with your API key wired in:

curl -fsSL https://app.typed.cloud/install.sh | bash -s <your-typed-key>

On Windows, run that inside WSL Ubuntu, or this from PowerShell (it needs Git for Windows and offers to install it with scoop):

$env:TYPED_INSTALL_KEY='<your-typed-key>'; irm https://app.typed.cloud/install.ps1 | iex

The dashboard pre-fills the command with your key + a Copy button. Once installed, typed and t are on your PATH:

typed                  # launches `claude` against typed.cloud
t                      # one-keystroke alias, identical behavior
typed --resume <id>    # any flags pass through

The launcher sets every required env var, picks the right defaults, and skips Claude Code's first-run "approve this key?" prompt.

If you would rather wire it up by hand, typed speaks the Anthropic API directly. Set three environment variables and any Anthropic-API-compatible client switches over:

export ANTHROPIC_BASE_URL=https://api.typed.cloud
export ANTHROPIC_AUTH_TOKEN=<your typed key>   # bearer/gateway pattern
export ANTHROPIC_MODEL=typed++

ANTHROPIC_AUTH_TOKEN (which Anthropic SDKs send as Authorization: Bearer) is the recommended variable: typed.cloud is a gateway, and Claude Code understands the bearer pattern as a proxy token, so it skips the interactive "approve this key?" prompt that fires when you use ANTHROPIC_API_KEY instead. typed.cloud accepts both header shapes, so existing configs continue to work, but new setups should prefer ANTHROPIC_AUTH_TOKEN.

The third line is required. typed uses its own model identifiers -- typed++ and typed (paid, hosted) and typed-local (free, run on your own machine); if your client sends an upstream model name like claude-3.5-sonnet, typed returns a 400 with the full list of valid identifiers in the error body. See section 1a below for how to choose.

Pick a hosted id for this path. The snippet above uses typed++, the premium hosted tier; the launcher fills in typed, the economy hosted tier, for Claude Code when nothing names one. The free typed-local id is local-only: it runs on a llama-server on your own machine and is reachable only through the first-party typed CLI pointed at that server (see the local setup guide). Sending it to api.typed.cloud from a hand-wired client returns a 503 config error, not a free response. Since 23 September 2026 a bare typed is the economy hosted tier, served and billed on this path; before that date it was the free local id, and a typed request to api.typed.cloud was refused with that same 503 rather than processed.

The canonical list is also available programmatically: the Anthropic SDK's client.models.list() works against typed (or curl https://api.typed.cloud/v1/models -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN"). SDK consumers that introspect the model surface see every canonical id in Anthropic-API shape -- the two paid hosted models, typed++ and typed, and the free local one, typed-local; back-compat aliases stay valid for /v1/messages but are intentionally omitted from discovery.

We run our own day-to-day on Claude Code against typed. Other Anthropic-API-compatible clients (Cursor, Cline, Roo Code, OpenClaw, and any tool that reads ANTHROPIC_AUTH_TOKEN or ANTHROPIC_API_KEY from the environment) should work the same way. The API is wire-compatible. If your client doesn't pick up the env vars, set the same values in its config file.

You can flip back at any time by swapping the variables. Nothing else changes.

1a. Choosing a model

typed exposes exactly three model ids: two hosted and paid, one local and free. Set ANTHROPIC_MODEL to one of them.

typed++ and typed -- paid, hosted. Both run on our infrastructure and draw on the same monthly budget on every plan. They are two different models, not two settings of one.

Model ID Use when Quality Latency Cost per request
typed++ The premium hosted tier, for serious work: refactors, feature work, code review, and the hard problems too. High (cache-reliable) Fast (typically 3-8s) Dialed to "high" reasoning effort. Cache-reliable, so most coding-client traffic gets cache discounts and runs noticeably cheaper per request.
typed The economy hosted tier, the cheapest paid option, and the typed CLI's default with a key on a paid plan: high-volume routine work where cost matters more than depth Mid Faster Cheapest paid tier, so the same budget stretches furthest here. Same 1M context window as typed++.

Until 23 September 2026 the economy tier was spelled typed-- (and typed++high before 20 September 2026); both spellings still resolve to typed. Before that date a bare typed named the free local tier below, and sending it to api.typed.cloud was refused with a 503. Since then it is served and billed as the economy tier.

typed-local -- free, runs on your own machine. The model runs on a llama-server you point typed at; typed is out of the request path entirely, which is why the tier costs nothing and needs no account. See the local setup guide.

Model ID Use when Notes
typed-local Everyday coding on your own hardware, offline work, anything you would rather not send anywhere The mixture-of-experts local model, and the tier the first-party CLI picks when nothing names one and no paid plan is on the account -- which is what a fresh install without a key selects. With a paid plan the CLI picks typed instead, and a launch that falls back to Claude Code as the client gets typed regardless -- see "Which default you actually get" below.

The three ids are not a ladder. typed is a lighter model than typed++, which is why it costs less, and the free typed-local tier is a different model again on different hardware, so do not read it as a step "below typed". Pick by where you want the work to run and what you want to spend, not by position.

Both hosted tiers carry the same 1M context window and both accept image input, so the difference between them is quality and price. We deliberately do not re-route a request onto a different model to make it fit, because that would quietly bill you for a model you did not choose.

1b. Running typed and Claude side by side

The launcher does not replace claude or modify your shell; it only sets env vars on the subprocess it spawns. Both commands stay on your PATH and run as independent processes:

  • claude: runs Claude Code against Anthropic directly with your Claude subscription auth. Unaffected by the typed install.
  • typed (or t): spawns Claude Code with ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN / ANTHROPIC_MODEL pointed at typed.cloud.

You can run as many of either at the same time across separate terminal panes / tmux windows / IDE terminals. Each invocation is its own process with its own environment:

Scenario Result
Pane A: claude, Pane B: typed Independent. Pane A hits Anthropic, Pane B hits typed.cloud.
N typed panes simultaneously All share one typed.cloud account / one monthly quota pool. No concurrency lock on the API side; you just burn quota proportionally.
M claude panes simultaneously All share one Anthropic subscription / one rate-limit budget. Same logic.
Mix of typed and claude panes typed.cloud's rate-limiter and Anthropic's are completely separate; the two backends don't see each other.

Toggle without rewriting aliases. Three subcommands let you flip the launcher's behavior without setting / unsetting env vars by hand:

typed status     # show current routing state + api key source
typed disable    # writes a flag file; subsequent `typed`/`t` exec `claude` WITHOUT typed env vars
                 # (your Claude subscription is used instead)
typed enable     # remove the flag file; route through typed.cloud again

Useful for bouncing between accounts mid-session without uninstalling or rewriting shell rc files.

One caveat (Claude Code's, not typed's). Claude Code writes session state to a per-project directory. Two claude-family processes (whether claude or typed/t) running in the same project directory at the same time can race on transcript / session files. Running them in different project directories, the natural pattern when you'd want both, is fine.

Context windows by model. On both hosted tiers, text and image-bearing (multimodal) requests share the same context window. The free local tier is the exception:

Model Context window
typed++ 1M tokens
typed 1M tokens
Multimodal (any image-bearing request on either hosted tier) 1M tokens
typed-local whatever your own llama-server was started with -- typed reads the real value at startup and sizes the session to it

1M covers the typical envelope of a Claude Code session with room to spare, and it applies uniformly across both hosted tiers: same window for text and for image-bearing requests, with no opt-in header, no model change, and no per-request configuration. (For reference, current-generation Claude models, Opus 4.6/4.7/4.8 and Sonnet 4.6, include a 1M window at standard pricing; current Haiku 4.5 and legacy Sonnet 4.5 still cap at 200K.)

The free local tier is bounded by your own hardware, not by us. Its window is whatever you gave llama-server, which is normally far smaller than the hosted figure above, and on that tier the requested max_tokens comes out of the window -- see the local setup guide.

Reading the window off the response. Every served /v1/messages response carries an x-typed-context-window header: the context window, in tokens, for the route that served the request. A client can read the figure from the server instead of copying it from this page.

Which default you actually get. There is no server-side default. The API requires model on every request: /v1/messages returns a 400 -- model and max_tokens are required -- when the field is missing or is not a string, and it never picks a tier on your behalf. "The default" is therefore a CLIENT-side concept. Since 23 September 2026 the rule is: With no typed.cloud key the CLI runs typed-local (free, on your machine). With a key on a paid plan it runs typed (hosted economy, billed to your plan). Name typed++ when you want the strongest tier. In full: an explicit choice wins (--model, --effort, ANTHROPIC_MODEL); otherwise the tier a Yaw Mode pane has set wins; otherwise a pin (.typed, ~/.config/typed/model); otherwise the first-party CLI runs typed when a key is on the machine and the account holds a paid plan, and typed-local in every other case (no key, or a key whose account has no plan); and a client that is not the first-party CLI gets typed.

What you launch Tier when nothing names one
typed / t after a successful install, or after typed cli on (first-party CLI -- what the installer selects) typed with a paid plan on the account; typed-local -- the free local model, no account required -- without one
typed / t spawning Claude Code (the fallback client, e.g. after typed cli off or a partial install), or any other client through the launcher typed -- a hosted tier regardless of plan, because the free typed-local tier is local-only and needs the first-party CLI; the API decides what your key is entitled to
The first-party CLI invoked directly same as the first row
Anything else (raw SDK, curl, another gateway) none -- send model yourself or take the 400

"A paid plan" means any plan, on any status: the CLI records the answer from GET /v1/me/plan on your machine (~/.config/typed/plan) when you run typed login, typed doctor or the first launch after logging in, and refreshes it in the background about once a day. typed status prints the default your next launch would use and why (which plan it recorded and when, or that no key is on the machine, or that a key is present but its plan has not been recorded yet -- typed doctor records it).

What outranks the built-in default, highest first: a client --model flag, an --effort word, an explicit ANTHROPIC_MODEL (the launcher only fills it in when you have not set it: ANTHROPIC_MODEL="${ANTHROPIC_MODEL:-<default>}"), the tier a Yaw Mode pane exports, a .typed repo pin, then a global pin from typed use <tier>. Run typed status to see the tier your next launch would use and where that value came from.

Reach for typed++ when you want the premium hosted tier for serious work -- cache-reliable, fast, enough reasoning headroom for feature work and hard problems alike. Stay on typed for routine work, where a quick answer with less depth is enough and the budget should stretch. Run typed-local when you would rather the work never left your machine.

Legacy aliases that work for backward compatibility. None of them is a separate tier; each one is another spelling of one of the three ids above:

Alias Resolves to
typed-- typed (the name the economy tier had from 20 to 23 September 2026)
typed++xhigh typed++ (the name this tier had until the 20 September 2026 restructure)
typed++high typed (the name the economy tier had until the 20 September 2026 restructure)
typed-xhigh typed-local (the spelling of the free local tier before 19 August 2026)
typed-max typed-local
typed-plus-plus typed++ (the spelled-out form)
typed-high typed
typed-medium typed++
typed-minus-minus typed (the spelled-out form of the economy tier's former typed-- spelling)
typed-pro typed++
typed-long-context typed++ (legacy alias from when typed had a distinct long-context tier; there is no separate long-context rung today, and a request is never moved to a larger one)

typed-local was itself an alias of the free tier (spelled typed from 19 August to 23 September 2026, and typed-xhigh before that) and is now that tier's canonical id, so a pin on it keeps meaning exactly what it meant. A pin on a bare typed written before 23 September 2026 meant the free local tier; the launcher and the CLI rewrite such a pin to typed-local once, and say so, rather than launching it on the economy tier unannounced.

typed++max is retired and is NOT an alias. It was a third hosted tier, removed in the September 2026 restructure. It was a distinct model at a distinct price, so resolving it onto typed++ would have silently changed both what served your request and what you were billed for. Sending it now returns a 400 that names typed++ as the tier to move to and lists the valid ids, and the launcher refuses a saved pin on it with the same guidance rather than guessing a replacement. Usage you ran on it before then still shows on your usage page, labelled as retired.

typed-low and typed-fast are gone and are NOT aliases. Both were removed on 2026-08-19, for the same reason: each named a combination no surviving model reproduces -- the typed++ rung with reasoning switched off entirely -- so mapping either onto a survivor would have silently changed what you were billed for and how deeply your prompt was reasoned about. Sending either now returns a 400 like any other unknown id, and a saved pin on either is treated as stale. If you were on one of them, pick typed++ for the premium hosted tier, typed for the cheapest paid rung, or typed-local to move off paid entirely. (typed-fast still appears on billing history from before the rename -- that is a historical label on old rows, not an id you can send.)

Anything else, including all claude-*, gpt-*, or typos, returns a 400 with the full list of valid IDs in the error body.

1c. The typed CLI: a first-party client, opt-in

Everything above runs your existing client against typed.cloud. typed also ships a first-party coding agent CLI, built in-house alongside the backend. It is opt-in:

typed cli on       # subsequent typed / t runs launch the typed CLI
typed cli off      # back to the previous default
typed cli          # print the current state

(typed tui on|off is an interchangeable alias for typed cli on|off and still works. typed beta on|off is a deprecated alias for the same thing and still works for now.)

Claude Code against typed.cloud remains fully supported either way. typed cli off reverts completely, and TYPED_CLIENT in the env overrides the choice for a one-shot run. Nothing else about your setup changes.

What carries over unchanged. The typed CLI reads the same configuration surfaces Claude Code does, so an existing setup works without rewriting anything:

  • CLAUDE.md discovery (user, project, .claude/, CLAUDE.local.md), including @import lines.
  • AGENTS.md, the client-neutral instruction file other coding agents read, from the project root and its ancestors. If a directory has both files and they differ, typed loads both (other tools stop at the first), so keep the pair identical if you switch between tools. AGENTS.md files below the directory you start in are not read; start in the package directory to pick one up.
  • settings.json permission rules: the same allow / deny / ask arrays and Tool(prefix*) syntax, merged from the same files.
  • Skills (.claude/skills), slash commands (.claude/commands), agents (.claude/agents), and hooks (nine events, with the same stdin-JSON and exit-code semantics).
  • Agent Skills standard skills from .agents/skills (project) and ~/.agents/skills (user), the cross-client layout other coding agents install into, including grouping folders and the standard frontmatter. Skills with a description are also catalogued for the model, which loads one by reading its SKILL.md when a task matches.
  • MCP servers from .mcp.json and .claude.json (stdio, Streamable HTTP, and SSE, with the same local > project > user precedence; a project .mcp.json loads once you trust the project), and CLAUDE_CONFIG_DIR everywhere ~/.claude would apply. A server that changes its tools mid-session (notifications/tools/list_changed) is re-listed before the next request.
  • The headless interface: -p, --output-format json, with the same envelope field names your scripts already parse.

What's different. The first-party client exists because some things work better when the client and the backend are designed together:

  • Deterministic permissions. Tool gating is rules-only, no model-in-the-loop safety classifier. Denials are reproducible and name the exact rule; one keystroke at a gate saves an allow rule for that command shape to your project settings.
  • /undo. Per-session shadow snapshots are taken before the first write to each file, so you can rewind a bad run even on a dirty tree, without git ceremony.
  • Live token HUD. The status line shows the model tier, the run state and the live session token counts (<in> in / <out> out) as the server reports them per request, not a client-side estimate; /status reports the session in tokens too. No surface in the CLI prints spend in dollars.
  • Verification-cycle cap. The loop short-circuits the repeat-the-same-check spiral (re-running an identical command or re-reading an unchanged file past a small budget) and steers the model to either conclude or make a change. On by default; tune or disable with TYPED_VERIFY_CYCLE_CAP.

First-run note: the typed CLI asks before using project config it discovers, and hooks are opt-in at that prompt (they execute arbitrary commands, so they never run silently). Full reference: typed doctor, typed --help, and the typed CLI reference.


2. What works identically

  • Most coding workflows: refactoring, debugging, code generation, code explanation, test writing, doc writing.
  • Image input: paste screenshots, design mockups, error messages. typed accepts the same multimodal content shape Claude does.
  • Context window: 1M tokens on both hosted tiers. The same 1M applies to image-bearing (multimodal) requests, with no opt-in header, no model change, and no per-request configuration. (The free local tier is bounded by your own server -- see "Context windows by model" above.) Claude's current-generation Opus 4.6/4.7/4.8 and Sonnet 4.6 include a 1M window at standard pricing; current Haiku 4.5 and legacy Sonnet 4.5 cap at 200K.
  • MCP servers: if your client already mounts MCP servers against Claude, the same configuration works against typed.
  • Prompt caching: available on both hosted tiers. Most reliably engages on typed++; typed caches on a best-effort basis. (The free local tier caches in your own llama-server, not in ours.)

3. What's different

typed is a frontier-class model, not Claude.

  • Edge cases will differ. Most coding work feels identical; some won't. We encourage spot-checking on the workflows that matter to you.
  • Knowledge cutoff varies by underlying model. For recent libraries, frameworks, and APIs published after the model's training, paste relevant docs into the prompt.
  • No Claude artifacts (browser-rendered code previews). Keep Claude for that surface if you rely on artifacts heavily.
  • Billing structure differs. typed bills monthly. By default, requests past quota return a 429 with a one-click top-up prompt; opt in to auto-top-up in your dashboard if you'd rather skip the prompt and auto-charge instead. Claude resets every 5 hours and weekly. Different shapes; better fit for some workflows.
  • Coding-only product surface. typed is tuned for coding workflows. General chat and creative writing will work but are not the design target.

4. Sales-final policy (clearly disclosed)

All sales final. Cancellations take effect at the next renewal. The remainder of your paid period is served normally, and your API key keeps working until the period ends.

We process discretionary refunds on a case-by-case basis for billing errors or extended service outages on our side. Email [email protected]. We do not promise an automatic refund window because cost-of-goods scales with usage, and we would rather be honest about the economics than build the refund into the price.

If you are not sure typed will work for your workflow, the right thing to do is pay for one month, try it for a week, and cancel before renewal if it does not suit you. The remaining three weeks of that month still serve normally. That is the trial.


5. Getting started

  1. Sign up at app.typed.cloud and pick a plan.
  2. The dashboard shows a one-line installer for the typed launcher (with your API key already inlined). Copy/paste, run, done.
  3. Run typed (or t) from any project directory. Usage appears in your dashboard within a few seconds of the first request.

If you would rather configure manually, set the three env vars from section 1 (ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL) in your shell or your client's config and run your normal client.

If you hit problems, email [email protected]. Include your account email and the approximate UTC timestamp of the failed request. That is enough for us to find the trace.


Last updated 2026-09-21.