Skip to content

    The inference layer for the agentic economy.

    One key and one model-agnostic endpoint for every agent you run — Claude Code, Cline, Aider, OpenCode, or your own. OpenAI-compatible, with native Anthropic Messages.

    OpenAI · Anthropic · xAI · Moonshot • Try it for free

    The problem

    Who reads the error?

    A person who sees "service unavailable" waits a minute, checks a status page, or fixes the typo. An agent can't. It obeys the status code, and a 503 says try again, so it retries into its budget.

    A typical gateway, for a typo'd model id

    retryable
    HTTP/1.1 503 Service Unavailable
    
    {
      "detail": "service unavailable, contact support"
    }

    Reads as an outage. The agent backs off and tries again, and again.

    Gatewayz, for the same request

    terminal
    HTTP/1.1 400 Bad Request
    
    {
      "error": {
        "type": "invalid_request_error",
        "code": "model_not_found"
      }
    }

    A client error with a stable code. The agent stops, and your logs say why.

    How we build it

    Three rules the gateway keeps

    Each one is something you can check against the API, not a promise on a page.

    Model agnostic

    We don't sell a model of our own, so we have no reason to steer you toward one.

    An alias resolves to its exact snapshot. An unknown id gets a 400, never a substitute you're billed for.

    model resolution
    "claude-sonnet-4-5"  ->  claude-sonnet-4-5-20250929
    "claude-sonet-9"     ->  400 model_not_found

    Errors your software can act on

    Nobody is watching when an agent runs, so every failure has to say what to do next in the status code itself — stop, or try again.

    Ask for a model that doesn't exist and the answer is 400 model_not_found — a typo in the request, not an outage. Spend past a key's cap and the answer is 402 request_cap_exhausted, and an empty balance gives 402 insufficient_credits. Neither clears by waiting, so an agent stops instead of retrying into its budget.

    The two worth retrying look different on purpose: 429 when you are rate limited, and 5xx when a provider fails. And if a provider dies mid-answer, the stream ends with an explicit error event rather than simply stopping, so a half-finished reply is never mistaken for a complete one.

    Stop, or retry

    Stop and change something
    Unknown model, key cap reached, or no credits left. Retrying the same request cannot help.
    Wait and try again
    Rate limited, or an upstream provider failed. The same request can succeed later.

    Under the hood

    How a request flows

    1. 01

      Agent

      Your code, Claude Code, or any OpenAI-compatible client

    2. 02

      Key and cap check

      Valid key, credits left, under its spend cap

    3. 03

      Exact model resolution

      Alias to snapshot, or a 400

    4. 04

      Provider

      Same-model failover, never a different model

    5. 05

      Metered response

      Billed per token, served by a named build

    See how it works →

    Integrate

    Works with your agent

    Two environment variables for Claude Code. One base URL for everything else.

    Claude Code · shell
    export ANTHROPIC_BASE_URL=https://api.gatewayz.ai
    export ANTHROPIC_AUTH_TOKEN=<your Gatewayz key>
    claude

    Claude Code (native Anthropic Messages)

    Running an organization made of agents? FlashyOS's agent CLI is built on Gatewayz

    Your own agent works the same way: set the base URL, keep your code.

    OpenAI SDK · Python
    from openai import OpenAI
    
    client = OpenAI(
        base_url="https://api.gatewayz.ai/v1",
        api_key="<your Gatewayz key>",
    )
    
    reply = client.chat.completions.create(
        model="anthropic/claude-sonnet-4-6",
        messages=[{"role": "user", "content": "Hello"}],
    )
    print(reply.choices[0].message.content)

    Cline, Aider, OpenCode, Continue — via the OpenAI-compatible endpoint

    Simple integration

    Three lines to change, then you're on

    Claude Code and every model in the catalog work with the Gatewayz API. Point any OpenAI-compatible client at one base URL and use one key. Change providers by changing the model string — the rest of your code stays as it is.

    curl · shell
    curl https://api.gatewayz.ai/v1/chat/completions \
      -H "Authorization: Bearer $GATEWAYZ_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "anthropic/claude-sonnet-4-6",
        "messages": [{"role": "user", "content": "Hello"}]
      }'

    Prefer the Anthropic format? Native Messages lives at /v1/messages, with tool use and prompt caching passed through.

    One base URL
    Anything that speaks the OpenAI API works unchanged, including the official SDKs.
    One key, every provider
    One balance and one bill across OpenAI, Anthropic, xAI and Moonshot.
    Swap models in one string
    Ask for a model by name and you get that model, or a clear error — never a quiet substitute. How resolution works

    Autonomous organizations

    The inference layer for autonomous organizations

    An organization made of agents makes every decision through an inference call. Gatewayz is where those calls run.

    FlashyOS's agent CLI ships Gatewayz as its only inference provider.

    FlashyOS agent CLI · shell
    npx @flashyos/agent init --inference gatewayz

    Any autonomous organization on the FlashyOS network can run on the same endpoint.

    • A key and a cap per agent

      Each agent or role gets its own key and request cap, so one runaway agent can't drain the rest.

    • Errors an agent can act on

      An unknown model is a 400 and a spent cap is a 402. Neither looks like a failure worth retrying.

    • Usage attributed to the work

      Tag a request with x-gatewayz-tag, then read calls, tokens and cost for that tag from /v1/usage.

    Pricing

    Pricing, honestly

    • Per token, at the provider's list price plus a routing fee.
    • One balance across every provider we serve.
    • Per-key caps, so an agent can't run past its budget.

    See pricing →

    Security & data

    What we keep

    • Plain API calls store no prompt or completion content.
    • We keep billing metadata only: key, model, tokens, cost.
    • Retention windows are published.

    Read security & data →

    Frequently Asked Questions

    Everything you need to know about the Gatewayz inference layer

    Still have questions? Contact our team →

    Give your agents one key.

    Create a key, set a cap, and point your agent at the gateway.