Your app, tested by an agent you own.

Suitest is an open-source, MCP-native QA platform. Generate, run, and publish end-to-end tests with video evidence. Self-hosted.

Latest: launcher 0.15.0 ยท MCP server 0.15.0 โ€” release notes

suitest

Running in one command.

One command boots the whole platform on your laptop โ€” dashboard, SQLite, and your IDE agent wired in. Scale up to Docker or Helm when the team does.

npx @suiflex/suitest onboard

Full platform on your laptop: web dashboard + SQLite + IDE MCP wiring. Connect your workspace model in Settings before running tests. Node 18+ and uv.

Prompt to published report.

One conversation in your IDE becomes a full test lifecycle.

  1. Prompt

    Tell your IDE agent "test this app with Suitest". The MCP server analyzes your app or crawls a URL.

  2. Execute

    The run engine drives a real browser through every generated case. Deterministic, repeatable, CI-friendly.

  3. Publish

    Cases, runs, video recordings, and traces land in your dashboard. Failures come with evidence attached.

Failures your agent can actually fix.

Evidence is recorded for humans. get_failure_context serializes it for the model that is going to repair the code.

get_failure_context3.1 KB of 8 KB budget
## checkout-flow: FAIL at step 4/7

**Step 4**: click "#submit-btn"
**Error**: TimeoutError: locator not visible after 30s

**DOM at failure** (excerpt around selector):
<form id="checkout">
  <button id="submit-button" disabled>Pay now</button>
</form>

**Console** (errors and warnings only):
[error] POST /api/cart 500 (Internal Server Error)

**Network** (non-2xx only):
POST /api/cart returned 500

**Evidence**: video, screenshot, trace

Sample payload. Two bugs in one read: the id issubmit-button, not submit-btn, and the cart API is returning 500.

  1. A test fails

    run_tests comes back 9/10. Most tools stop here and hand you a red X.

  2. The agent asks why

    get_failure_context returns a budgeted markdown bundle: the failed step, a DOM excerpt around the selector, console errors, and non-2xx requests. Under 8 KB, built for a context window.

  3. The agent fixes your code

    Wrong selector and a 500 from the cart API, both visible in one payload. The agent edits the app, not just the test.

  4. Re-run until green

    Deterministic replay confirms the fix. The run, video, and history publish to the dashboard.

  5. How the failure bundle works

Watch a real run.

A PRD-generated suite executing against the bundled demo app: API and browser steps and screenshots included. Connect your workspace model before replaying it.

Suitest demo: a PRD-generated test suite runs green against the Brewly demo app, showing live logs, API steps, and browser steps with screenshots
Replay it locally: make demo, then openlocalhost:3000.

Features

Built for the way agents work now.

MCP-native, IDE-first

Claude Code, Cursor, and Codex talk to Suitest through one MCP server. Analyze, generate, run, and publish without leaving the editor.

{
  "mcpServers": {
    "suitest": {
      "command": "npx",
      "args": ["-y", "@suiflex/suitest-mcp"]
    }
  }
}

Black-box, gray-box, white-box

Test a URL with no repo access, generate from your source, or run the pytest, Vitest, or Jest suite you already have. One TCM, and every case says which one it was.

Evidence on every run

Each case ships with a video recording and traces. A failing test is a bug report you can watch.

Deterministic runs

Runs use the validated workspace model and remain fully replayable. Same input, same result, every time.

A real test-management UI

Cases, runs, steps, and history live in a dashboard your whole team can read, not in one developer's terminal.

24 tools. One MCP server.

Everything the agent needs to take a test suite from first crawl to published report, exposed as plain MCP tools.

Lifecycle

Analyze a repo, generate cases, run them, read failures, publish.

  • bootstrap_project
  • analyze_project
  • generate_test_cases
  • generate_backend_tests
  • generate_frontend_tests
  • run_tests
  • run_backend_tests
  • run_frontend_tests
  • get_failure_context
  • generate_report
  • sync_tcm

Blackbox

Test any web app from a URL and credentials. No repo access needed.

  • blackbox_discover_app
  • blackbox_detect_login
  • blackbox_perform_login
  • blackbox_crawl_routes
  • blackbox_analyze_page
  • blackbox_build_interaction_graph
  • blackbox_generate_playwright_tests
  • blackbox_run_playwright_tests
  • blackbox_collect_evidence
  • blackbox_publish_results
  • blackbox_summarize_findings

White-box

Run the pytest, Vitest, or Jest suite already in the repo and publish it as cases, runs, and coverage.

  • whitebox_discover_tests
  • whitebox_run_tests

Arguments, return envelopes, and examples for every tool live in theMCP tool reference.

Runs where your code runs.

Laptop, server, cluster, or nothing but a CI runner. Same engine, same config file.

Local bundle

A solo developer, one laptop

The full platform in one command: web dashboard, API on SQLite, run supervisor, and your IDE's MCP config. No Docker; connect a workspace model before running tests.

npx @suiflex/suitest onboard
Local bundle guide

Docker Compose

A team on its own server

The whole platform: Postgres, Redis, MinIO, API, web dashboard, and runner. Prebuilt images are pulled from ghcr.io. Your data never leaves your infrastructure.

git clone https://github.com/suiflex/suitest && cd suitest
cp .env.example .env && make docker-up
Docker guide

Kubernetes

Platform teams, air-gapped environments

Helm chart for production, including cluster-hosted models for generation and diagnosis with zero egress.

helm install suitest infra/helm/suitest \
  -f infra/helm/suitest/values.yaml
Kubernetes guide

CI only, no server

Projects that just want a merge gate

The GitHub Action runs the lifecycle inside the workflow and reports to the PR. No Suitest server required at all.

- uses: suiflex/suitest/action@main
  with:
    config: suitest.config.json
CI guide

One product. Bring your own model.

Provider location does not change the product. Validate the workspace model once, then Suitest uses it for MCP execution and AI workflows.

MANAGEAvailable immediately

Sign in, organize workspaces, and manage test cases before connecting a model.

CONNECTBring your model

Save any supported hosted or self-hosted provider in Settings, then validate it once.

EXECUTELLM ready

Run MCP tools, execute suites, generate tests, and diagnose failures with the same feature set.

Every pull request gets a verdict.

One composite action runs the suite, upserts a single PR comment with per-failure excerpts, and gates the merge with its exit code.

.github/workflows/e2e.yml
name: e2e
on: pull_request

jobs:
  suitest:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npm run start & npx wait-on http://localhost:3000
      - uses: suiflex/suitest/action@main
        with:
          config: suitest.config.json
PR comment (updated in place, example)

Suitest: 9/10 passed

passauth: login with valid credentials
passdashboard: renders project stats
failcheckout: timeout at step 4, excerpt attached

Evidence and full failure detail in the run report.

  • 0 all passed
  • 1 test failure, PR check goes red
  • 2 infra error, no false verdict

Setup, exit codes, and local preview with --dry-run in theCI guide.

One tool instead of three.

Teams usually stitch together a case manager, a runner, and an AI testing service. Suitest is the union of those roles, and you host it.

CapabilitySuitestTestSpritePlaywright MCPTestRail
Open source, self-hostedYesNoYesNo
Manual TCM before connecting an LLMYesNoYesYes
Test management UI (cases, suites, history)YesPartialPartialNoYes
Deterministic, replayable runsYesPartialPartialYesNo
Agent generates and runs tests from the IDEYesYesPartialPartialNo
Agent-readable failure context for auto-fixYesNoNoNo
Video and trace evidence per runYesYesPartialPartialNo

Capability summary based on public documentation, July 2026. Details and caveats in the FAQ.

Questions, answered.

Longer answers live in the docs FAQ.

Do I need an OpenAI or Anthropic API key?

Not specifically. Manual test management works before an LLM is connected. To use MCP, runs, and AI workflows, connect and validate any supported provider in workspace settings: a hosted API, a self-hosted model, or a supported sign-in provider.

What does Suitest cost?

Nothing. Suitest is Apache-2.0 licensed and self-hosted. There is no hosted service to subscribe to and no feature gate tied to payment.

Which IDEs and agents does it work with?

Any MCP client. Claude Code, Cursor, and Codex are tested first-class: one npx command registers the server, and npx @suiflex/suitest-mcp init writes the config for your IDE automatically. Want the full local platform with the dashboard? npx @suiflex/suitest onboard boots everything and wires your IDE in one step.

Can it test an app I don't have the source for?

Yes. The blackbox engine needs a URL and test credentials. It discovers routes, detects and performs login, builds an interaction graph, generates Playwright tests, and runs them, all without repo access.

What evidence is recorded for each run?

Per-test video, screenshots, DOM snapshots, console logs, network HAR, and execution traces. Failures also produce an agent-readable failure bundle so your coding agent can fix the underlying bug.

Does it replace TestRail or Playwright?

It covers both roles: a test management UI with cases, suites, runs, and traceability, plus a deterministic runner that drives Playwright, HTTP APIs, and Postgres through MCP providers. You can adopt either half first.

What data leaves my machine?

Suitest sends LLM requests only to the provider explicitly configured for the workspace. Choose a self-hosted provider to keep inference on your infrastructure. Provider secrets stay encrypted on the Suitest server and are never sent to the MCP client.

How does it run in CI?

A composite GitHub Action runs the suite, updates a single PR comment with results and failure excerpts, and fails the check when tests fail: exit code 0 for pass, 1 for test failures, 2 for infrastructure errors.

Apache-2.0 licensed.
Self-hosted by default.

Your stack, your LLM, your data. Run the whole platform on a laptop or a cluster.