116 detection rules · Semantic intent analysis · Proven on production systems

Your AI agent's output will fail in production.

A malformed field. A credential in a code block. A tool response that injects a command. A schema that drifts. These aren't theoretical — they're happening right now in agent systems that look fine in demos.

Free · No obligation

Get a free 1-page scan summary

Send your agent endpoint, repository, or an exported trace log. You get back a concise summary of what an automated surface scan finds — and where a manual Correctover audit goes deeper.

Get my free scan summary →

Correctover — AI Reliability

Ready for the full manual audit? I test your agent's output layer across 7 dimensions with 116 detection rules. You get a PDF report with every issue, reproduction steps, and concrete fixes — all async over email, no meetings.

Request the full audit →

What I test

Structure & Schema

Every output path that can break downstream parsers — malformed JSON, missing fields, type mismatches, enum drift, nested object errors. Tested across normal, edge-case, and adversarial inputs.

Security

Credential leakage in outputs (AWS/GCP/Azure/SSH tokens in code blocks), prompt injection through tool responses, unsafe content pass-through, command injection vectors. Not keyword matching — semantic intent analysis that traces whether exec() is sandboxed, whether dictionary keys are constant, whether a code path is a legitimate engine or an exploit.

Identity & Integrity

Agent spoofing, unverified tool call sources, missing identity attestation, output tampering, response injection. If your agent can't prove who it is and that its output wasn't modified, it can't be trusted with production systems.

Latency & Cost

P50/P99 outliers under load, unbounded response times, token explosion on long contexts, loop amplification, budget breach paths. Agents that work in demos can cost 10x expected in production.

Why this isn't a checklist

Anyone can read a public standard and build a checkbox. What sets this audit apart:

Black-box audit — no source code required

"No source code needed — ship us the published package, we audit it."

Proprietary agent, closed-source MCP server, or a vendor who can't share their repository? The audit runs on what you actually ship and deploy — an npm package tarball, a Docker image, or a compiled binary. Minifiers rename local variables; they don't rename API calls, permission checks, error strings, environment variable names, or URLs. The security semantics survive the build.

Measured on real published artifacts (September 2026 feasibility test, no source code used at any point):

0 of 14 checks failedAgainst a 9.36MB minified commercial CLI bundle: 8 of 14 automated surface checks fully effective, 5 partially effective, 1 converted to vendor-SBOM review — none fully failed.
86,408 · 268 · 182String literals, distinct environment variable names, and outbound URLs all recovered in plaintext from the same minified bundle.
21MB extractedPlaintext JavaScript recovered from inside a 206MB compiled native binary (Bun build) — 240,000+ readable lines through the same audit pipeline.

Concrete security mechanisms were reconstructed entirely from published packages: sandbox isolation setup, subprocess command-injection guards, SSRF preflight logic, credential flow and environment handling, and the full MCP protocol state machine — each conclusion backed by extracted code excerpts in the report. Even stripped native modules yield a supply-chain bill of materials from binary strings: dependency versions, toolchain fingerprints, build provenance.

Honest coverage boundaries

Black-box review covers 80–95% of the JavaScript, configuration, and protocol layers, at roughly 2× the time of source-based review. Deeply native logic in compiled Rust/C++ (beyond strings, symbol tables, and boundary enumeration) sits at 30–40%. Every report states explicitly which conclusions come from the published artifact and which would benefit from supplemental materials — a lockfile, a Docker image, MCP config samples, source maps, or runtime logs. Those are optional enhancers; the audit does not require them to start.

Send a package for black-box audit →

Certification for allowlists — signed third-party evidence

"Scanners flag. Certificates clear."

Enterprises are starting to run automated filtering and allowlists over the skills, plugins, and MCP servers their agents are allowed to install. One vendor reports its platform filters out about 27% of the add-ons and skills it finds online (TechCrunch); researchers measured 17,800+ public AI add-ons across 6.7 million installations pulling instructions from unverified external sources (SiliconAngle). Automated filters flag components by signal — they don't issue archivable evidence that a specific version was independently reviewed.

If your component is blocked or stuck in a customer's allowlist review, a scanner verdict isn't something you can appeal with. Independent third-party audit evidence is. Correctover audits the exact published artifact under our 116-check manual audit methodology — source code not required — and issues a signed CCS receipt:

For component vendors: ship us the published npm package, Docker image, or binary and get a signed receipt your customers' security teams can verify on their own. Request certification for your component →

For enterprise security teams: independent sampling audits plus non-repudiable cryptographic evidence answer the compliance question scanners leave open — who verified the verifier. Ask about independent sampling audits →

Third-party figures per TechCrunch (Sept 1, 2026) and SiliconAngle (Sept 1, 2026). The CCS receipt format follows the Correctover Conformance Shape, published as an individual Internet-Draft — not an RFC, and not an IETF endorsement.

What you get

Process

1

Day 1 — Scope

You send sample agent outputs, tool definitions, and expected schemas. Or I test against your staging endpoint.

2

Days 2–4 — Deep verification

116 rules across all 7 dimensions, including adversarial inputs, edge cases, and semantic code path analysis.

3

Day 5 — Delivery

Full report with findings, fixes, and verification profile.

Guigui Wang — Correctover

I build runtime verification for AI agent systems. My verification engine (ccs-verifier on PyPI) runs at sub-millisecond latency (P50 <10µs) and has been integrated into production agent frameworks. I've contributed security fixes to EMILIA (PR merged) and CrewAI (PR open). The underlying conformance framework is archived on Zenodo (DOI 10.5281/zenodo.21783723), while the 116 rules and semantic analysis engine are proprietary — built from testing real systems, not from reading specifications.