For the complete documentation index, see llms.txt. This page is also available as Markdown.

LayerLens Documentation

Welcome to LayerLens documentation. Three customer-facing experiences, one platform — pick a path and ship faster.

LayerLens is the AI evaluation platform for teams that need to know — before customers do — whether a model, prompt, or agent is good enough to ship. Browse the world's largest public catalog of model evaluations, run private evaluations on your own data, score traces with deterministic rules and LLM graders, and ship AI you can trust.

Three customer-facing experiences

app.layerlens.ai — anonymous browsing of 175+ models, 52+ benchmarks, and 2,000+ public evaluations. Compare any two models head-to-head. No sign-up required.

Start with LayerLens Public →

app.layerlens.ai — after you sign in. The logged-in workspace where teams run private evaluations, build graders, score traces, manage graders, run agentic evaluations, and govern AI quality across an organization.

Start with LayerLens Premium →

pip install layerlens --extra-index-url https://sdk.layerlens.ai/package — the Python client SDK (v1.3.0) for programmatic evaluation, grader orchestration, and trace ingestion. 138 sample programs.

Start with the SDK →

Pick a path

I'm a builder

Ship faster with evaluations wired into your dev loop.

I'm an operator

Run live AI in production with quality gates and continuous evaluation.

I'm a researcher

Explore 175+ models against 52+ benchmarks, browse 2,000+ public evaluations.

I'm an admin or buyer

Set up your org, manage seats and credits, evaluate enterprise readiness.

The LayerLens Workflow

Five stages from raw model to governed production. Every Premium capability maps to one of them.

Select → Build → Observe → Evaluate → Improve. This is the spine. Every concept, how-to, tutorial, and recipe in this documentation pins to one of these stages. Learn the workflow →

What's new

  • Q1 2026 model leaderboard — refreshed quarterly with GPT-5.3, Claude Opus 4.6, Gemini 3.1 Pro/Flash, and 200+ more.

  • Agentic evaluations — pre- and post-deployment quality gates for multi-step agents. Read the announcement.

  • GEPA grader optimization — automatically tune your LLM graders against ground-truth labels. Read the concept.

Need help?

  • In-app: click the Assistant icon (Premium only) for context-aware help.

  • Docs feedback: open an issue or reach out via the in-app feedback form.

Where to next

Last updated

Was this helpful?