LayerLens Documentation
Welcome to LayerLens documentation. Three customer-facing experiences, one platform — pick a path and ship faster.
LayerLens is the AI evaluation platform for teams that need to know — before customers do — whether a model, prompt, or agent is good enough to ship. Browse the world's largest public catalog of model evaluations, run private evaluations on your own data, score traces with deterministic rules and LLM graders, and ship AI you can trust.
Three customer-facing experiences
app.layerlens.ai — anonymous browsing of 175+ models, 52+ benchmarks, and 2,000+ public evaluations. Compare any two models head-to-head. No sign-up required.
pip install layerlens --extra-index-url https://sdk.layerlens.ai/package — the Python client SDK (v1.3.0) for programmatic evaluation, grader orchestration, and trace ingestion. 138 sample programs.
Pick a path
I'm a researcher
Explore 175+ models against 52+ benchmarks, browse 2,000+ public evaluations.
I'm an admin or buyer
Set up your org, manage seats and credits, evaluate enterprise readiness.
The LayerLens Workflow
Five stages from raw model to governed production. Every Premium capability maps to one of them.
What's new
Q1 2026 model leaderboard — refreshed quarterly with GPT-5.3, Claude Opus 4.6, Gemini 3.1 Pro/Flash, and 200+ more.
Agentic evaluations — pre- and post-deployment quality gates for multi-step agents. Read the announcement.
GEPA grader optimization — automatically tune your LLM graders against ground-truth labels. Read the concept.
Need help?
In-app: click the Assistant icon (Premium only) for context-aware help.
Email: [email protected]
Docs feedback: open an issue or reach out via the in-app feedback form.
Where to next
New here? → What is LayerLens?
Want to ship today? → Getting Started
Researching for purchase? → How LayerLens compares
Building a business case? → Pricing
Last updated
Was this helpful?