Continuous, language-agnostic health scans of whole repositories by
specialized LLM agents.
Deterministic scores, tracked findings, and copy-pasteable fix prompts for
your coding agent. Self-hosted.
Quick start · How it works · Documentation
Warning
Proof of concept. CodeTend works end to end and is used on real repositories, but its authentication boundary is single-tenant and the full LLM scan is an external-service integration. Read Limitations before relying on it.
CodeTend is not an autonomous pull-request bot. It registers repositories and, on a schedule or on demand, makes a fresh checkout of the configured branch and runs fifteen specialized scanner agents over it: one per dimension of technical debt, from architecture and duplication to reliability, type safety, AI slop, dependency health and security. Their structured findings become deterministic scores and grades, every finding is tracked across scans, and each dimension produces one aggregated fix prompt you paste into Claude Code, Codex or another coding agent.
- 🧭 Fifteen dimensions, one registry. Architecture & Modularity, Duplication & Abstraction, Dead & Obsolete Code, Complexity & Maintainability, Tests & Testability, Reliability & Error Handling, Documentation, Domain & API Design, Type Safety & Data Contracts, Consistency / Vibe Debt, AI Slop & Noise, Dependencies & Build Health, Vulnerable Dependencies, CI & GitHub Actions Security and Security Hygiene. Each scanner's prompt says what it owns and what its neighbours own. A bounded independent review consolidates verified cross-scanner duplicates before scoring.
- 📐 Numbers the model never touches. Every scanner starts at 100 and loses a fixed penalty per open finding scaled by confidence. Grades A–F follow fixed thresholds. A high-confidence critical open finding caps the overall score at 39/F; security findings require independently confirmed exploitability. The LLM proposes findings, never scores.
- 🔁 Findings with a lifecycle. A finding is a stable fingerprint, not a
line number. On every rescan the scanner verifies open findings within its budget and explicitly defers those it cannot verify,
and the app derives
new → active → improved → resolved → regressedfrom its own persisted results. Resolution requires an explicit re-verification verdict and is never inferred from absence in a bounded scan. - 🛡️ Real vulnerability data. Google's OSV Scanner matches exact lockfile versions; CodeTend computes CVSS locally, joins FIRST EPSS and the CISA KEV catalog, and keeps raw severity separate from contextual priority.
- 🧪 Evidence, not vibes. Security findings carry source/control/sink evidence, an attack path, and optional executable validation that runs only in a networkless, capability-dropped container, never on the host.
- 🩹 Review-gated patches. Generate a unified diff for one finding in a disposable clone, re-run its reproducer, then approve or reject the stored proposal. Download approved diffs and apply them through your normal review and CI process; CodeTend never writes generated patches to the source repository.
- 📊 Cost you can see. Provider-reported tokens per scanner, per scan and per repository, with list-price estimates, a UTC-day budget cap, and concurrency limits.
- 🔒 Yours, end to end. PostgreSQL you run, any OpenAI-compatible model endpoint you choose, your own login gate in front, no telemetry.
- 🤖 Ask your coding agent. An optional OAuth-protected MCP endpoint lets Codex and other clients inspect repositories, scans, findings, and knowledge; a separate owner-approved scope can start or cancel scans.
Repository health at a glance. All data shown is illustrative.
Drill into score trends, findings and the repository knowledge model.
- Add a repository from the repositories available to the configured GitHub App or token, or use the URL fallback for another Git host.
- Scan now, or configure the global cron and queue cooldown in Settings. Choose an input-token investigation budget, whole repository or selected paths, and an optional cost ceiling. Each scanner orients from repository structure and chooses evidence appropriate to its own dimension.
- Watch the phase update live: cloning → indexing → knowledge → scanning → reconciling.
- The repository page shows the overall score and grade, per-scanner scores, new/improved/resolved/regressed counts, active findings, scan history, token and cost usage, and change-over-time charts.
- Each scanner page shows its trend, findings with evidence and recommendations, and the aggregated fix prompt with a copy button.
- Mark a finding fixed with a note describing the change, or mark it false positive or accepted risk with context. A reported fix is verified by the next scan and regresses if rediscovered. Dispositions are durable; a scanner can only reopen one by citing a concrete code change that contradicts its note.
- Download the manifest, findings, investigation report, Markdown report or SARIF for any scan.
When the Helm deployment enables mcp.enabled, CodeTend serves MCP on its own
origin and uses the existing Hodor owner session for OAuth consent. Add it to
Codex with:
codex mcp add codetend --url https://codetend.example.com/api/mcp
codex mcp login codetendThe read scope can inspect repositories, scans and findings and generate a current coding-agent fix prompt for a repository scanner. The separately granted write scope can start or cancel scans, mark findings as fixed, mark them as false positives or accepted risks, and reopen manually triaged findings. Use Settings → AI access (MCP) to pause access, enable individual tools, or revoke a client. See configuration for the proxy and security boundary.
flowchart TB
subgraph INTAKE["1 · CodeTend app — intake"]
direction LR
TRIGGER["Schedule or Scan now"] --> QUEUE["Create scan and enqueue it"]
end
QUEUE -->|authenticated request| RUN
subgraph EVE["2 · Eve runtime — clones, sandboxes and model calls"]
direction LR
RUN["Durable, resumable scan"] --> PREP["Clone and build context<br/>GitNexus · knowledge · dependency graph<br/>budgeted file sample"]
PREP --> SCANNERS["15 specialized scanners in parallel<br/>with isolated executable validation"]
PREP --> OSV["OSV audit of exact<br/>lockfile versions"]
SCANNERS --> BUNDLE["Structured evidence<br/>coverage and cost usage"]
OSV --> BUNDLE
end
BUNDLE -->|sealed scan result| PROCESS
subgraph RESULTS["3 · CodeTend app — results in PostgreSQL"]
direction LR
PROCESS["Enrich and prioritize<br/>Reconcile duplicates<br/>Score deterministically"]
PROCESS --> OUTPUTS["Dashboard and finding lifecycle<br/>Fix prompts · SARIF · Markdown"]
end
The app and the Eve runtime are separate processes sharing
a data directory and an authenticated HTTP link. The app owns PostgreSQL; Eve
owns clones, read-only sandboxes and model calls. Scanners get read_file,
glob, grep, a virtual shell, and optionally
GitNexus code-intelligence
tools over MCP. Details in docs/architecture.md.
| Layer | Choice |
|---|---|
| App | Bun, TanStack Start + Router, React 19, shadcn/ui (Base UI), Tailwind v4 |
| Persistence | PostgreSQL + Drizzle with committed migrations |
| Agent runtime | Eve: durable workflow tools, declared subagents, just-bash sandbox, MCP connections, structured outputs |
| Models | Any OpenAI-compatible endpoint; Claude and others through a gateway such as OpenRouter |
| Vulnerability data | Google OSV Scanner, FIRST EPSS, CISA KEV, local CVSS 2.0–4.0 |
| Code intelligence | GitNexus over MCP, optional |
| Tooling | Vite, Biome, Knip, bun test, import-boundary check |
| Deployment | Docker Compose; multi-arch images on GHCR |
Docker Compose starts PostgreSQL, applies migrations, starts the Eve runtime and the app.
cp .env.example .env
# Replace every change-me value and set OPENAI_API_KEY.
docker compose up --buildOpen http://localhost:3000. The bundled topology publishes only the app on loopback; PostgreSQL and Eve stay on the Compose network. The app itself has no login of its own — put a login-gating proxy or another trusted network boundary in front before exposing it beyond loopback.
Prerequisites: Bun 1.4, Node 26 (for Eve), Git, PostgreSQL (Docker is fine),
an API key for an OpenAI-compatible endpoint, and
OSV Scanner. Optionally
GitNexus: npm i -g gitnexus or mkdir -p .tools && (cd .tools && npm i gitnexus).
cp .env.example .env
bun install && (cd eve && npm install)
docker compose up -d postgres # or point DATABASE_URL at your own database
bun run eve:build && bun run eve:start # terminal 1: Eve runtime on :2000
bun run dev # terminal 2: app on :3000, migrations applied automaticallyscripts/start-dev.sh supervises both processes from one terminal and
restarts Eve in place after a rebuild. scripts/start-preview.sh --build
serves the production build instead.
All available commands
| Command | Purpose |
|---|---|
bun run dev |
App with automatic migrations |
bun run preview:serve |
Production build + Eve for a stable shared preview |
bun run check |
Biome formatting/lint and import boundaries |
bun run test |
Unit tests for scoring, pricing, vulnerabilities, artifacts |
bun run test:database |
Migrated-schema and active-scan constraint smoke test |
bun run typecheck |
Type-check the app |
bun run eve:typecheck |
Type-check the Eve agent |
bun run lint:deadcode |
Find unused code and dependencies with Knip |
bun run build |
Build production assets |
bun run verify |
Fast local gate: formatting, tests, types, build and dead-code checks |
bun run eve:build / bun run eve:start |
Compile and serve the Eve agent |
bun run db:generate |
Generate a migration after editing src/db/schema.ts |
bun run db:migrate |
Apply committed migrations |
bun run scripts/cli.ts add <name> <url> [branch] |
Scripting helper |
CI additionally validates the Compose and Helm deployment manifests, audits production dependencies, applies migrations to PostgreSQL, runs the database lifecycle suite, and checks that the generated route tree is current. Run those environment-dependent checks when changing deployment, dependency, database, or route-generation behavior.
| Guide | What is inside |
|---|---|
| Architecture | Scan flow, scanners, scoring, finding lifecycle, vulnerability enrichment, patches, knowledge, cost tracking |
| Configuration reference | Every environment variable and its default |
| Operations | Deployment shape, validation boundary, health, backup, monitoring |
| Data handling | What is processed, sent to providers and retained |
- The app has no login of its own; it expects a login-gating proxy or equivalent network boundary in front. A multi-customer deployment needs an identity provider, organizations and repository-level authorization on top of that.
- CodeTend mints and caches its own GitHub App installation tokens from
GITHUB_APP_ID/GITHUB_APP_PRIVATE_KEY, resolving the installation for each repository; it also accepts a plain, externally-issuedGITHUB_TOKEN. - Executable finding validation is disabled by default. When enabled, Eve must have Docker access; the bundled Compose deployment deliberately does not mount the Docker socket.
- The scanner sandbox is
just-bash(virtual shell, no real toolchains, no network). Switcheve/agent/sandbox/sandbox.tstodocker()when scanners must run real language tooling. - Scanners sample large repositories; very large monorepos may exceed a single scan's context and should be split.
- GitNexus supports a fixed set of languages; other languages are analyzed from the filesystem only.
- The full LLM scan is an external-service integration. Deterministic parts (vulnerability parsing, priority, scoring, repository access, accounting, database invariants) are covered by tests without spending model credits.
Open-source software for teams who would rather measure their debt than
argue about it.
Star it on GitHub

