XAKERBANK · PUBLIC BENCHMARK

Бенчмарк, якому
є що доводити.

Реальні анонімізовані інциденти, ізольовані fixtures та три незалежні reviewer-агенти для кожної відповіді.

14 інцидентів3 reviewers0 runtime secrets
01

Каталог інцидентів

Не toy-задачі — мінімальні відтворення реальних відмов

SB-001medium

Silent skip у фоновому asyncio-консюмері

Пояснити відсутність фінальних логів без вигаданого зависання чи пропущеного await.

pythonasynciologgingdeduplication
SB-002easy

AI-транскрипт дописано в кінець Python-файлу

Знайти точну причину PartialParsing і запропонувати prevention gate.

pythonsyntaxtool-callpre-commit
SB-003medium

Правильний pre-commit guard лежить у неактивному каталозі

Пояснити, чому syntax guard існує у репо, але не запускається Git.

githookspythonconfiguration-drift
SB-004hard

Портал показує стару confirmed-знахідку після incomplete scan

Відрізнити stale verified fallback від актуального стану коду.

securitystatefreshnessaudit-pipeline
SB-005hard

SQL identifier allowlist маскує інші injection paths

Відокремити safe identifier від невалідованого direction і values.

pythonsqlsecuritydata-flow
SB-006hard

OSV приписує репозиторію пакет і версію, яких у ньому немає

Відрізнити CVE коду від систематичного artifact-association bug.

osvsupply-chainscannerprovenance
SB-007medium

Sandboxed scanner job exits 0 every day but never actually scans

A systemd oneshot audit job reports SUCCESS daily, but the underlying SAST scanner has been silently failing on every repo for over a day.

systemdsandboxingsilent-failureciobservability
SB-008hard

Guard рестартує щоцикл, але інстанс ніколи не відновлюється й не ескалює в алерт

Race condition (рестарт без очікування готовності) ховає збереження стану через непіймане виключення — down_streak застряг назавжди, ескалація ніколи не спрацьовує.

pythonrace-conditionexception-handlingself-healstate-persistence
SB-009medium

Dual-AI review dashboard shows CONFLICT when both reviewers found the same bugs

Two independent code-review agents flag identical findings (same id/file/line) but the dashboard still shows a red 'conflict' badge because their severity labels differ by one notch.

comparison-logicfalse-positiveaudit-pipelineobservability
SB-010hard

Fix looked successful, but the next scheduled job broke because the agent used root instead of the service account

An agent with root SSH access fixes a bug, runs the test suite, commits and pushes — everything reports success. Hours later the repo's own scheduled deploy job starts failing with a permission error nobody can explain from the diff.

gitpermissionsroot-accessagent-behaviorsilent-failure
SB-011medium

Marking a finding "resolved" in the tracker didn't make it disappear from the dashboard the operator was actually watching

A known false positive was suppressed in config two days ago and the fix was committed. The finding still shows as unresolved on the public status dashboard. The database row can be flipped to resolved in one query and looks like a fix — but the dashboard is rendered from a completely separate, file-based snapshot of the last successful full pipeline run, and no full run has succeeded since the suppression landed.

observabilitytwo-tier-statesilent-failureagent-behaviorfalse-positive-suppression
SB-012medium

Manually calling the merge function with the right argument "proved" the fix — the real endpoint still hardcodes the old one

A per-day coverage dashboard merges two sources: submissions recorded in a database table and files found by scanning a legacy network drop folder. An earlier incident (the folder scan had been silently disconnected) was "fixed" by adding a line that computes the scanned files into a local variable — but the very next line, the actual call into the merge function, still passes a hardcoded empty list literal instead of that variable. The engineer who verified the fix called the merge function directly from a script, manually supplying the correctly-computed value as an argument, got the right output twice, and closed the ticket. The production endpoint was never exercised end-to-end during verification, so it kept returning the old, wrong result.

silent-failureverification-methodologyagent-behaviorregressiontwo-source-merge
SB-013hard

An unauthenticated request crashes a worker with RecursionError — no valid key, no login, no auth-log entry

A JWT auth dependency's except clause only catches InvalidTokenError, on the assumption that a malformed or forged token always produces one. The installed JWT library parses the token header (json.loads on client-supplied data, no depth limit) before it ever checks the signature — so a deeply nested header alone, with no valid key and no prior login, raises an uncaught RecursionError that is not a subclass of InvalidTokenError. Support sees sporadic worker crashes with an empty auth log for that minute and initially suspects application logic; the real fix is a dependency version bump (the vendor's own changelog fixed exactly this, without spelling out that it was pre-auth reachable).

securityunauthenticated-dosexception-hierarchydependency-cveorder-of-operations
SB-014hard

Two suspects paused, CPU load unchanged — the real culprit was a third process nobody was watching

A union filesystem mounted over a local branch and a remote SFTP-backed branch keeps 20-30% CPU on the remote-storage client and 5-9% on the union mount itself, hours on end. An engineer suspects the two background tree-walkers, pauses both (SIGSTOP) — no change in CPU. Only a syscall-level audit rule (not application logs) reveals every getxattr/lgetxattr call in the window comes from a single pid: the union mount process itself, querying the security.capability extended attribute on every file it touches. The remote branch's backend does not support xattrs, so every query fails (EOPNOTSUPP), and the failure itself is the extra work — not a bug in either suspect, and not visible from either suspect's own logs or strace. The fix is a documented-but-not---help'd mount option with an exact, easily-mistyped name.

performanceroot-causesyscall-tracingred-herringexact-syntaxinfra
02

Лідерборд

Якість окремо від стабільності провайдера

#МодельЯкістьНадійністьConfirmedInfra errorAvg latency
1Claude Sonnet 5100%100%20—
2GLM-5.2100%100%10—
3Gemini Flash100%100%10—
4Mistral Large0%100%00—
5Qwen2.5 Coder 7B0%100%00—
6Llama 3.1 8B0%100%00—
7DeepSeek R1 8B0%100%00—
8Gemma 4 31B0%0%01—
03

Останні відповіді

  • Нові agent-panel результати ще не записані
TRUST MODEL

Один агент не судить іншого одноосібно.

Correctness розмічає атомарні факти. Evidence звіряє file:line. Adversarial впливає лише відтвореним контрприкладом. Reducer лишає replayable trace.