Reliability Diagnostic
One system, 3-5 days
I map how it fails: what breaks, under what load, what it costs you when it does, and the order to fix it in. You get a prioritised failure-mode report and a call to walk through it.
A seat on a small team shipping AI product, or one system of yours made reliable in weeks.
Email me or book a call.
A staff, senior or founding seat on a small remote team shipping AI product, where I own the problem end to end: working out what to build, building it, running it and carrying the pager. As an employee or through my company, whichever fits the team. Co-founding is a real option too, for a problem we'd both want to own for years.
One system, 3-5 days
I map how it fails: what breaks, under what load, what it costs you when it does, and the order to fix it in. You get a prioritised failure-mode report and a call to walk through it.
One system, one week
The top two or three failure modes from a Diagnostic, fixed and shipped, with a before-and-after note. If you already know what is broken, start here.
One system, 3-4 weeks
The deep version, for a system about to matter: failure-mode report, guardrails implemented, an eval or monitoring suite wired into CI, and before-and-after numbers. For AI systems that means evals and LLM-as-judge calibration, prompt-injection hardening, fail-closed defaults, cost and latency. For Rails, infrastructure and data pipelines it means p95 and cost, CI unblock, legacy rescue, concurrency correctness. Scope and price per SOW.
Prices are on this page because your scarcest resource is a meeting. 50% on signature, balance on delivery. Fixed price means fixed scope: anything discovered mid-mission is quoted, not absorbed. No hourly retainers, and no feature work slipped inside an audit.
The same missions taken as subcontract, billed through you on your rate card, white-label if you want. You keep the relationship; your client gets me rather than a bench and a PM layer. Renewals are new contracts.
All feedback from the team and client has been very positive. Will certainly reach out in the future when we have another opportunity
Amanda Bizzinotto, Director of Operations (OmbuLabs)Staff-level: the engineer you trust with what nobody owns. I find what needs doing, build it, run it and carry the pager, and the team around me gets faster too: at Hivebrite, deploys doubled, builds halved and the CI bill came down €160k a year for 40 engineers. When something breaks I follow it to the root, even upstream, with fixes merged in Rails, Rack, RuboCop, Sentry's Ruby SDK and NextDNS's Go resolver. Ten years on million-user platforms; lately, LLM systems that run in production on their own.
Europe-based and fully remote, working async-first. I keep live overlap across European, US, and APAC business hours through the year, so a distributed team gets real-time hours regardless of where it sits. Distributed is my default, not an adjustment: I grew up across six countries on four continents, in the French school system throughout, natively bilingual in English and French, and have worked from five continents since.
Observability tells you what your agent did; evals tell you whether it was good, and most teams have the first without the second. The work is closing that gap: tracing every model call, a graded eval suite (schema and format assertions first, then a validated LLM-as-judge measured against human labels), schema-validated output instead of trusting prompts, retries and guardrails, and only then an automatic improvement loop. The order matters: optimise against an unvalidated judge and you build a system confidently tuned toward being wrong. The goal is an agent or pipeline that runs unattended without quietly breaking - one that says 'I don't know' instead of guessing.
Email me, or book a call. More context in the CV.