Availability · open now

Available for a full‑time seat, or a fixed‑scope mission.

A seat on a small team shipping AI product, or one system of yours made reliable in weeks.

Email me or book a call.

Full-time seat

A staff, senior or founding seat on a small remote team shipping AI product, where I own the problem end to end: working out what to build, building it, running it and carrying the pager. As an employee or through my company, whichever fits the team. Co-founding is a real option too, for a problem we'd both want to own for years.

Proof, checkable now

Fixed-scope missions

€2,500 fixed

Reliability Diagnostic

One system, 3-5 days

I map how it fails: what breaks, under what load, what it costs you when it does, and the order to fix it in. You get a prioritised failure-mode report and a call to walk through it.

€6,000-8,000 fixed

Hardening Sprint

One system, one week

The top two or three failure modes from a Diagnostic, fixed and shipped, with a before-and-after note. If you already know what is broken, start here.

From €18,000

Production Reliability Audit

One system, 3-4 weeks

The deep version, for a system about to matter: failure-mode report, guardrails implemented, an eval or monitoring suite wired into CI, and before-and-after numbers. For AI systems that means evals and LLM-as-judge calibration, prompt-injection hardening, fail-closed defaults, cost and latency. For Rails, infrastructure and data pipelines it means p95 and cost, CI unblock, legacy rescue, concurrency correctness. Scope and price per SOW.

Terms

Prices are on this page because your scarcest resource is a meeting. 50% on signature, balance on delivery. Fixed price means fixed scope: anything discovered mid-mission is quoted, not absorbed. No hourly retainers, and no feature work slipped inside an audit.

For consultancies

The same missions taken as subcontract, billed through you on your rate card, white-label if you want. You keep the relationship; your client gets me rather than a bench and a PM layer. Renewals are new contracts.

All feedback from the team and client has been very positive. Will certainly reach out in the future when we have another opportunity

Amanda Bizzinotto, Director of Operations (OmbuLabs)

What I take on

Best fit

  • An AI feature that works in the demo, then breaks or quietly bleeds cost under real load
  • Taking an AI/LLM prototype to production reliability: evals, observability, schema-validated output, guardrails
  • Hardening agentic and MCP systems that break once nobody is watching
  • Rails, infrastructure, or data-pipeline reliability: p95 and cost, CI unblock, legacy rescue, concurrency correctness

Not a fit

  • On-site or relocation roles
  • Pure ML research or model training
  • Staff augmentation through an agency, with no ownership of the outcome
  • Maintenance-only work with no ownership

Practicalities

Engagement
A staff, senior or founding engineering seat, fully remote, as an employee or through my company; co-founding a company; or fixed-scope B2B missions (defined start and end, own tools, own hours)
Location
Remote, Europe-based, async-first
Timezone
Live overlap across European, US, and APAC hours through the year

Questions

What kind of engineer are you, and at what level?

Staff-level: the engineer you trust with what nobody owns. I find what needs doing, build it, run it and carry the pager, and the team around me gets faster too: at Hivebrite, deploys doubled, builds halved and the CI bill came down €160k a year for 40 engineers. When something breaks I follow it to the root, even upstream, with fixes merged in Rails, Rack, RuboCop, Sentry's Ruby SDK and NextDNS's Go resolver. Ten years on million-user platforms; lately, LLM systems that run in production on their own.

Where are you based and which time zones do you cover?

Europe-based and fully remote, working async-first. I keep live overlap across European, US, and APAC business hours through the year, so a distributed team gets real-time hours regardless of where it sits. Distributed is my default, not an adjustment: I grew up across six countries on four continents, in the French school system throughout, natively bilingual in English and French, and have worked from five continents since.

What does taking an LLM prototype to production reliability involve?

Observability tells you what your agent did; evals tell you whether it was good, and most teams have the first without the second. The work is closing that gap: tracing every model call, a graded eval suite (schema and format assertions first, then a validated LLM-as-judge measured against human labels), schema-validated output instead of trusting prompts, retries and guardrails, and only then an automatic improvement loop. The order matters: optimise against an unvalidated judge and you build a system confidently tuned toward being wrong. The goal is an agent or pipeline that runs unattended without quietly breaking - one that says 'I don't know' instead of guessing.

Get in touch

Email me, or book a call. More context in the CV.