Research
Longer-form
investigation.
Extended work on agentic system reliability, evaluation methodology, and the governance frameworks that increasingly shape how these systems get built.
The current line of work
One question runs through everything below: what has to be true before an autonomous system can be put in front of real consequences? Not whether it performs well on a benchmark — whether its failures are bounded, whether anyone can tell what it did, and whether a person still has time to intervene when it matters.
These are analytical frameworks derived from published engineering practice and regulatory text, not empirical studies. Where a claim rests on external evidence, the source is cited in the article.
Reliability & Governance
Reliability
Why Agentic AI Systems Fail at Scale: A Reliability Budget
Per-step accuracy compounds into task failure. Budgeting reliability across the whole path, rather than tuning one step, is the lever that moves.
Reliability
From AI Prototype to Production System: The Architecture of Recoverability
Idempotency, durable event history, replay semantics and compensation — and why the recovery log and the audit trail are the same artifact.
Governance
Human Oversight Fails When It Is Designed as a Button
What the EU AI Act and the NIST AI RMF actually require, and the latency budget that decides whether an oversight process can function at all.
In progress
Open questions
Evaluation that survives contact with production traces rather than offline benchmarks; mapping governance frameworks onto engineering backlogs; and the containment properties that decide how much autonomy a system can safely be given.