Google deprecated rule-based attribution. Step-level agent credit assignment is about to re-run that experiment.
Latest
Every KellerAI paper, newest first.
20 papers • 8 releases
Ad allocation and LLM eval converged on the same budgeted-bandit problem, a decade apart, without citing each other.
Meta's Robyn makes agreement with experiments a fitting objective. Judge calibration should steal the pattern, and the estimand warning that comes with it.
A regulatory clearance authorizes one use, not a model — and renders every other use unvalidated by construction.
A model reported at 95% accuracy says nothing about which 5% it gets wrong — and one missed cancer is not a thousand false alarms.
Clinical-AI autonomy is not the removal of the clinician; it is the guaranteed-reachable fallback that licenses the autonomy.
Two reviewers who fail together are one reviewer.
A demo quarter is not a backtest. Authority is priced in failure data, not favorable runs.
The vendor ran the eval. You still own the governance.
Aviation stopped asking whether a twin-engine jet could cross an ocean and started asking how far it had earned the right to fly. AI agents need the same envelope.
The ETOPS rule is not "fly farther." It is "never fly past a reachable safe harbour" — and the same rule should govern every autonomous agent.
Why a wider autonomy budget is something you earn from failure-rate data: the ETOPS lesson for AI agents.
The only honest statement of how autonomous a system is, is the statement of where its envelope ends.
SOTIF and the fault-free hazard: a system can execute its specification perfectly and still be lethally wrong.
The minimal-risk maneuver is not where autonomy fails. It is what licenses autonomy at all.
Assurance for an AI agent must attach to what it does, not to the model it runs.
The Federal Reserve published the model-risk inventory schema. A bank can just use it.
The AI field has been asking the wrong question. Aviation and banking each solved the underlying engineering problem decades ago — under regulatory compulsion, at enormous cost, and with a precision the AI industry has yet to borrow.
Why machine-enforced reasoning chains are a durable competitive advantage.
AI systems are starting to improve themselves. Here is why that is dangerous, and what a fix looks like.