Skip to main content
kellerai.blog

Autonomy is priced, not granted

Why a wider autonomy budget is something you earn from failure-rate data: the ETOPS lesson for AI agents.

KellerAI White Paper · Engineering Discipline & Verification · Jun 2026

Context

AI teams set agent autonomy by feel: let it run an hour, let it touch production. Aviation and banking abolished that instinct: a twin-engine jet earns its extended-range minutes only against a demonstrated, continuously monitored, low engine-failure rate; a risk model earns its limits only by being backtested.

The Finding

To justify a wider autonomy budget you must measure and demonstrate a stable, low rate of undetected failures on the task class over accumulated runs, and price the rare tail event, not the average. Provable reliability is not a tax on autonomy; it is the dividend that buys the direct route.

Tags:
AI agent autonomy & reliability accountingETOPS in-flight-shutdown-rate and Basel backtestingTail-aware governance and the abstention discipline
Paper Details
CategoryEngineering Discipline & Verification
AudienceEngineering, platform, and risk leaders
MethodCross-discipline analysis · reliability-accounting framing
Length~2,150 · 9 min
Sections5
DateJun 2026
AuthorsKellerAI
Read the full paper
In plain language

The problem on your desk

Most AI teams set an agent's autonomy by feel. It demoed well, so they let it run unattended for an hour, touch a test system but not live ones, propose software changes but not put them into production. Authority gets handed out the way a parent extends a curfew: on remembered good behavior, never written down. That is exactly the instinct two of the most consequence-heavy industries on Earth, aviation and banking, deliberately abolished, because guessing kept killing people and losing money. And the gap is now on the risk officer's desk. The banking supervisors who write the rules for risk models have not yet written one for AI agents. The new 2026 bank model-risk guidance, SR 26-2, keeps traditional quantitative and machine-learning risk models in scope, but it deliberately leaves the newer generative and agentic AI (AI systems that can take actions across tools and workflows) to each firm's existing controls. Buying an agent does not transfer that duty. For the risk or model-risk owner the gap is not philosophical: without a measured failure rate, there is no evidence for which human approval gates were removed, which were kept, and why. The buyer earns the autonomy, or no one does.

What the solution is

Reliability accounting: treat an agent's autonomy as a budget that is priced in measured failure-rate data, not asserted in a launch announcement. The unit is the rate of undetected failure (errors that reach a decision-maker without any warning). It is the same discipline aviation and banking already run, applied to AI agents, and it collapses into five deployable steps: - Define the task and what counts as an 'undetected failure' precisely, so the rate that follows means something. - Measure that failure rate over many real runs (in banking this is called backtesting: checking a model's promises against what actually happened; an autonomy claim with no ongoing measurement is, in banking terms, an unvalidated model in production). - Scale the autonomy granted to a demonstrated, stable, low rate, with a tighter bar for higher-stakes actions (drafting text earns a looser bar than pushing code straight to production). - Price the tail, not just the average: the tail is the rare bad outcomes that a good average hides, so govern by the worst credible case rather than the typical day. - Abstain at the point of no return: refuse to take an irreversible action when it rests on a predicted rather than an observed condition, or when no fallback is still reachable.

Why it works

Three fields that never shared a committee converged on this exact accounting, which is why it is structure and not fashion. Aviation prices how many single-engine diversion minutes a twin-engine jet may be from a runway in its in-flight engine-shutdown rate, tightening from about 0.05 to 0.02 to 0.01 shutdowns per thousand engine-hours as the permitted minutes grow, tracked continuously for years. Banking prices a risk model's authority by checking it against reality: it counts how often real losses blew past the loss limit the model promised, and runs a green/yellow/red regime that pulls authority back when failures pile up. It later went further, pricing the average loss on the days that breach the limit, not just the limit itself, so the rare disaster is counted. The third field is AI's own research on answer-or-abstain methods, which let a system measure when it should act and when it should hold back. The deployable AI version also needs independent checking, what banking calls effective challenge (independent review with authority to push back) and aviation builds into independent monitoring: the party that checks the model must be separate from the party that built it, so a model never marks its own homework. That moves AI autonomy from 'scary new thing' to a discipline regulators already recognize in adjacent model-risk work.

The bottom line

Provable reliability is not a tax on autonomy; it is what buys it. When Air New Zealand flew the first scheduled service at the longest over-ocean diversion limit then operating on the Boeing 777, that earned budget let it fly the direct line across the empty Southern Ocean instead of a fuel-wasting detour: the reliability it demonstrated is what let it bank the savings. The same dividend is on offer in AI: a measurably reliable agent earns fewer human checkpoints, more tasks finished end-to-end, and lower supervision cost. The natural champion is the risk or model-risk owner, because reliability accounting is the record they use to defend an autonomy decision to an examiner. The machinery already exists: an independent monitor can measure and bound an agent's undetected-failure rate on a defined task, against accumulated real-run data, and reset the measurement when operating conditions change, while the agent refuses to act rather than guess when that check is uncertain. The in-depth companion lays out the full deployable discipline. It is ready to put to work now.

Section 01

The Reliability Ledger

Ask most AI teams how much autonomy an agent should have, and the answer is a vibe. It feels reliable. It passed the demo. We’ll let it run for an hour and watch. Authority gets handed out the way a parent extends a curfew — on accumulated good behavior, loosely remembered, never written down. That is exactly the instinct two of the most consequence-heavy industries on Earth abolished, deliberately, because it kept killing people and losing money.

A twin-engine airliner is not allowed to fly far from a runway because a regulator likes the look of it. It is allowed to fly far from a runway because the airplane-and-engine combination has demonstrated a specific, low engine-failure rate across hundreds of thousands of fleet engine-hours. The permission is a line item. It is denominated in a number — the rate at which engines quit in flight — and the number is tracked continuously, for years, before the airline gets to fly the longer route. Autonomy, in aviation, is literally an entry on a reliability ledger.

Banking does the identical thing with a different vocabulary. A risk model does not earn the right to set a trading limit because its builders trust it; it earns that right by being backtested — its predicted loss bound measured against what actually happened, day after day, with authority contracting automatically when the failures pile up. Both fields converged, independently, on the same discipline: measure the rate of failure, demonstrate it is stable and low over real experience, and only then widen the budget. The move AI has not yet made is to recognize that an agent’s autonomy is the same kind of quantity — priced in failure-rate data, not asserted in a launch announcement.

Autonomy is not granted. It is priced. The unit is the measured rate of undetected failure on a task class, demonstrated stable and low over accumulated operating experience — and the price includes the tail.

The load-bearing reframe
Section 02

How Aviation Prices Autonomy: The IFSD Rate

For decades, a twin-engine airliner could not legally fly more than sixty minutes’ flying time from the nearest adequate airport. The reasoning was crude but defensible: with only two engines, lose one and you are flying on the remaining one, and the further you are from a runway the longer you are exposed. The sixty-minute rule kept twins hugging the coastlines while four-engine jets took the direct ocean routes.

Then came ETOPS — Extended-range Twin-engine Operational Performance Standards — and the door to longer routes opened in 1985. But it did not open on confidence. It opened on data. The governing quantity is the in-flight shutdown rate: how often, per thousand engine-hours of fleet operation, an engine has to be shut down in flight. A longer diversion tier — the maximum single-engine flying time a twin may legally be from a runway — is granted only when the airplane-engine combination demonstrates an in-flight-shutdown rate at or below a target, and the target gets tighter as the minutes grow: on the order of 0.05 shutdowns per thousand engine-hours for the shorter tiers, tightening to 0.02 at 180 minutes and 0.01 for the longest-range ones. Higher authority demands a lower demonstrated failure rate. That is consequence-scaling, written into the rule.

And the watching never stops. The reporting requirement keeps the world fleet under monitored surveillance — tracking continues beyond 250,000 engine-hours of fleet operating experience until a stable shutdown rate is shown. A tier earned is not a tier owned forever; drift the rate upward and the authority is in question. The lesson for AI is almost embarrassingly direct. The autonomy budget is a function of a measured, stable, low failure rate over accumulated experience — not of a demo that went well once, and not of a rate you stopped measuring the day after you shipped.

Extended range is not a reward for good engineering. It is what a demonstrated, continuously monitored, low failure rate buys. Stop measuring the rate and you have stopped earning the minutes.

The ETOPS lesson
Section 03

How Banking Prices the Same Thing: Backtested Exceptions

Banking reached the identical accounting from the other side of the world, under its own regulatory pressure. A bank that runs a risk model stating, say, a 99% one-day Value-at-Risk is making a precise probabilistic claim: on a normal day, losses should exceed this number only about one time in a hundred. Regulators do not take that claim on faith. They backtest it — they count the days the actual loss blew through the stated bound, expect roughly two and a half such exceptions a year out of 250 trading days, and run a traffic-light regime around it. Stay in the green zone and the model keeps its authority. Drift into yellow or red — too many exceptions over the rolling window — and the bank is forced to recalibrate and to hold more capital against the model it can no longer fully trust.

The structure is identical to ETOPS, term for term. A stated tolerance scaled to consequence. An empirical failure rate measured against it over a rolling body of real experience. Authority that widens only on a demonstrated low rate and contracts automatically on drift. A bound, in banking, is not a claim. It is a measured, regularly verified commitment, and the verification is performed by someone structurally independent of the people who built the model — effective challenge, in the supervisory language of SR 11-7 and its 2026 successor, SR 26-2.

State the AI translation plainly. An AI vendor who asserts a low error rate without an ongoing regime to measure that rate against reality is, in banking terms, running an unvalidated model in production. That is not a selling point; it would likely draw an examiner finding. Autonomy you cannot backtest is autonomy you have not earned — you have merely asserted it, and an assertion is precisely what both aviation and banking spent decades learning never to accept.

Section 04

The Dividend: Reliability Pays

It is tempting to read all of this as overhead — reliability as a tax that careful people pay and bold people skip. Aviation tells the opposite story. Provable reliability is not the cost; it is what unlocks the cheaper, faster, better operation. The reliability is the thing that pays.

On 1 December 2015, Air New Zealand became the first airline ever to fly a scheduled ETOPS-330 service — Auckland to Buenos Aires, a Boeing 777-200ER on Rolls-Royce Trent 800 engines, the longest extended-range authority then operating on the 777. The airline did not leap there. It had flown 240-minute ETOPS routes on that airframe-engine combination from October 2014, accumulating roughly a year of operating experience before receiving the 330-minute approval in November 2015. And the payoff for that earned budget was geometric: a 330-minute diversion allowance let the twin fly the direct great-circle line across the empty Southern Ocean instead of bending the route into a fuel-wasting dogleg to stay within reach of a runway it would almost certainly never need.

That straight line is money. A modern twin burns substantially less fuel than a four-engine jet over the same sector — the gap runs to tens of percent per seat — and because engine-related costs are a major share of maintenance, carrying two engines instead of four compounds the saving across fuel, spares, and overhaul. Every diversion-minute earned through reliability data converts directly into fuel saved, time saved, and carbon not burned. The reliability the airline demonstrated is what let it bank the efficiency.

The AI parallel is exact. A provably reliable agent earns a wider autonomy budget, and a wider budget is the efficiency dividend: fewer human checkpoints in the loop, more tasks completed end-to-end without a hand-back, lower supervision cost per unit of work. The teams that measure their agents’ failure rates and demonstrate them stable and low are not slowing themselves down. They are the only ones who get to fly the direct route. Reliability accounting is how you bank the savings.

Provable reliability is not a tax on autonomy. It is the thing that buys it. The direct route — fewer checkpoints, faster completion — is the dividend the failure-rate data pays out.

The dividend
Section 05

Account for the Tail

There is a way to do reliability accounting badly, and it is the most natural way: price the average and ignore the tail. A low mean failure rate can sit comfortably on top of a rare, catastrophic failure mode that the average quietly buries. Aviation has a hard case that makes the point, and it has to be told accurately.

On 7 October 2013, a Royal New Zealand Air Force No. 40 Squadron Boeing 757 — not a civilian Air New Zealand aircraft — flew callsign NZ7571 from Christchurch toward Pegasus Field on the Ross Ice Shelf, Antarctica, with 130 people aboard. The 757 could not return to Christchurch without refuelling at Pegasus, so a point of safe return was computed before departure: a hard line past which the designed fallback — turning around and going home — was no longer reachable. Forecasters assured the crew the weather at Pegasus would improve and cleared the flight past that point. Roughly twenty minutes later, observations showed a fog bank had enveloped the runway and its approaches. The crew flew three approaches; on the third, at about 110 feet, they acquired the approach lighting and landed below published minima in near-whiteout. There was no damage and no injuries. It was a successful recovery — and that is the point, not a tragedy to dramatize. (It should not be confused with the 1979 Mount Erebus disaster, an unrelated Air New Zealand DC-10 navigation accident that killed 257.)

The inquiry found the crew’s decisions appropriate. What it faulted was the original risk assessment: it had gaps — no alternative approach procedures or aerodromes suitable for a 757, and an under-weighting of how quickly early-season Antarctic weather can turn. In other words, the plan committed past an irreversible point on a forecast that then diverged from reality, with the designed fallback already foreclosed by fuel and range, and an inadequate set of alternates behind it. The average mission to Antarctica is uneventful. The tail mission is the one that defines the risk.

Map that onto autonomous agents and it stops being a flying story. The NZ7571 failure mode, rendered in software, is an agent that commits past a rollback horizon on predicted rather than observed conditions, with no reachable safe-harbour and an under-specified fallback set — a long-running tool-use chain that crosses an irreversible operation because the upstream signal it was told to expect probably holds, and finds out twenty steps later that it did not. A reliability ledger that counts only the average task will price that mode at roughly zero, right up until it dominates the entire loss. Banking made exactly this correction when it moved from Value-at-Risk, which marks the loss quantile but says nothing about how bad losses get beyond it, to Expected Shortfall, which averages the tail beyond that quantile. The discipline is the same in software: govern the tail, and abstain — refuse to cross the irreversible point — when the forecast the decision rests on cannot be verified at the moment of commitment.

That is reliability you can bank: an autonomy budget earned against a measured failure rate, monitored continuously, scaled to consequence, and priced at the tail rather than the mean. The in-depth companion develops the full account — the ETOPS in-flight-shutdown ledger, Basel backtesting and the move to Expected Shortfall, the selective-prediction machinery that lets an agent measure and bound its own undetected-failure rate, and the deployable discipline that ties them together. Read it at Priced in Failure-Rate Data .

A low average failure rate is not a safe one if it hides a rare, irreversible failure mode. Price the tail, not the mean — and abstain when you would commit past the point of safe return on a forecast you cannot verify.

The tail rule