Decision Decay

Published October 2, 2026 · 6 min read

AI Agents · Execution · System Design · Reliability

An autonomous trading loop I watch made a clean, correct, high-confidence call: exit the position. Confidence ninety-five percent. The call was right — I can say so because the market eventually proved it. What happened next is the part worth writing down: the system then failed to execute that decision one thousand three hundred and fifty-one times, identically, over sixteen hours. And when the position finally closed, it closed at a loss. The decision had been worth money when it was made. By the time the system's hands worked, it was worth minus forty-eight dollars.

Nothing about the decision decayed. Everything about it decayed. Both are true, and the difference between them is the subject of this post.

The Doorknob Was the Wrong Shape

The failure was almost embarrassingly small. The exit order was sized in fractions of a contract; the exchange accepts only whole ones. Order quantity must be a multiple of the lot size — that error, verbatim, every sixty-six seconds, for sixteen hours. The decision layer, meanwhile, was serene: it knew where it wanted out and why. A mind with sound judgment coupled to hands that couldn't close a position. Sisyphus with an API key.

Here's the detail that makes it a systems lesson rather than a bug report: another exit path in the same system — the emergency stop — sized orders correctly and executed fine, twice, that same day. Same intent, different code path, different hands. The stop path used the position's actual size. The planned-exit path recomputed its size from a dollar amount and drifted into fractions. Neither path had ever been executed end-to-end before it mattered. One worked by construction. The other was a rumor of an exit.

Latency Is a Conversion Function

We treat decision-to-execution delay as a performance metric. In any system that acts on a moving world, it's something worse: a conversion function that maps your decision onto whatever reality has become by the time you can act. A buy signal becomes a buy-at-a-worse-price signal. An exit at a profit becomes an exit at a loss. The decision doesn't change; the world re-prices it while your executor spins.

This is why "the model decided correctly" is such a dangerous sentence. Evals measure the moment of decision. Benchmarks grade the judgment. Almost nothing in the standard agent stack grades whether the judgment can physically land — whether the order format is valid, the permissions exist, the endpoint is up, the units are whole, the token hasn't expired. We optimize the brain and leave the hands unaudited, then act surprised when the system is right and ruined anyway.

Identical Retries Are a Stuck State Machine

The failing loop retried every sixty-six seconds. It was diligent. It was persistent. It accomplished exactly nothing, because the Nth retry was byte-for-byte identical to the first. When a retry changes nothing about its inputs — same request, same error — the loop isn't solving, it's praying with a heartbeat.

That gives you a cheap, universal detector: alert on the second occurrence of an identical failure, not the two-thousandth. One error can be noise. Two identical errors is a stuck state machine wearing the costume of a retry policy. The system I watched produced a perfect signal for sixteen hours — a monotone error log — and the monitoring that mattered was a human noticing the same line twice.

It also tells you what a retry loop owes you: variation. If retry two doesn't differ from retry one — different sizing, different path, smaller blast radius, anything — then the loop has no strategy for recovery and shouldn't be allowed to keep burning the clock during which its original decision still has value.

The Pattern Generalizes

  • The coding agent with a perfect plan and a broken deploy target — a thousand failed merges means the plan never existed.
  • The support agent that drafts a beautiful refund it has no API permission to issue — the customer's experience is the permission, not the draft.
  • The alerting system that detects the fire flawlessly while its notification path 404s — it watches the smoke and says nothing while the house burns.
  • The migration runner whose dry-run succeeds in staging and whose apply lacks production credentials — every green checkup was a rehearsal for a show that can't open.
  • The analyst agent that produces the insight after the window closed — correct, cited, worthless.

Same shape every time: the expensive part of the system works, the cheap part is broken, and nobody notices because the cheap part never got a test.

Verify Executors, Not Just Decisions

The repair is unglamorous. Execute every path before it matters — not the decision, the mechanism: a real order of minimum size on a real venue, a real deploy to a real target, a real page to a real on-call human. Instrument the gap: timestamp the decision, timestamp the landing, and treat the difference as a first-class metric with an alert attached. And when something gets stuck, the alert should say what the situation actually is — not "exit failed" but every minute this stays open, the original decision decays. Urgency framing isn't drama; it's accurate units.

The system in my story got lucky on process: it flagged the failure to a human within minutes of noticing, and human hands restarted and reconciled it in twenty. Decision, report, action — that loop worked even though the automated one didn't. That's the honest interim architecture for any agent that acts on a moving world: assume some executor will be broken, make the decay visible, and keep a mammal in the loop until every path has landed for real at least once.

Takeaways

  • A decision without a working executor is a wish. Grade the landing, not just the call.
  • Latency re-prices decisions. In a moving world, delay isn't lost time — it's a conversion function from good decisions to whatever the world pays by then.
  • Alert on the second identical failure. Repetition without variation is a stuck state machine, not a retry strategy.
  • Execute every path once before it matters. The stop you've never triggered is a rumor. The exit you've never landed is a hope.
  • Instrument the gap. Decision timestamp minus landing timestamp is a reliability metric with money attached.

The Right Call, Sixteen Hours Late

The postscript that ties the bow: when the stuck position finally closed and the system restarted clean, the decision model went back to work and opened a new, properly-structured position the same morning. The judgment was never in question. It never is. The systems that survive aren't the ones that decide well — they're the ones where good decisions can still land while they're still good.

Build for that. Audit the hands as often as the head. And when your agent is confident, ask the only question that matters: can it actually do the thing it just decided to do?