Providers and failover

What AutoDev does when a provider reports a usage limit or cannot be reached. It routes the turn elsewhere instead of sleeping, and it sleeps only when nowhere is left.

How a limit is detected #

AutoDev runs three provider CLIs for now: claude, codex and antigravity, whose binary is agy. It does not predict a limit. It learns of one when a call reports it. Three things count:

A notice is short, and text a model wrote, when it is longer than 300 characters, is never scanned, so a report that only quotes a limit notice is not mistaken for one. The reset time is read out of the notice. When none can be read, claude waits an hour, and codex and agy probe again every 30 minutes.

A limit belongs to an engine #

A limit belongs to the engine that reported it, not to the participant that hit it. Every participant running on the same engine shares it, so a limited turn is never handed to a teammate on the same exhausted account, which would hit the same limit seconds later. An engine is a provider and, when a failover entry names one, a model of it: claude and claude:<model> are two engines.

First, the participant's own chain #

A participant can carry a failover chain in its file: provider[:model] entries, in order.

.autodev/participants/<id>.md
---
provider: claude
failover: [codex, claude]
---

When a participant's turn reports a limit and nothing was posted, AutoDev walks that participant's chain first. It moves to the next entry that is not known to be limited, and runs the turn there in the same step. If that entry is limited as well, it advances again. It writes <id>: <from> limited → <to>. An entry naming a CLI this machine lacks is skipped. On a hop the participant's pin does not apply, since it names a model of the primary provider, unless the entry names a model itself.

The chain comes first because a teammate often has nothing of its own to do: when the thread is waiting for a review or an answer from this participant, the others could only say they are waiting, and each of those is a paid turn.

There is no default chain: a participant without failover has only its teammates. The field lives in the participant's file. Edit prompt opens it in the panel, and autodev team edit <id> prints its path. Participants and lanes covers the file.

Then, a teammate #

When the chain is spent or empty, the turn goes to another participant whose engine is free. AutoDev writes <id> is rate limited — @<other> takes the turn to the log. The teammate is chosen by the same rules as any turn.

Two cases do not route to a teammate. The holder of a pending gate never hands its turn away, and a team of one has nobody to hand it to. Both use only their own chain. If the limit lands after the participant has already posted, the post stands: the limit is recorded, and the next turn routes around it.

When nobody can take the turn #

When every participant and every chain is limited, the run does not stop. The status becomes rate_limited, and the daemon sleeps until the soonest reset across all of them, not the reset of whichever failed last. AutoDev writes every participant limited — waiting for <provider>, back first at <time>.

The sleep runs in slices. Each slice refreshes the heartbeat, and a stop or a pause interrupts it. A message from you does not, since nothing can run before the reset. A machine that was suspended during the wait wakes at the reset, not after the sleep it had left. The time asleep is not charged to maxRuntimeMin.

A limit wait has no cap of its own. maxWaitMin and the doubling recheck belong to a wait move, which The shared thread describes.

Coming back #

At the start of each turn, AutoDev lifts the limits whose reset has passed. A participant that had moved along its chain goes back to the earliest entry that is free, which is not always the primary. AutoDev writes <id>: limit lifted — back on its primary provider.

The choice of provider is made when a turn starts and never during one. The record of who is limited is kept in memory for the life of the daemon. A restart forgets it and finds out again by probing.

When a call fails instead #

A usage limit is something a provider reports, with a time to wait for. A call can also fail without reporting one, and AutoDev tells two of those apart.

The network is down. When a call fails with a network error, such as an address that cannot be found or a connection that is refused, reset or timed out, and nothing was posted, the provider is out for every participant on it, as a limit puts an engine out. AutoDev tries it again after 1, 2, 4 and 8 minutes, then every 15 for as long as it lasts. It writes <provider> could not be reached; trying it again at <time>, and <provider> is reachable again when it is. Participants on another provider carry on. With every provider out, it writes No provider can be reached; trying again at <time>. This has no deadline: a night on which the internet is down for three hours picks up when it comes back.

The call fails some other way. The CLI crashes, times out, or ends its turn without a post. AutoDev waits longer after each failure in a row: a second after the first, then 1, 2, 4, 8, 15 and 30 minutes, then an hour each time. After 12 failures in a row, about five hours of waiting, the run pauses and writes Gave up after 12 failed attempts in a row, over five hours (<what failed>); run paused — resume to try again. A message from you cuts a wait short, and one call that works starts the count over. If a limited engine comes back sooner, the next attempt is made then. Unlike a limit wait, each attempt counts against maxIterations.

The five hours are there for a limit whose notice AutoDev cannot read yet. It lifts by itself, and a schedule that outlasts a usage window carries the run through it. A participant that ends three turns in a row without a post is a different case: retrying would spend quota to learn nothing, so the run pauses at once and names it.

What this does not do #

Participants and lanes covers the file the chain lives in, and Every config key lists maxIterations, maxRuntimeMin and maxWaitMin.