A new frontier model reportedly landed every 37.5 days23, on median, in 2023. In 2026 the median gap is reportedly 11 days24. That is a 3.4x speedup in three years, according to ARK Investment data built on Artificial Analysis figures through April 28, 202612.
Here is the strange part. Three of the four biggest labs, Anthropic, OpenAI, and xAI, publicly backed a proposed slowdown on September 1218 and 13. Nine days earlier, four frontier models shipped in about 72 hours. The same companies asked for a brake while pressing the gas.
One damaging admission before we start. The 1124-day figure is a median from a small sample of U.S. labs, and it wobbles. The 2025 number was 17 days, up from 13.5 in 2024, because the tracked sample grew. So I will not pretend the number is a law of physics.
But the direction is not ambiguous. Every year since ChatGPT, major releases across 1124 labs have risen: 22 in 2023, 58 in 2024, 92 in 2025, and 76 by September 1318 this year, per AI Release Tracker. If you are building on these models, you need a plan for that tempo. This article gives you one.
Brake, Bell, Blueprint
Every promise to slow down falls into one of three buckets. Sort the pledge before you react to it.
What the release data says about the slowdown that was promised.
A Brake changes behavior. It is binding, it has a trigger, and you can watch it fire. The June 2, 202602 executive order created a voluntary pre-release review of up to 30 days19 led by the NSA. That is a partial Brake at best, because "voluntary" means the lab holds the pedal.
A Bell announces intent. It makes noise, and noise has some value, but nothing stops. OpenAI matching Anthropic's evaluator pledge with "employee-like access" and no implementation details is a Bell. Meta reportedly issuing no position at all is silence, which is the absence of even a Bell.
A Blueprint builds the capacity to slow down later. The Pacing the Frontier statement, signed by 1,386 frontier-lab employees in July 2026, is a Blueprint. Read the text: it asks the U.S. government to support an international effort to "develop the technical and governance tools needed to deliberately pace the frontier." One published critic called it "a request for capacity to act, not itself a binding slowdown. Only a Brake changes your roadmap. Bells and Blueprints change your reading list. As of September 2026, the frontier has one soft Brake, several Bells, and one Blueprint. Design accordingly.
Eleven Days Is a Calendar, Not a Conscience
Strip away the safety language and look at what a release actually is. It is a product cycle. Product cycles answer to competitors, compute schedules, and training runs that finish when they finish.
Intel ran on this logic for roughly a decade with its tick-tock model. One year a new process node, the next year a new architecture, every year a launch. Nobody called that a safety posture. It was a factory rhythm, and the rhythm set the strategy, not the other way around.
The AI labs now run the same kind of clock, only faster. Anthropic's Michael Gerstenhaber described Claude gaps compressing from six months to two months. One treadmill analysis put the Opus interval at 26 weeks between Opus 4 and 4.5, then about 10 weeks after. Google reportedly halved its Pro interval from 34 weeks to 1318.
The clusters give the game away. Between February 5 and April 23, 2026, Anthropic, OpenAI, and Google shipped seven frontier models in 78 days. In September, Anthropic shipped on the 1st, Google and Meta on the 2nd, OpenAI on the 3rd. Nobody schedules four launches inside 72 hours by accident. They ship to close a competitor's window before it opens.
Now do the arithmetic the pledge asks you to skip. A 3019-day review clock against an 1124-day release median means roughly three models sit in review at once. Divide 30 by 11 and you get 2.7. The review does not slow the frontier. It runs in parallel with it, like a security line that never shortens the flight schedule.
Put the two side by side. A Brake costs you a launch. A pledge costs you a press release. When the field endorsed a slowdown within 48 hours on September 1218 and 13, the labs picked the cheaper signal, and the calendar kept moving.
Anthropic's move is the exception worth studying. Permanent third-party evaluator access is a structural change, not a statement. It binds future versions of the company, not just the current mood. I think it is the only item in the September news that qualifies as a partial Brake, and even it does nothing to lengthen the gap between releases.
There is a fair counterargument. Cadence is not risk. Five small hops spaced six days apart can be safer than one big jump after 30 days19, and labs do run incremental evals on each drop. Fair. But that argument concedes my point: engineering and competition set the interval, and safety fits itself inside whatever interval results.
It is unclear whether OpenAI ruling out a 2026 IPO changes any of this. Public markets push toward quarterly drumbeats, so staying private could loosen the clock. Or it could mean nothing, because the competitors setting the tempo are private too. I have not seen convincing evidence that ownership structure has ever slowed a shipping race.
Factory Rhythm Beats Pledge Language
A 3019-day review does not slow an 1124-day clock.
Divide the voluntary 30-day pre-release window from the June 2, 202602 executive order by the 11-day release median and roughly three models sit in review at once. The review runs alongside the frontier rather than in front of it, like a security line that never shortens the flight schedule.
Betting on the slowdown is the expensive mistake.
Design for churn and a slowdown arrives anyway, and you have overpaid for an abstraction layer. Hard-wire one vendor's prompt quirks and every 11 days a competitor gains a capability you cannot reach without a rewrite.
The median moved backward before it collapsed.
The 2025 figure was 17 days, up from 13.5 in 2024, because the tracked sample of U.S. labs grew. Treat 11 days as a direction, not a law of physics, while major releases across 11 labs climbed from 22 in 2023 to 92 in 2025 and 76 by September 1318 this year.
2031: Impermanence as Infrastructure
Pull back five years. The question is not whether the gap hits 8 days or drifts back to 17. The question is which mistake you can afford to make.
Design for churn and the slowdown arrives anyway: you paid for an abstraction layer you used less often than planned. Mild cost. Bet on the slowdown and churn continues: you hard-wired one vendor's prompt quirks into your product, and every 11 days24 your competitor gets a capability you cannot reach without a rewrite. Severe cost. That is asymmetric risk, and the asymmetry only points one way.
Nvidia offers the template for reading a tempo change. When the company moved its datacenter GPU architecture cadence from roughly every two years to every year, the buyers who won were not the ones who bought the best chip. They were the ones who built software that ran on whichever chip came next. CUDA outlived every individual card.
The Buddhist idea of impermanence is usually offered as comfort. Here it is an engineering spec. A model is a rental, not a purchase. The Fable 5 that opened Anthropic's 5 family on June 9 was followed by Sonnet 5 on June 3019. GPT-5.6 arrived July 9 as three named variants. None of these is your foundation. The layer you own above them is.
Beginner's mind, shoshin, turns into an edge in this world. The team that assumes each new release might beat its favorite model tests it. The team that assumes its favorite is still best skips the test and finds out from customers. The 70 percent rule applies: when a new model looks 70 percent likely to beat your current one on your evals, run the swap trial. Wait for certainty and the next release will make the question moot before you answer it.
My read is that by 2031 the frontier median stops mattering as a headline, because the winners will have stopped tracking it. They will track one number instead: hours from a new model's release to a go or no-go decision on their own eval suite. Salary buys furniture. Model-agnostic infrastructure buys your future.
Ship a Model-Swap Rehearsal by Friday
Here is what to build this week. None of it needs a CS degree. All of it needs the humility to assume your current model will not be your model in a month.
First, write the acceptance test before you touch a new model. An acceptance test is a fixed set of inputs with the outputs you require, scored automatically. Steal the habit even if you skip the tool.
Second, put a routing layer between your code and the vendor. A router is a thin piece of software where you name the model once, in one file, instead of in fifty. Swapping GPT-5.6 for Claude Sonnet 5 should be a one-line change plus a test run, not a sprint. If you work in C or C++, where these swaps get hairy, ByteAsk is a coding agent that compiles, debugs, and tests until the change actually holds.
Third, run one rehearsal swap on purpose, this Friday. Pick a model you are not using. Route 10 percent of traffic to it, run the acceptance test, and log what broke. Things will break. That is the point of a rehearsal, and finding it on Friday beats finding it during a real release on Tuesday.
Fourth, keep a swap ledger. Google Sheets canvas can turn a plain sheet into an interactive mini-app from one prompt, so a two-column log of release date and your decision becomes a dashboard in minutes. Track the days between a release and your verdict. Get that number under 1124 and the frontier's tempo stops being a threat.
Fifth, automate the watching. Gemini Spark can run web errands in Chrome, so set it to check the release trackers you trust and drop new entries into your ledger. You do not need to read every system card the day it lands. You need to know a card exists, and you need your test to run.
Get your reps in. The labs have shown you their calendar: 37.5 days23, then 13.5, then 17, then 1124. They may slow down someday, and if they do, your abstraction layer costs you a little overhead and nothing else. Until a real Brake appears, build like the 11-day clock is the only schedule anyone is honoring, because so far it is.
Rehearse a model swap before the next release forces one.
- Write the acceptance test first. Fix a set of inputs with the outputs you require and score them automatically, before you touch any new model. The test is the asset, not the model it happens to be pointed at.
- Name the model in exactly one file. Put a thin routing layer between your code and the vendor so swapping GPT-5.6 for Claude Sonnet 5 is a one-line change plus a test run, not a sprint.
- Run a rehearsal swap on Friday. Pick a model you are not using, route 10 percent of traffic to it, run the acceptance test, and log what broke. Then keep a two-column ledger of release date and your verdict, and drive the days between them under 1124.
The labs already published their calendar.
Sort every promise before you react to it: a Brake changes behavior, a Bell announces intent, a Blueprint builds the capacity to slow down later. As of September 2026 the frontier has one soft Brake, several Bells and one Blueprint signed by 1,386 employees, and the median gap still sits at 11 days24 against 37.523 in 2023. Anthropic's permanent third-party evaluator access is the one structural change in the news, and even it does nothing to lengthen the interval. So stop tracking the median as a headline and start tracking the only number you control: hours from a release to a go or no-go call on your own eval suite.
