Google went seven months without a new top-tier AI model. Its last one, Gemini 3.1 Pro Preview, shipped on February 19, according to Tech-ish. On September 30 it finally launched Gemini 4 Argon. Then it told almost everyone they could not use it.
Google's website says that Argon went first to a small group of cybersecurity partners through Google's Fairwind Program. Moneycontrol says the reason is hacking risk. Google has set no public release date. Meanwhile Ars Technica reports that Argon already runs inside Google and is projected to save over 300 14 TiB of data-center memory once rolled out.
The benchmark story is also thinner than the headlines. Argon tops 13 01 of 19 tests in Google's own table. It loses five, and two of those losses are by 9 20 points or more. I have not used Argon, and neither have you, so every number here comes from Google or from reporters reading Google's tables.
By the end of this piece you will have a plan for building around a model you cannot touch yet.
The Four-Rung Release Ladder
Every frontier model now has two dates. One is the day it exists. The other is the day you can call it from your code. I call the space between them the Four-Rung Release Ladder, and Argon shows every rung.
What Argon would cost you, if you could call it today
Rung one is the lab itself. Anadolu Agency reports that thousands of Google employees already use Argon for specialized coding and research work. Ars Technica adds that Argon agents moved more than 800,000 lines of the Fuchsia Zircon kernel from C/C++ to Rust.
Rung two is the vetted outsider. Fairwind started on September 2, 2026 24, and reportedly counts more than 650 partners, including CrowdStrike and Palo Alto Networks. Only some of them get Argon, and those who do reportedly get it with cyber guardrails relaxed so they can find and patch real vulnerabilities.
Rung three is the paying customer. Launch coverage says paid API users and Google AI Ultra subscribers should get access before the general public. Rung four is everybody else, which includes most startups. Google told Anadolu it will open access "as soon as possible," a promise with no date attached.
The ladder comes down to one rule. Your product lives on the rung you can reach today. Plan for the bottom rung and treat anything above it as upside.
Reading Argon's 13 Wins Honestly
A win count hides the size of each win. On Google's table, a 1-point lead and a 10-point loss each count as one result. The margins are what matter, so read them line by line.
Argon's clearest wins are in knowledge work. On the Harvey Legal Agent Benchmark it scored 19.6% 17, against 6.7% for the next model. On AutomationBench it posted 51.3% 18, while the closest rival hit 42.5%. One report says its DeepSWE v1.1 score of 77.9% 04 beats GPT-6 Astra and both Claude models.
Several wins are close calls. That is a 1.9 25-point lead. On Vibe Code Bench, its 91.9% 05 trailed Claude Sonnet 5.5's leading 92.39% by about 0.5 points.
The losses come on hard engineering. On FrontierSWE v2, GPT-6 Astra scored 65.5% 06 to Argon's reported 55.0%, a 10.5-point gap. On Terminal-Bench 4.0, Claude Opus 5.5 scored 66.4% 20 to Argon's 57.4%, a 9.0-point gap. Both tests lean on long, tool-heavy engineering work, which looks a lot like what most coding agents do in production.
The strangest number sits right next to the gate. Google restricted Argon because of its cyber power. Yet on CWE-bench v1, the vulnerability benchmark, Argon scored 68.0% 21 and reportedly tied for first. The model judged too risky for a broad release does not own its cyber benchmark outright.
Google's evaluation page has a caveat worth reading. It says scores for rival models come from those providers' own reports. Argon's DeepSWE and Terminal-Bench scores are self-computed, with a mini-swe agent harness used for DeepSWE. Different harnesses can move scores, so a 1.6-point gap deserves far less weight than a 10.5 06-point one.
My read is that the frontier has become a portfolio. Astra and Opus each lead one hard engineering test on this table. Argon leads knowledge work and legal agents. No single model is best at every workload a builder cares about.
Access is turning into an asset in its own right. Firms like CrowdStrike and Wiz can qualify for a gated program because they bring monitoring and accountable human operators. A two-person security startup cannot show that today. So the newest cyber capability reaches incumbents first, an asymmetric advantage handed out by the lab.
You will never have full information about a gated model before you have to decide. A useful rule is to act once you have about 70% of what you would want. On Argon you already have that much: the margins and the missing date.
I think the gate is part safety and part strategy. Google gets adversarial feedback from hundreds of trusted partners while it builds demand among paying customers. Whether this becomes standing policy for every Google flagship, or stays a one-off for an unusually dual-use model, is still unclear.
Narrow leads, odd gates, uneven access
Argon loses where coding agents live.
GPT-6 Astra beat Argon 65.5% to 55.0% on FrontierSWE v2, and Claude Opus 5.5 led Terminal-Bench 4.0 by 9.0 20 points. Both tests lean on long, tool-heavy engineering work that resembles production coding agents.
The gated model only ties on cyber.
Google restricted Argon over hacking risk, yet on CWE-bench v1 it scored 68.0% and reportedly tied for first. Whether gating becomes standing policy or stays a one-off for a dual-use model is still unclear.
Incumbents climb the ladder first.
Fairwind reportedly counts more than 650 partners, including CrowdStrike and Palo Alto Networks, who bring monitoring and accountable operators. A two-person security startup cannot show that today, so new capability reaches incumbents first.
2031: Access Becomes a Product Spec
The best precedent is on Google's own calendar. In May 2026, Google said Gemini 3.5 Pro would roll out the following month. Business Insider reports it was postponed several times internally because of poor performance, and it never shipped. Any team that pinned a summer roadmap to it lost the summer.
Around that gap, things moved fast. Tech-ish dates Claude Fable 5.1 to September 1 and GPT-6 Astra to two days later. Anthropic shipped Claude Opus 5.5 on September 22, and Argon followed on September 30. That is four frontier launches from three labs inside 30 days 26.
Assume that burst was unusual and the real pace is a quarter of it. That still means one frontier launch a month, or about 60 27 between now and the end of 2031. If half of them arrive gated, builders will read about roughly 30 28 models before they can call them. The public frontier might always trail the real one by a rung or two.
Capability is compounding at the same time. Anadolu reports Argon's output limit jumps from 64,000 tokens 15 to 1 million, about 15.6 29 times in one generation. Bigger single outputs mean longer autonomous runs, and that is the kind of power labs will want to stage.
The gates keep multiplying. Axios reports Google is taking part in the US government's voluntary pre-release process, a second checkpoint beyond Google's own review. Fairwind itself was set up just 28 days 30 before Argon launched. By 2031 I expect investors to ask a startup which access programs it belongs to, right after they ask about its burn rate.
The math favors cheap optionality. Suppose a model router and a test suite cost one engineer two weeks. A roadmap pinned to a model that slips one quarter stalls for about 13 weeks 31. Two weeks of insurance against 13 weeks of exposure pays off every time a launch slips.
Wire a Model Router Before Argon Ships
You do not need a CS degree for this. You need a spreadsheet, a few API keys and an afternoon. Take a deep breath and take it step by step.
First, write down 20 real tasks from your product. Pull them from support tickets and past bugs. Write the pass condition next to each one in plain words. This is your eval suite, which just means a fixed test you rerun on every model.
Then put every model call behind one function. A router is that single function plus a setting that names which model to use. Switching models should mean editing one line of config.
Next, run the 20 tasks on each model you can reach today. Log the pass rate and the cost per completed task, which includes the retries. If your work looks like FrontierSWE or Terminal-Bench, include GPT-6 Astra and Claude Opus 5.5, since they led those tests on Google's table.
Argon's slot in your router should carry a price written in pencil. Reported launch pricing is about $2 23 per million input tokens and $10 per million output tokens, and none of it is contractual yet. A task with 200,000 32 input tokens and 50,000 output tokens would cost about $0.40 plus $0.50, or $0.90. One full 1-million-token output would cost about $10 33 in output alone.
Last, add an entry for Argon in your config and leave it switched off. When access reaches your rung, run the same 20 tasks and compare. If it wins your workload by a wide margin, flip it on. If it wins by 1.9 25 points, the size of its Vals Index lead, weigh price and latency hard before you migrate.
Expect the first run to break. A prompt that works on one model often fails on another, and that failure is the data you wanted. Share your pass rates with your team, or learn in public and post them. Then rerun the suite every time a new model lands, because at one launch a month, the next one is already on its way.
Wire a model router before Argon ships
- Write a 20-task eval suite. Pull 20 real tasks from support tickets and past bugs, and write a plain-words pass condition next to each one.
- Put every model call behind one function. Make switching models a one-line config edit, then log pass rate and cost per completed task, retries included, for GPT-6 Astra, Claude Opus 5.5 and anything else you can reach.
- Add Argon to your config, switched off. Price it in pencil at $2 23 input and $10 output per million tokens. When access reaches your rung, rerun the 20 tasks and flip it on only if it wins your workload by a wide margin.
Build on the rung you can reach, and keep a slot open for the one you cannot.
Argon shows that a frontier model now has two dates, the day it exists and the day you can call it. Its 13 01 wins hide narrow margins, and its losses fall on the engineering work most agents actually do. Two weeks spent on a router and an eval suite protect you against the 13 weeks 31 a slipped launch can cost. At one frontier launch a month, the next gated model is already on its way.
