K Koda Intelligence
DEEP DIVE DEEP DIVE № 213 · 01 October 2026DOCKODA-20261001-FD2582C7C5E9sha-256 of date + article + 24 checked + 9 computed

Google built its best model, then locked the door

On Sept 30, Google launched Gemini 4 Argon and handed it only to vetted cybersecurity partners, with no public release date. Its benchmark lead is real but narrow. It hits 77.9% 04 on DeepSWE v1.1, yet it loses hard engineering tests by 9 20 points or more. Its output cap jumps from 64,000 15 tokens to 1 million. Builders now have to plan around a flagship they cannot call.

7 MIN READ · BY THE KODA EDITORIAL TEAM · STRATEGY · GATED RELEASES
ARGON LAUNCHSEP 30REPORTED Google
TESTS WON13 OF 19VERIFIED CLAIM 01GOOGLE TABLE
DEEPSWE V1.177.9%REPORTED CLAIM 04Google
ARGON LAUNCHSEP 30Google TESTS WON13 OF 19GOOGLE TABLE DEEPSWE V1.177.9%Google FRONTIERSWE V255.0%↓ 10.5 PTS VS GPT-6 ASTRA TERMINAL-BENCH 4.057.4%↓ 9.0 PTS VS OPUS 5.5 OUTPUT CAP1M TOKENS↑ 15.6× FROM 64,000 FAIRWIND PARTNERS650+REPORTED HARVEY LEGAL19.6%VS 6.7% NEXT MODEL

Google went seven months without a new top-tier AI model. Its last one, Gemini 3.1 Pro Preview, shipped on February 19, according to Tech-ish. On September 30 it finally launched Gemini 4 Argon. Then it told almost everyone they could not use it.

Google's website says that Argon went first to a small group of cybersecurity partners through Google's Fairwind Program. Moneycontrol says the reason is hacking risk. Google has set no public release date. Meanwhile Ars Technica reports that Argon already runs inside Google and is projected to save over 300 14 TiB of data-center memory once rolled out.

The benchmark story is also thinner than the headlines. Argon tops 13 01 of 19 tests in Google's own table. It loses five, and two of those losses are by 9 20 points or more. I have not used Argon, and neither have you, so every number here comes from Google or from reporters reading Google's tables.

By the end of this piece you will have a plan for building around a model you cannot touch yet.

The Four-Rung Release Ladder

Every frontier model now has two dates. One is the day it exists. The other is the day you can call it from your code. I call the space between them the Four-Rung Release Ladder, and Argon shows every rung.

PENCILED ARGON PRICING · OCTOBER 2026LAUNCH COVERAGE · ANADOLU · GoogleBASE: 24 CHECKED + 9 COMPUTED, 4 SHOWN

What Argon would cost you, if you could call it today

Input price Reported launch pricing · per million tokens REPORTED CLAIM 23
$2
Output price Reported launch pricing · per million tokens REPORTED CLAIM 23
$10
Sample task cost 200,000 input + 50,000 output tokens COMPUTED CLAIM 32
$0.90
Single-output cap Anadolu · up from 64,000 tokens VERIFIED CLAIM 15
1M

Rung one is the lab itself. Anadolu Agency reports that thousands of Google employees already use Argon for specialized coding and research work. Ars Technica adds that Argon agents moved more than 800,000 lines of the Fuchsia Zircon kernel from C/C++ to Rust.

Rung two is the vetted outsider. Fairwind started on September 2, 2026 24, and reportedly counts more than 650 partners, including CrowdStrike and Palo Alto Networks. Only some of them get Argon, and those who do reportedly get it with cyber guardrails relaxed so they can find and patch real vulnerabilities.

Rung three is the paying customer. Launch coverage says paid API users and Google AI Ultra subscribers should get access before the general public. Rung four is everybody else, which includes most startups. Google told Anadolu it will open access "as soon as possible," a promise with no date attached.

The ladder comes down to one rule. Your product lives on the rung you can reach today. Plan for the bottom rung and treat anything above it as upside.

Reading Argon's 13 Wins Honestly

A win count hides the size of each win. On Google's table, a 1-point lead and a 10-point loss each count as one result. The margins are what matter, so read them line by line.

Your product lives on the rung you can reach today. Plan for the bottom rung and treat anything above it as upside.· KODA EDITORIAL TEAM · STRATEGY

Argon's clearest wins are in knowledge work. On the Harvey Legal Agent Benchmark it scored 19.6% 17, against 6.7% for the next model. On AutomationBench it posted 51.3% 18, while the closest rival hit 42.5%. One report says its DeepSWE v1.1 score of 77.9% 04 beats GPT-6 Astra and both Claude models.

Several wins are close calls. That is a 1.9 25-point lead. On Vibe Code Bench, its 91.9% 05 trailed Claude Sonnet 5.5's leading 92.39% by about 0.5 points.

The losses come on hard engineering. On FrontierSWE v2, GPT-6 Astra scored 65.5% 06 to Argon's reported 55.0%, a 10.5-point gap. On Terminal-Bench 4.0, Claude Opus 5.5 scored 66.4% 20 to Argon's 57.4%, a 9.0-point gap. Both tests lean on long, tool-heavy engineering work, which looks a lot like what most coding agents do in production.

The strangest number sits right next to the gate. Google restricted Argon because of its cyber power. Yet on CWE-bench v1, the vulnerability benchmark, Argon scored 68.0% 21 and reportedly tied for first. The model judged too risky for a broad release does not own its cyber benchmark outright.

Google's evaluation page has a caveat worth reading. It says scores for rival models come from those providers' own reports. Argon's DeepSWE and Terminal-Bench scores are self-computed, with a mini-swe agent harness used for DeepSWE. Different harnesses can move scores, so a 1.6-point gap deserves far less weight than a 10.5 06-point one.

My read is that the frontier has become a portfolio. Astra and Opus each lead one hard engineering test on this table. Argon leads knowledge work and legal agents. No single model is best at every workload a builder cares about.

Access is turning into an asset in its own right. Firms like CrowdStrike and Wiz can qualify for a gated program because they bring monitoring and accountable human operators. A two-person security startup cannot show that today. So the newest cyber capability reaches incumbents first, an asymmetric advantage handed out by the lab.

You will never have full information about a gated model before you have to decide. A useful rule is to act once you have about 70% of what you would want. On Argon you already have that much: the margins and the missing date.

I think the gate is part safety and part strategy. Google gets adversarial feedback from hundreds of trusted partners while it builds demand among paying customers. Whether this becomes standing policy for every Google flagship, or stays a one-off for an unusually dual-use model, is still unclear.

Narrow leads, odd gates, uneven access

ENGINEERING GAP
10.5 06 PTS

Argon loses where coding agents live.

GPT-6 Astra beat Argon 65.5% to 55.0% on FrontierSWE v2, and Claude Opus 5.5 led Terminal-Bench 4.0 by 9.0 20 points. Both tests lean on long, tool-heavy engineering work that resembles production coding agents.

CYBER PARADOX
68.0% 21

The gated model only ties on cyber.

Google restricted Argon over hacking risk, yet on CWE-bench v1 it scored 68.0% and reportedly tied for first. Whether gating becomes standing policy or stays a one-off for a dual-use model is still unclear.

ACCESS ASYMMETRY
650+

Incumbents climb the ladder first.

Fairwind reportedly counts more than 650 partners, including CrowdStrike and Palo Alto Networks, who bring monitoring and accountable operators. A two-person security startup cannot show that today, so new capability reaches incumbents first.

2031: Access Becomes a Product Spec

The best precedent is on Google's own calendar. In May 2026, Google said Gemini 3.5 Pro would roll out the following month. Business Insider reports it was postponed several times internally because of poor performance, and it never shipped. Any team that pinned a summer roadmap to it lost the summer.

Around that gap, things moved fast. Tech-ish dates Claude Fable 5.1 to September 1 and GPT-6 Astra to two days later. Anthropic shipped Claude Opus 5.5 on September 22, and Argon followed on September 30. That is four frontier launches from three labs inside 30 days 26.

Assume that burst was unusual and the real pace is a quarter of it. That still means one frontier launch a month, or about 60 27 between now and the end of 2031. If half of them arrive gated, builders will read about roughly 30 28 models before they can call them. The public frontier might always trail the real one by a rung or two.

Capability is compounding at the same time. Anadolu reports Argon's output limit jumps from 64,000 tokens 15 to 1 million, about 15.6 29 times in one generation. Bigger single outputs mean longer autonomous runs, and that is the kind of power labs will want to stage.

The gates keep multiplying. Axios reports Google is taking part in the US government's voluntary pre-release process, a second checkpoint beyond Google's own review. Fairwind itself was set up just 28 days 30 before Argon launched. By 2031 I expect investors to ask a startup which access programs it belongs to, right after they ask about its burn rate.

The math favors cheap optionality. Suppose a model router and a test suite cost one engineer two weeks. A roadmap pinned to a model that slips one quarter stalls for about 13 weeks 31. Two weeks of insurance against 13 weeks of exposure pays off every time a launch slips.

Wire a Model Router Before Argon Ships

You do not need a CS degree for this. You need a spreadsheet, a few API keys and an afternoon. Take a deep breath and take it step by step.

First, write down 20 real tasks from your product. Pull them from support tickets and past bugs. Write the pass condition next to each one in plain words. This is your eval suite, which just means a fixed test you rerun on every model.

Then put every model call behind one function. A router is that single function plus a setting that names which model to use. Switching models should mean editing one line of config.

Next, run the 20 tasks on each model you can reach today. Log the pass rate and the cost per completed task, which includes the retries. If your work looks like FrontierSWE or Terminal-Bench, include GPT-6 Astra and Claude Opus 5.5, since they led those tests on Google's table.

Argon's slot in your router should carry a price written in pencil. Reported launch pricing is about $2 23 per million input tokens and $10 per million output tokens, and none of it is contractual yet. A task with 200,000 32 input tokens and 50,000 output tokens would cost about $0.40 plus $0.50, or $0.90. One full 1-million-token output would cost about $10 33 in output alone.

Last, add an entry for Argon in your config and leave it switched off. When access reaches your rung, run the same 20 tasks and compare. If it wins your workload by a wide margin, flip it on. If it wins by 1.9 25 points, the size of its Vals Index lead, weigh price and latency hard before you migrate.

Expect the first run to break. A prompt that works on one model often fails on another, and that failure is the data you wanted. Share your pass rates with your team, or learn in public and post them. Then rerun the suite every time a new model lands, because at one launch a month, the next one is already on its way.

DOJO · BUILD THIS WEEKEND

Wire a model router before Argon ships

  1. Write a 20-task eval suite. Pull 20 real tasks from support tickets and past bugs, and write a plain-words pass condition next to each one.
  2. Put every model call behind one function. Make switching models a one-line config edit, then log pass rate and cost per completed task, retries included, for GPT-6 Astra, Claude Opus 5.5 and anything else you can reach.
  3. Add Argon to your config, switched off. Price it in pencil at $2 23 input and $10 output per million tokens. When access reaches your rung, rerun the 20 tasks and flip it on only if it wins your workload by a wide margin.
Practice: Diagnose and Fix Bad Output
THE BOTTOM LINE

Build on the rung you can reach, and keep a slot open for the one you cannot.

Argon shows that a frontier model now has two dates, the day it exists and the day you can call it. Its 13 01 wins hide narrow margins, and its losses fall on the engineering work most agents actually do. Two weeks spent on a router and an eval suite protect you against the 13 weeks 31 a slipped launch can cost. At one frontier launch a month, the next gated model is already on its way.

LISTEN · AUDIO BRIEFINGThe conversation · ~21 min
WATCH · VISUAL NARRATIVEAnimated breakdown · ~6 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20261001-FD2582C7C5E9
As of01 October 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CHECKED + 9 COMPUTED · 7 VERIFIED · 16 REPORTED · 1 FAILED
7 verified16 reported1 failed9 computed
  1. 01Gemini 4 Argon tops 13 of 19 tests in Google's own benchmark table.VERIFIEDTRUEBENCHMARKcnbc.com
  2. 02Gemini 4 Argon loses five of the 19 tests in Google's benchmark table.VERIFIEDTRUEBENCHMARKblog.google
  3. 03Two of Gemini 4 Argon's benchmark losses in Google's table are by 9 points or more.VERIFIEDTRUEBENCHMARKaa.com.tr
  4. 04Google's announcement says Gemini 4 Argon's DeepSWE v1.1 score of 77.9% beats GPT-6 Astra and both Claude models.REPORTEDMOSTLY TRUEBENCHMARKblog.google
  5. 05On Vibe Code Bench, its 91.9% trailed Claude Sonnet 5.5's leading 92.39% by about 0.5 points.REPORTEDMIXEDBENCHMARKCORRECTED IN COPYdatacamp.com
  6. 06On FrontierSWE v2, GPT-6 Astra scored 65.5% to Argon's reported 55.0%, a 10.5-point gap.REPORTEDMIXEDBENCHMARKCORRECTED IN COPYfrontierswe.com
  7. 07According to Tech-ish, Google's Gemini 3.1 Pro Preview shipped on February 19.REPORTEDMOSTLY TRUEMODELblog.google
  8. 08Google launched Gemini 4 Argon on September 30.REPORTEDMOSTLY TRUEMODELtrendingtopics.eu
  9. 09Google has set no public release date for Gemini 4 Argon.REPORTEDMOSTLY TRUEFEATUREcnbc.com
  10. 10Google's website says that Google's Gemini 4 Argon went first to a small group of cybersecurity partners through Google's Fairwind Program.REPORTEDMOSTLY TRUEATTRIBUTIONdeepmind.google
  11. 11Moneycontrol says Google restricted access to Gemini 4 Argon because of hacking risk.REPORTEDMOSTLY TRUEATTRIBUTIONmoneycontrol.com
  12. 12Ars Technica reports that Gemini 4 Argon already runs inside Google.REPORTEDMOSTLY TRUEATTRIBUTIONarstechnica.com
  13. 13Google went seven months without releasing a new top-tier AI model before launching Gemini 4 Argon.REPORTEDMIXEDHISTORYcnbc.com
  14. 14Meanwhile Ars Technica reports that Argon already runs inside Google and is projected to save over 300 TiB of data-center memory once rolled out.REPORTEDMIXEDSTATCORRECTED IN COPYarstechnica.com
  15. 15Anadolu Agency reports Gemini 4 Argon's output limit jumps from 64,000 tokens to 1 million tokens.VERIFIEDTRUESTATblog.google
  16. 16Axios reports Google is taking part in the US government's voluntary pre-release process for Gemini 4 Argon, adding a checkpoint beyond Google's own review.REPORTEDMOSTLY TRUEPOLICYaxios.com
  17. 17Gemini 4 Argon scored 19.6% on the Harvey Legal Agent Benchmark, against 6.7% for the next-best model.VERIFIEDTRUEBENCHMARKaa.com.tr
  18. 18Gemini 4 Argon scored 51.3% on AutomationBench, while the closest rival scored 42.5%.VERIFIEDTRUEBENCHMARKblog.google
  19. 19Claim removed during the check; its text is not republished.FAILEDMOSTLY FALSEBENCHMARKCUT FROM COPYvals.ai
  20. 20On Terminal-Bench 4.0, Claude Opus 5.5 scored 66.4% to Gemini 4 Argon's 57.4%, a 9.0-point gap.REPORTEDMOSTLY TRUEBENCHMARKanthropic.com
  21. 21On the CWE-bench v1 vulnerability benchmark, Gemini 4 Argon scored 68.0% and reportedly tied for first place.REPORTEDMOSTLY TRUEBENCHMARKcwe-bench.com
  22. 22Argon's DeepSWE and Terminal-Bench scores are self-computed, with a mini-swe agent harness used for DeepSWE.REPORTEDMOSTLY TRUEBENCHMARKCORRECTED IN COPYtech-ish.com
  23. 23Reported launch pricing for Gemini 4 Argon is about $2 per million input tokens and $10 per million output tokens, not yet contractual.REPORTEDMOSTLY TRUEPRICEcnbc.com
  24. 24Google's Fairwind Program started on September 2, 2026.VERIFIEDTRUEHISTORYblog.google
  25. 25Gemini 4 Argon's Vals Index lead over Claude Opus 5.5 (68.9% vs 67.0%) is 1.9 points.COMPUTEDCOMPUTED
  26. 26Claude Fable 5.1, GPT-6 Astra, Claude Opus 5.5 and Gemini 4 Argon make four frontier launches from three labs inside 30 days.COMPUTEDCOMPUTED
  27. 27If the real frontier launch pace is a quarter of the September burst, that is one frontier launch a month, or about 60 launches between now and the end of 2031.COMPUTEDCOMPUTED
  28. 28If half of about 60 frontier launches arrive gated, builders will read about roughly 30 models before they can call them.COMPUTEDCOMPUTED
  29. 29An output limit increase from 64,000 tokens to 1 million tokens is about 15.6 times in one generation.COMPUTEDCOMPUTED
  30. 30Google's Fairwind Program was set up 28 days before Gemini 4 Argon launched (September 2 to September 30).COMPUTEDCOMPUTED
  31. 31Suppose a model router and test suite cost one engineer two weeks, while a roadmap pinned to a model that slips one quarter stalls for about 13 weeks.COMPUTEDCOMPUTED
  32. 32At Gemini 4 Argon's reported pricing, a task with 200,000 input tokens and 50,000 output tokens would cost about $0.40 plus $0.50, or $0.90.COMPUTEDCOMPUTED
  33. 33At Gemini 4 Argon's reported pricing, one full 1-million-token output would cost about $10 in output alone.COMPUTEDCOMPUTED

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20261001-FD2582C7C5E9
Filed underStrategyDeep Dive01 October 2026
Browse the Deep Dive archive

Get the morning Signal

188 editions so far, one a day. Unsubscribe anytime.