Skip to content
K Koda Intelligence
DEEP DIVE DEEP DIVE № 147 · 25 July 2026DOCKODA-20260725-FBDD159A0870sha-256 of date + article + 2 claims

Your next hire is an agent
with your logins

OpenAI shipped ChatGPT Work on July 9, 2026, three days after workspace agent runs stopped being free and started drawing down credits. The company says nearly 100% of its own teams now run on ChatGPT Work and Codex, and Codex has passed 5 million weekly users, more than 1 million of them non-programmers. The output is not a chat reply. It is finished sheets, slides, docs, and shareable web apps, built across your connected apps and files.

4 MIN READ · BY THE KODA EDITORIAL TEAM · TOOLS · AGENTS
CHATGPT WORK LAUNCHJUL 9NOT MEASUREDOPENAI 2026
AGENT RUNS METEREDJUL 6NOT MEASUREDWORKSPACE CREDITS
CODEX WEEKLY USERS5MNOT MEASUREDOPENAI
CHATGPT WORK LAUNCHJUL 9OPENAI 2026 AGENT RUNS METEREDJUL 6WORKSPACE CREDITS CODEX WEEKLY USERS5MOPENAI NON-CODER USERS1M+OPENAI PLUGIN DIRECTORY1,400+OPENAI PLUS AGENT CEILING40/MOFINDSKILL ENTERPRISE APPS W/ AGENTS40%GARTNER 2026 AGENTS IN PRODUCTION51%SURVEY WORK

OpenAI shipped ChatGPT Work on July 9, 2026. Three days earlier, on July 6, workspace agent runs stopped being free and started drawing down workspace credits. That order matters more than the launch video.

Here is the second number. OpenAI says nearly 100% of its own teams, finance and sales included, now run on ChatGPT Work and Codex. Codex passed 5 million weekly users, and more than 1 million of those people are not programmers.

And the output is not a chat reply. Per OpenAI's own launch materials, you brief it and get back finished sheets, slides, docs, and shareable web apps. Reuters described it plainly on July 9: an agent "designed to execute tasks across different applications and files."

So the product category quietly changed. You are no longer buying a smarter answer machine. You are buying a coworker with your logins. That trade has real upside and a real bill attached, and most teams have not read the fine print on either.

The Handoff Line

Every AI task you own sits on one side of a single line.

AGENT ADOPTION LEDGER · JULY 2026OPENAI · GARTNER · FINDSKILL · SURVEY WORKBASE: 2 CHECKED CLAIMS, 4 SHOWN

Four numbers that show chat becoming task execution.

OpenAI internal teams on ChatGPT Work and Codex OpenAI launch materials · finance and sales included NOT MEASURED
~100%
Enterprise apps shipping task-specific agents by end of 2026 Gartner · up from under 5% in 2025 NOT MEASURED
40%
Enterprises with agents already in production Survey work · roughly 74% past the pilot stage NOT MEASURED
51%
Agent messages per month on Plus FindSkill reporting · shared allowance with Codex NOT MEASURED
40

On the conversation side, you stay present. You prompt, you read, you edit, you prompt again. The output is text and you are the assembly line. Value scales with your attention, which means it does not scale.

On the handoff side, you leave. You write a brief, walk away, come back to a file. The output is an artifact, and your job shifts from producing to checking.

Call it the Handoff Line. Everything on your task list is Chatty, Choppy, or Clean. Chatty tasks need your judgment at every step, so keep them in chat. Choppy tasks are handoffs with no template and no source of truth, which is where agents burn credits and hand back confident garbage. Clean tasks repeat, have a known shape, and pull from data the agent can actually reach.

The whole 2026 agent boom is one sentence: the number of Clean tasks just went up. Not because the models got mystical, but because the connectors got boring. ChatGPT Work gathers context from Slack, Gmail, Google Drive, CRM platforms, and internal knowledge bases, and OpenAI's unified plugin directory now lists more than 1,400 plugins. Plumbing, not magic.

The Only 20% Worth Automating First

Here is the 20% that gets 80% of the result. Ignore the flashy demo. Find the recurring task that takes a competent person about four hours and produces a structured, predictable output.

An agent designed to execute tasks across different applications and files.· REUTERS · JULY 9, 2026

The weekly competitor report. The monthly board deck. The quarterly vendor analysis. The onboarding doc someone updates whenever a process changes. Four hours a week is 208 hours a year, which is five working weeks of a salaried person spent reformatting things that already exist somewhere.

Run every candidate task through three questions before you hand it over. Does it have a template, meaning a previous version you can point at and say "like this"? Does it have a source, meaning the data lives in a system the agent can read rather than in someone's head? Does it have a checker, meaning a specific human who can spot a wrong answer in under ten minutes?

No template, no source, or no checker, and you are automating a guess. That is how you get a Ferrari deliverable: gorgeous slides, clean charts, no engine underneath. A Tractor output is ugly and correct, and I will take the Tractor every single time on a board deck. The Unicorn, pretty and correct, only shows up after you have fed the thing your actual format.

Now the part the launch coverage soft-pedals. This agent is a 500 IQ intern who has your password. Current agent designs typically run with the same permissions as the user who launched them, so a compromised agent looks a lot like a compromised employee account across mail, drives, and repos. Prompt injection, where hidden instructions inside a webpage or document hijack the agent, is still an unsolved problem, not a patched one.

So scope before you scale. An ounce in pre is worth a pound in post. Give the agent read access to the folder it needs, not the drive. Route anything that sends, pays, or publishes through an approval step. Log agent actions separately from human ones, because if you cannot tell them apart later, you cannot audit anything.

Then watch the meter. Because agents take real actions in real browsers, single runs can stretch from a few minutes to well over half an hour, and long runtime is not the same as progress. On Plus, reporting from FindSkill puts the ceiling around 40 agent messages a month, shared with Codex. Whether that allowance survives contact with a team that actually likes the product is anyone's guess.

2031. Zoom out. Chat rented your attention. Agents rent your permissions. Those are wildly different businesses with wildly different switching costs.

The market math is already loud. Gartner expects 40% of enterprise applications to ship task-specific agents by the end of 2026, up from under 5% in 2025. Survey work puts about 51% of enterprises with agents already in production and roughly 74% past the pilot stage.

Behavior is moving with it. Across the agent ecosystem, the 99.9th percentile session length roughly doubled between October 2025 and January 2026, from under 25 minutes to over 45. People are switching from human-in-the-loop to human-on-the-loop.

Three signals inside the same shift

OUTPUT CHANGED
1,400+

The deliverable is an artifact, not a transcript.

ChatGPT Work returns finished sheets, slides, docs, and shareable web apps after a single brief. It pulls context from Slack, Gmail, Google Drive, CRM platforms, and internal knowledge bases, and OpenAI's unified plugin directory now lists more than 1,400 plugins. The connectors got boring, which is exactly why the Clean task list grew.

PERMISSION RISK
500 IQ

A 500 IQ intern who has your password.

Current agent designs typically run with the same permissions as the user who launched them, so a compromised agent looks like a compromised employee account across mail, drives, and repos. Prompt injection, where hidden instructions inside a webpage or document hijack the agent, remains unsolved rather than patched. Scope read access to the folder, not the drive.

LONGER LEASH
45 MIN

Sessions doubled as humans stepped back.

Across the agent ecosystem, the 99.9th percentile session length roughly doubled between October 2025 and January 2026, from under 25 minutes to over 45. Single runs can now stretch from a few minutes to well over half an hour. Long runtime is not the same as progress, and metered credits make that difference expensive.

DOJO · BUILD THIS WEEKEND

Hand off one Clean task and audit the bill.

  1. Find your four-hour task. Pick the recurring job that takes a competent person about four hours and produces structured output, like the weekly competitor report or monthly board deck. Four hours a week is 208 hours a year, five working weeks of reformatting things that already exist somewhere.
  2. Run the template, source, checker test. Before handing anything over, confirm there is a previous version to point at, data the agent can actually read, and a specific human who can spot a wrong answer in under ten minutes. Miss any one of the three and you are automating a guess.
  3. Scope permissions before you scale. Give read access to the folder the agent needs, not the drive, and route anything that sends, pays, or publishes through an approval step. Log agent actions separately from human ones, and watch your meter against the roughly 40 agent messages Plus reportedly allows per month.
Train the full skill in The Dojo
THE BOTTOM LINE

Chat rented your attention. Agents rent your permissions.

The launch video is about capability, but the July 6 metering change is about the business model, and the permission model is about your risk. Sort every task into Chatty, Choppy, or Clean, then hand over only the Clean ones with a template, a source, and a checker attached. Prefer the Tractor output that is ugly and correct over the Ferrari deck that is gorgeous and hollow, because the Unicorn only appears after you have fed the system your real format. Gartner expects 40% of enterprise applications to ship task-specific agents by the end of 2026, up from under 5% in 2025, so the question is not whether you adopt but how narrowly you scope. Teams that treat the agent like a new employee account, logged and approval-gated, will still be running theirs in 2031.

LISTEN · AUDIO BRIEFINGThe conversation
WATCH · VISUAL NARRATIVEAnimated breakdown · ~2 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20260725-FBDD159A0870
As of25 July 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE2 CLAIMS CHECKED · 2 FAILED
2 failed
  1. 01$10.91FAILEDFALSESTAT
  2. 02Claim removed during the check; its text is not republished.FAILEDMOSTLY FALSESTATCUT FROM COPY

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20260725-FBDD159A0870
Filed underToolsDeep Dive25 July 2026
Browse the Deep Dive archive

Get the morning Signal

163 editions so far, one a day. Unsubscribe anytime.