OpenAI shipped ChatGPT Work on July 9, 2026. Three days earlier, on July 6, workspace agent runs stopped being free and started drawing down workspace credits. That order matters more than the launch video.
Here is the second number. OpenAI says nearly 100% of its own teams, finance and sales included, now run on ChatGPT Work and Codex. Codex passed 5 million weekly users, and more than 1 million of those people are not programmers.
And the output is not a chat reply. Per OpenAI's own launch materials, you brief it and get back finished sheets, slides, docs, and shareable web apps. Reuters described it plainly on July 9: an agent "designed to execute tasks across different applications and files."
So the product category quietly changed. You are no longer buying a smarter answer machine. You are buying a coworker with your logins. That trade has real upside and a real bill attached, and most teams have not read the fine print on either.
The Handoff Line
Every AI task you own sits on one side of a single line.
Four numbers that show chat becoming task execution.
On the conversation side, you stay present. You prompt, you read, you edit, you prompt again. The output is text and you are the assembly line. Value scales with your attention, which means it does not scale.
On the handoff side, you leave. You write a brief, walk away, come back to a file. The output is an artifact, and your job shifts from producing to checking.
Call it the Handoff Line. Everything on your task list is Chatty, Choppy, or Clean. Chatty tasks need your judgment at every step, so keep them in chat. Choppy tasks are handoffs with no template and no source of truth, which is where agents burn credits and hand back confident garbage. Clean tasks repeat, have a known shape, and pull from data the agent can actually reach.
The whole 2026 agent boom is one sentence: the number of Clean tasks just went up. Not because the models got mystical, but because the connectors got boring. ChatGPT Work gathers context from Slack, Gmail, Google Drive, CRM platforms, and internal knowledge bases, and OpenAI's unified plugin directory now lists more than 1,400 plugins. Plumbing, not magic.
The Only 20% Worth Automating First
Here is the 20% that gets 80% of the result. Ignore the flashy demo. Find the recurring task that takes a competent person about four hours and produces a structured, predictable output.
The weekly competitor report. The monthly board deck. The quarterly vendor analysis. The onboarding doc someone updates whenever a process changes. Four hours a week is 208 hours a year, which is five working weeks of a salaried person spent reformatting things that already exist somewhere.
Run every candidate task through three questions before you hand it over. Does it have a template, meaning a previous version you can point at and say "like this"? Does it have a source, meaning the data lives in a system the agent can read rather than in someone's head? Does it have a checker, meaning a specific human who can spot a wrong answer in under ten minutes?
No template, no source, or no checker, and you are automating a guess. That is how you get a Ferrari deliverable: gorgeous slides, clean charts, no engine underneath. A Tractor output is ugly and correct, and I will take the Tractor every single time on a board deck. The Unicorn, pretty and correct, only shows up after you have fed the thing your actual format.
Now the part the launch coverage soft-pedals. This agent is a 500 IQ intern who has your password. Current agent designs typically run with the same permissions as the user who launched them, so a compromised agent looks a lot like a compromised employee account across mail, drives, and repos. Prompt injection, where hidden instructions inside a webpage or document hijack the agent, is still an unsolved problem, not a patched one.
So scope before you scale. An ounce in pre is worth a pound in post. Give the agent read access to the folder it needs, not the drive. Route anything that sends, pays, or publishes through an approval step. Log agent actions separately from human ones, because if you cannot tell them apart later, you cannot audit anything.
Then watch the meter. Because agents take real actions in real browsers, single runs can stretch from a few minutes to well over half an hour, and long runtime is not the same as progress. On Plus, reporting from FindSkill puts the ceiling around 40 agent messages a month, shared with Codex. Whether that allowance survives contact with a team that actually likes the product is anyone's guess.
Three signals inside the same shift
The deliverable is an artifact, not a transcript.
ChatGPT Work returns finished sheets, slides, docs, and shareable web apps after a single brief. It pulls context from Slack, Gmail, Google Drive, CRM platforms, and internal knowledge bases, and OpenAI's unified plugin directory now lists more than 1,400 plugins. The connectors got boring, which is exactly why the Clean task list grew.
A 500 IQ intern who has your password.
Current agent designs typically run with the same permissions as the user who launched them, so a compromised agent looks like a compromised employee account across mail, drives, and repos. Prompt injection, where hidden instructions inside a webpage or document hijack the agent, remains unsolved rather than patched. Scope read access to the folder, not the drive.
Sessions doubled as humans stepped back.
Across the agent ecosystem, the 99.9th percentile session length roughly doubled between October 2025 and January 2026, from under 25 minutes to over 45. Single runs can now stretch from a few minutes to well over half an hour. Long runtime is not the same as progress, and metered credits make that difference expensive.
2031
Zoom out. Chat rented your attention. Agents rent your permissions. Those are wildly different businesses with wildly different switching costs.
The market math is already loud. Gartner expects 40% of enterprise applications to ship task-specific agents by the end of 2026, up from under 5% in 2025. Survey work puts about 51% of enterprises with agents already in production and roughly 74% past the pilot stage.
Behavior is moving with it. Across the agent ecosystem, the 99.9th percentile session length roughly doubled between October 2025 and January 2026, from under 25 minutes to over 45. People are switching from human-in-the-loop to human-on-the-loop, which is a pol
Hand off one Clean task and audit the bill.
- Find your four-hour task. Pick the recurring job that takes a competent person about four hours and produces structured output, like the weekly competitor report or monthly board deck. Four hours a week is 208 hours a year, five working weeks of reformatting things that already exist somewhere.
- Run the template, source, checker test. Before handing anything over, confirm there is a previous version to point at, data the agent can actually read, and a specific human who can spot a wrong answer in under ten minutes. Miss any one of the three and you are automating a guess.
- Scope permissions before you scale. Give read access to the folder the agent needs, not the drive, and route anything that sends, pays, or publishes through an approval step. Log agent actions separately from human ones, and watch your meter against the roughly 40 agent messages Plus reportedly allows per month.
Chat rented your attention. Agents rent your permissions.
The launch video is about capability, but the July 6 metering change is about the business model, and the permission model is about your risk. Sort every task into Chatty, Choppy, or Clean, then hand over only the Clean ones with a template, a source, and a checker attached. Prefer the Tractor output that is ugly and correct over the Ferrari deck that is gorgeous and hollow, because the Unicorn only appears after you have fed the system your real format. Gartner expects 40% of enterprise applications to ship task-specific agents by the end of 2026, up from under 5% in 2025, so the question is not whether you adopt but how narrowly you scope. Teams that treat the agent like a new employee account, logged and approval-gated, will still be running theirs in 2031.