Ryan Sael asked Claude Opus 5.5 to show him what is inside the buildings it runs in. One hour and 53 minutes later he had a 3D datacenter you can overload: push a rack too hard and you watch its GPUs throttle and the heat climb out through the roof. His bill for the run was $38.99. The same run without caching would have cost him about $486.
That was one of 17 finished things people shipped with Opus 5.5 in the four days after Anthropic released it on September 22. A camera-focus lab that 2.9 million people watched. A kart racer and a paint shooter, each from what their makers call a single prompt. A 78-second animated film in one HTML file. A Prince of Persia port that got within two pixels of the original game. Sixteen bug-fix pull requests found by formal proofs.
We cut the week into a two-and-a-half-minute film, which you can watch on this page, plus a vertical cut on YouTube Shorts. This piece is about what the receipts say, because the receipts are the story.
The Receipt Test
Launch weeks are loud. Every new model gets a flood of demos, and most of them tell you nothing you can use. So here is the filter we used to pick the 17: a build counts only if it comes with a receipt. A receipt has four lines. How long it took. How much it cost. What the model was actually asked to do. And a finished thing you can open, play or read yourself.
What a finished thing cost this week
The camera lab has all four: one prompt, one hour and 26 minutes, $25.66 in API cost, and a live site. The datacenter has all four, plus the most useful line of the week: 202 calls, 119 million input tokens, and 99% of that input served from cache. Senko's Minecraft and Warcraft clones have them too: about 45 minutes each, $11 to $14 per game, and both games playable next to the same prompts run on Fable 5.1 and GPT-6 Astra.
Builds that failed the test did not make the film. The biggest AI video we found this week, at 12.5 million views, never said which model made it, so it is not here.
The line most people skip
Look at the datacenter receipt again. $38.99 against about $486 is a 12.5x difference, and it comes from one mechanism. An agent resends the whole session on every step, so most of what it reads is text it has already seen. Opus 5.5 charges $0.20 per million tokens to read from cache, which is one twentieth of the $4 it charges for fresh input, and 60% less than Opus 5 charged for the same cache read.
If you run a budget, that is the number to watch. The headline price of a model tells you what a token costs. The cache hit rate tells you what a finished thing costs. In this week's receipts, the finished things landed between $11 and $39.
The patch-review test from a developer who posts as gwd on Hacker News makes the same point from the other direction. Twelve patches, 14 known issues. Opus 5.5 found 8 for $15.40. Fable 5.1 found 7 for $66.34. That is $1.93 per issue found against $9.48, a gap of almost five times, and the cheaper model found more.
Where the week's builds landed
Games and toys from one prompt
A kart racer, a paint shooter, a lawn-mowing game a stream chat joined live, a boat ride through feudal Japan. Treat one prompt as a claim: no maker published a full transcript.
Films written as code
A three-minute history of AI in about an hour, a 78-second film in one HTML file, and a product that turns a link into a launch film.
Engineering you can check
Formal proofs found real bugs in the Agent SDK, and a 1989 game port was checked pixel by pixel against the original. Fewest views, most signal.
Three kinds of build
The week sorted itself into three groups. The first is the one that went viral: games and toys from a single prompt. A kart racer, a paint shooter, a lawn-mowing game his stream chat joined live, a boat ride through feudal Japan with weather and a day and night cycle. Treat "one prompt" as a claim, not a measurement. None of the game makers published a full transcript.
The second group is films written as code. Chubby made a three-minute history of AI in about an hour, roughly 7,400 lines of code with every frame drawn in SVG and canvas. Jurly made a 78-second risograph film in about 45 minutes, in one HTML file with no images, video or libraries. Someone has already turned the idea into a product: LaunchVideo takes a link and returns a launch film in about four minutes. Most of our own film for this piece was made the same way.
The third group got the fewest views and matters most. Boris Cherny, who leads Claude Code at Anthropic, had Opus 5.5 model the Agent SDK's hardest state machines in the Lean proof assistant. It found counterexamples, reproduced them as real bugs, and fixed them: "A couple short prompts = 16 PRs". Priyan R has spent a year porting the 1989 Prince of Persia to C# with each new model. On one prompt, Opus 5.5 found the original room-drawing routine, wrote an unpacker for the compressed game file, and checked its work pixel by pixel against the real game running in DOSBox. The pixels that differed on the first screen went from 8,429 to 2, and those two are a torch flame caught at a different moment.
What the receipts do not prove
Every number here is the builder's own. We opened each post, checked the author, the date and the exact wording, and did not rerun any build. Some claims carry their own caveats. Priyan notes that the key knowledge came from the SDLPoP community's reverse-engineering work, and the model said so itself. gwd corrected a typo in his own results: the Opus 5 cost was $58.19, not $15.19, and "best and cheapest" holds only with the fix. Two of the builders work at Anthropic. We flag each of these in the film.
That is the real skill this week asks of anyone who manages AI work. When a finished thing costs less than a team lunch, building stops being the bottleneck. Choosing what to build, and checking that it is right, becomes the job.
The zoom out
For most of the last three years, the argument about AI and work was about whether the output was good enough. This week's receipts move the argument. A focus lab that teaches optics, a verified set of state machines, a pixel-accurate port: these are good enough to publish, and they came back in hours for tens of dollars. The open question is no longer whether a model can finish the job. It is whether your team has a way to ask for finished things and to check them when they arrive.
That is a management problem, not a model problem, and it is one you can start on this week. Pick one thing you would normally brief to a person. Write the prompt with the acceptance test inside it. Keep the receipt.
Run your own receipt
- Pick one finished thing. Choose something you would normally brief to a person, and write the acceptance test into the prompt itself.
- Keep the receipt. Log the time, the calls, the input tokens and the cache hit rate. The datacenter run shows 99% of input can come from cache.
- Check it against ground truth. Diff the result against something real, the way the game port was checked against the original running in DOSBox.
When a finished thing costs less than lunch, checking it is the job.
This week's receipts show models finishing real work in hours for tens of dollars. The teams that win next are the ones that can ask for finished things precisely and verify them fast. Start with one receipt this week.
