Koda Score = weighted blend: capability 35, ease 20, value 25, momentum 20.
Computer Use, Browser Use, Skills and Files APIs for building real agent loops
The four pieces you were duct-taping together to make an agent actually do work are now GA on one API, and one of them drives web apps by page structure instead of guessing pixel coordinates.

This is for engineering teams who already have a concrete, boring automation problem: an internal web tool with no API, a document pipeline, a procedure nobody has written down properly. The move from preview to GA matters because you can now put these in a production path without betting on a deprecation notice, and Browser Use targeting page elements rather than screen coordinates is the difference between a demo and a workflow that survives a layout change. Adopt it for one narrow, auditable task first; the tooling is ready, your permissions model probably is not.
App control for agents that need to operate software directly, moved out of preview into general availability.
A new tool for agents that drive web applications, described as more robust because it targets page structure and elements instead of guessing screen coordinates.
Upload, version and reuse internal procedures and workflows as capabilities your agents can call.
Programmatic document work with higher limits and managed storage, so file handling stops being your own plumbing problem.
Access to all Claude models from Haiku 4.5 up to Fable 5, with usage-based pricing and prompt caching on every tier.
Batch jobs plus automatically increasing rate limits as your usage builds, which matters when agent loops multiply calls.
Residency choices including US-only inference, offered at 1.1x pricing for input and output tokens.
Web search and fetch, code execution and memory sit alongside the new tools inside the same platform surface.
Concrete setups pulled from the research, not feature-list hand-waving.
Internal tools engineer at a company running legacy vendor software
Point Browser Use at the internal web tool that has no API, define the extraction task in natural language, and log every action the agent takes to an audit trail before anyone trusts the output.
Structured data out of a system you were never going to get integration budget for, with a record of exactly what the agent clicked.
Ops lead who owns three undocumented runbooks
Encode each runbook as a Skill via the Skills API, version it, and have agents call the skill instead of re-explaining the procedure inside a sprawling system prompt.
One place to fix a process, and agent behavior that changes when you ship a new skill version rather than when someone edits a prompt.
Backend team handling contract or report intake
Use the Files API for upload and managed storage, then run analysis with Opus 5 or Sonnet 5 depending on how much judgment the documents demand.
Less custom storage code and a clean split between cheap triage and expensive reasoning.
Security-conscious platform team
Give the agent a task-specific credential with the minimum permissions for one workflow, run it under Computer Use, and review the action log daily for a week before widening scope.
A repeatable permissions pattern you can defend in review, instead of discovering later that the agent had admin.
Cheapest tier. Prompt caching at $1.25 / MTok write, $0.10 / MTok read.
Mid tier. Prompt caching at $2.50 / MTok write, $0.20 / MTok read.
Cited by Box, Factory and Zapier for enterprise analysis and automation work. Faster speeds available at higher pricing.
Top of the listed lineup, cited by Cursor, Cognition, Replit and Genspark on coding and long-horizon evals.
The pricing ladder is wide enough that a well-designed agent loop can triage on Haiku and escalate to Opus or Fable only when it matters, but agent loops burn tokens by nature and the top tiers get expensive fast without prompt caching discipline.
Pricing as captured on 2026-08-25. Check the live site before you commit.
Real reactions surfaced during research. Paraphrased faithfully, linked to source.
Too thin to call: the research surfaced only article-style coverage and reposts, with no verified reactions from Reddit, Hacker News, X, G2 or Product Hunt, so there is no street sentiment to report yet.
The closest rival on the agent loop plus desktop and browser control angle.
PICK IT WHENYou are already standardized on OpenAI models and would rather not run a second vendor for agent execution.
Comparable option for web-task automation with agentic models.
PICK IT WHENYour data and identity already live in Google Cloud and web automation is the only piece you need.
Deterministic browser automation framework, no model in the loop.
PICK IT WHENThe workflow is stable and repeatable and you want exact, testable behavior rather than natural-language tasking.
Time to first value: about an afternoon for a first audited workflow
Create an account at platform.claude.com and pull an API key, then read the tool docs for Browser Use and Skills.
Pick one narrow workflow, such as extracting data from an internal web tool with no API, and issue a task-scoped credential with minimum permissions.
Run the loop against a Sonnet 5 or Haiku 4.5 baseline, log every action the agent takes, and review the log before you widen scope or upgrade models.
Computer Use, the Skills API and the Files API moved from preview to general availability, and Browser Use is new. The practical difference is that you can now build these into a production path rather than treating them as experiments.
Computer Use is broader app control, while Browser Use is aimed at agents driving web applications. The coverage describes it as more robust for web tasks because it targets page structure and elements instead of guessing screen coordinates.
Start on Sonnet 5 at $2 / MTok input and $10 / MTok output, or Haiku 4.5 for triage, and escalate only where judgment matters. Opus 5 is the tier enterprise customers like Box, Factory and Zapier cite for analysis and automation work, and Fable 5 is the top tier cited on coding and long-horizon evals.
Permissions. An agent that can drive a browser or an app inherits whatever access you hand it, so scope credentials per task and log every action from day one. Treat the first workflow as an audit exercise, not a productivity win.
Pick the single workflow that costs your team the most manual clicking, wire it with scoped credentials and a full action log, and only then talk about scaling the pattern.
Field research: 1 pages scraped · 3 search passes · 12 community sources. Reviewed by the Koda desk on 2026-08-25.