K Koda Intelligence
Subscribe
Cover image for the 05 October 2026 Signal: DeepSeek's V4.1 Flash scored 81.1 on LiveBench, leaving Anthropic's leading model only roughly 3% ahead at 83.4.
THE SIGNAL · 05 OCT 2026 · 8 MIN READ

DeepSeek cuts the gap with U.S. AI models to roughly 3%

Bloomberg Intelligence ranked DeepSeek V4.1 Flash sixth globally on LiveBench for September, and it now sits within roughly 3%1 of Anthropic's leading model. Model shoppers should treat price as a live variable again.

AIMODELSSPACE
21 of 29 figures cited10 receipts in the register
  1. 3%1Bloomberg
  2. 1 Billion2Reuters
  3. 458,9673The Rundown AI
The register
Lead Story

The day's defining move.

China · Bloomberg
Lead storyChina 3%1 Bloomberg · 04 OCT 2026
China·Bloomberg·04 OCT 2026

DeepSeek gains cut gap with U.S. to 3%1

Bloomberg Intelligence ranked DeepSeek's V4.1 Flash sixth globally on LiveBench for September, scoring 81.1 against 83.4 for Anthropic's leading model. The September-launched model cut the reported gap between top Chinese and U.S. models to roughly 3%1. The report says the gains point to further market share wins for Chinese contenders.

Why it mattersWith V4.1 Flash scoring 81.1 against Anthropic's 83.4 on a single September LiveBench reading, test DeepSeek against Anthropic on your own workload and per-token price before renewing a model contract.

Continue reading
Markets

The Nasdaq led the indexes at +1.19% against +0.73% for the S&P 500, a spread that shows the Nasdaq outpacing the broader benchmark, while WTI crude slipped 0.55%, Bitcoin rose 2.11% and the crypto Fear & Greed index read 70 (Greed).

SPONSOR KODAFrom the desk

Put your product in front of AI operators.

Koda publishes a fact-checked AI intelligence briefing every morning: site, email, podcast, and video. Founding sponsors lock today's rates for 6 months as the audience compounds.

View the media kit →
Listen & Watch

Daily broadcasts.

Podcast · Video · Infographic
YouTube · Short

Today's Signal Short

Visual

Intelligence map

Markets

6 levels at the close.

Market TerminalLAST CLOSE 02 OCT 2026 · READ 2026-10-05 · Yahoo Finance
S&P 500 7,722.724 +0.73%4 6D, window short 7,651.54 low / 7,743.41 high
Nasdaq 27,190.865 +1.19%5 6D, window short 26,797.54 low / 27,190.86 high
Bitcoin $86,555.136 +2.11%6 6D, window short $83,553.85 low / $86,555.13 high
Ethereum $2,725.947 +1.44%7 6D, window short $2,668.15 low / $2,725.94 high
Crude Oil (WTI) $90.618 -0.55%8 6D, window short $89.38 low / $92.87 high
Crypto Fear & Greed 70 / 1009 Greed 0 to 100 index 25 fear / 50 neutral / 75 greed · alternative.me
Today's Focus

What else the day turns on.

01

DoorDash Wins FAA Drone Certification

DoorDash has won FAA Part 135 air carrier certification, letting it run its own drone delivery fleet, DoorDash Air, alongside its marketplace. Drones will take mid-distance orders of 3 to 5 miles, and human Dashers will keep the shorter, trickier ones.

02

Altman Frames Anthropic Policy Divide

In an interview published October 4 and 5, Sam Altman said AI's upside warrants tolerating certain risks and criticized calls for tighter regulatory restrictions. He framed the issue as a fundamental policy divide with Anthropic, openly positioning OpenAI as the pro-access lab.

The Lab

Tool of the day, researched.

All Lab reports
GPTBots home page, captured for the 5 October 2026 Lab report Tool of the day 6.710/ 10
AI Platform Pricing not published about 10 minutes

GPTBots

GPTBots says its agents automate 90% of customer service issues, and the cheapest way to learn whether your tickets fall in that 90% is an internal policy bot that no customer ever sees.

GPTBots bundles a no-code agent builder, knowledge bases trained on your own documents, and channel deployment behind one login, and its G2 page reports 4.8 out of 5 from 70 reviews. The catch is that plan prices are absent from its site and developer documentation draws repeated complaints. Support and operations teams should pilot it on one contained internal agent first, then decide on customer-facing use once they have their own accuracy numbers.

Capability 7.0
Ease 7.5
Value not scored
Momentum 5.5

Built a working bot in about 10 minutes with no coding at all.Paraphrased · g2.com

Also on the radar
Coding

JetBrains Air: run several coding agents side by side in your IDE

JetBrains Air is a multi-agent coding environment for JetBrains IDEs that can call Codex, Copilot, Claude, Junie, and other agent providers from one place. Assign the same scoped task to two providers, then compare their diffs before you merge. Keep one agent on implementation and another on tests so each one's output checks the other's.

Try it
Productivity

Capy Desktop: keep long coding jobs running after you close the laptop

Capy Desktop orchestrates coding agents across macOS, Windows, and Linux. It offers remote shell and file control, and jobs persist in the cloud. Start a long refactor or test run before you leave your desk, then check progress and pull files from another machine. This works well for overnight migrations that would otherwise tie up your laptop.

Try it
Build

Atlassian Governed Agent Loops: take Jira tickets to pull requests with an audit trail

Now in open beta, Governed Agent Loops automates Jira work from a backlog item through to a pull request and applies governance controls along the way. It ships alongside Code Context, which indexes multiple repositories so agents can see how services connect. Pilot it on a small set of well-specified bug tickets first, and review which approvals the governance layer actually enforces.

Try it
Coding

ByteAsk: a coding agent that compiles and tests its own C and C++ changes

ByteAsk is a coding agent from Y Combinator's F2026 batch built for C and C++. It builds, debugs, and tests every change with your real toolchain until the change holds, so it doesn't stop at a plausible-looking diff. It runs from the terminal or inside VS Code, Neovim, and Emacs, and a fully on-prem option suits teams working on trading, automotive, or chip code that can't leave the building.

Try it
The Arena

3 labs moved this week.

  1. OpenAIAltman argues AI's benefits justify accepting some risk thestar.com.my · 05 OCT 202601
  2. Google DeepMindGemini consumer tiers get reshuffled starting October 9 sharjah24.ae · 04 OCT 202602
  3. AnthropicClaude Sonnet 5.5 reportedly ships ahead of planned IPO finance.sina.com.cn · 05 OCT 202603
Builder Radar

1 receipt from the repos, one paper.

Primary sources, pulled by index
Receipts
  1. Behind: JetBrains Air: run several coding agents side by side in your IDE DocsJetBrains IDEs

    Use Claude Code with JetBrains IDEs including IntelliJ, PyCharm, WebStorm, and more Claude Code integrates with JetBrains IDEs through a dedicated plugin, providing features like interactive diff viewing, selection context sharing, and...

    code.claude.com
Paper of the day

Incident-Arena: Getting agents to the last nine of reliability

Andre Fu, Malik Drabla, Leon Liu et al. · arXiv 2026-09-30

AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response.

While useful for control, they leave uncertainty about measured performance transferring to real-world systems. Furthermore, frontier models are already capable of detecting evaluation regimes ( Needham et al., 2025), therefore artificial interfaces, small...
Read the paper

Receipts come from the Firecrawl Developer Index (READMEs, docs, issues, merged PRs), first-party sources only. The paper comes from the Firecrawl Research Index: arXiv, indexed in the last 7 days and submitted this month or last, picked by today's research theme. Quotes are verbatim.

THE DOJO · BUILD TODAY

Price your models by the work they actually ship

  1. 01

    Score models by cost per accepted task. Log which outputs your team keeps, divide model spend by that count, and treat any benchmark gap under 2.7 points as a tie.

  2. 02

    Trial GPTBots on one document set. Load a single internal knowledge base and deploy to one channel, then request a written quote before adding seats.

  3. 03

    Pilot Atlassian Governed Agent Loops beside JetBrains Air. Run both on one low-risk repo for a week and note each point where a human has to sign off before an agent action lands.

The Bottom Line

The frontier is crowded enough that switching costs now matter more than leaderboard rank.

DeepSeek closing to roughly 3%1 of the U.S. leaders lands just as Anthropic reportedly ships Claude Sonnet 5.5 ahead of its planned IPO, which means every top lab now has a reason to compete on price and release cadence. Builders gain leverage from that pressure only if their stack can move between providers quickly. Keep prompts and evals provider-agnostic, and keep a cheap challenger model wired in so a better offer becomes a config change.

Source: bloomberg.com
Receipts

Every number, and where it came from.

21 of 29 figures cited, 10 receipts: 1 verified, 8 reported, 1 estimated
  1. 013%DeepSeek gains cut gap with U.S. to 3%Bloomberg04 OCT 202603B search corroborationV
  2. 021 BillionMerz Pledges €1 Billion More for Ukraine.Reuters04 OCT 2026as publishedR
  3. 03458,967Rundown cites email costs across 458,967 readers as reason to prune its listThe Rundown AI04 OCT 2026as publishedR
  4. 047,722.72S&P 500 close, +0.73% on the sessionYahoo Finance02 OCT 2026yfinance closeR
  5. 0527,190.86Nasdaq close, +1.19% on the sessionYahoo Finance02 OCT 2026yfinance closeR
  6. 06$86,555.13Bitcoin close, +2.11% on the sessionYahoo Finance05 OCT 2026yfinance closeR
  7. 07$2,725.94Ethereum close, +1.44% on the sessionYahoo Finance05 OCT 2026yfinance closeR
  8. 08$90.61Crude Oil (WTI) close, -0.55% on the sessionYahoo Finance04 OCT 2026yfinance closeR
  9. 0970 / 100Crypto Fear & Greed close, Greed on the sessionalternative.me05 OCT 2026as publishedR
  10. 106.7 / 10GPTBots, Koda score across four dimensions04R dossier05 OCT 202604R dossier scoringE
V
Verified: the stat gate corroborated this figure against an independent search before publication.
R
Reported: the figure is carried as published by the linked source and was not independently corroborated.
E
Estimated: the figure is approximate or hedged, either in the source or by the stat gate.

Method names the computation path behind the row, from a closed vocabulary: as published, yfinance close, 03B search corroboration, 04R dossier scoring, not recorded.

A figure is money, a percentage, a spelled magnitude, a multiple, a rate carrying its unit, a score over its denominator, a thousands separated number or an index close. Bare integers and years are not counted, so a phrase like "over 100 integrations" carries no mark. A chart's own scale line is the axis of the level above it and is counted once, with that level.

Get the morning Signal

192 editions so far, one a day. Unsubscribe anytime.

Forward this to one operator you work with. Your referral link is in every email; milestones at koda.community/refer.