KKoda IntelligenceDaily Signal
DEEPSWE (%)62.7↑ 12.8%·TERM-BENCH 2.187.9↑ 1.2%·NEMOTRON MOE30B· FLAT·S&P 5007,748.50↑ 0.26%·NASDAQ26,588.49↑ 0.54%·BTC$63,385.76↓ 0.26%·ETH$1,875.61↓ 0.31%·FEAR INDEX29↓ 2.8%·
THE SIGNAL · 13 AUG 2026 · 5 MIN READ

DeepSeek's agentic score just quintupled

V4 Pro 0813 pushes DeepSWE from 12.8% to 62.7% and Terminal-Bench 2.1 to 87.9, while API prices rise in the same breath. Agentic coding is now the scoreboard that matters.

AIBENCHMARKSLOCAL INFERENCE
KODA PROFrom the desk

The Operator Tier is coming.

A weekly operator deep dive, the full Dojo Pro prompt packs, and the complete prompt database. Founding members lock the launch price forever.

Join the founding list →
Lead Story

The day's defining move.

Model Release · Sina Tech
Lead Story
Model Release·Sina Tech·13 August 2026

DeepSeek V4 Pro Posts Huge Agentic Jump

DeepSeek formally released V4 Pro 0813 with reported gains concentrated in agentic and coding work: DeepSWE climbing from 12.8% to 62.7% and Terminal-Bench 2.1 reaching 87.9. The jump lands days after DeepSeek warned customers that its API prices are expected to rise, making the model a test of whether buyers...

Continue reading arrow_outward
smart_toyAI

DeepSeek's V4 Pro 0813 lifts DeepSWE from 12.8% to 62.7% and hits 87.9 on Terminal-Bench 2.1, days after the lab told customers API prices are expected to rise.

publicWorld

Trump declared on August 12 that the United States controls the Strait of Hormuz and will "keep it," describing a "wall of steel" naval posture while the chokepoint stays effectively closed to normal traffic.

trending_upMarkets

Mood is Fear: a closed Hormuz and rising frontier API rates both point the same direction on input costs.

boltWild Card

Three of today's five AI stories are about cheapness rather than intelligence: NVIDIA's Nemotron 3.5 Lightning activates 3B of 30B parameters, Unsloth Desktop claims 70% less VRAM, and DeepSeek is testing whether buyers will pay more anyway.

Markets

Market snapshot.

query_statsMarket TerminalLive
S&P 500 7,748.50 arrow_drop_up+0.26%
Nasdaq 26,588.49 arrow_drop_up+0.54%
Bitcoin $63,385.76 arrow_drop_down-0.26%
Ethereum $1,875.61 arrow_drop_down-0.31%
Crude Oil (WTI) $82.17 arrow_drop_down-1.24%
Crypto Fear & Greed 29 arrow_drop_downFear
Today's Focus

The three signals that move the day.

01

Agentic Coding Becomes The Benchmark

DeepSeek's headline gains are concentrated in DeepSWE and Terminal-Bench 2.1, not chat quality, and NVIDIA aimed Nemotron 3.5 Lightning at code review, tool use and security alert monitoring. The competitive axis has moved from what a model knows to what it can finish unattended. Price per completed task, not per token, is the number to watch.

02

Local Inference Gets A Desktop

Unsloth Desktop landed on August 11 as a free open-source app for macOS, Windows and Linux, covering GGUF, MLX, diffusion image and video, and audio models with claimed 2x faster training at 70% less VRAM. Paired with a 30B model that activates 3B parameters per token, the on-premises floor is rising just as hosted API prices signal an increase.

03

Google Reorganizes While Shipping Hardware

Koray Kavukcuoglu sits at the center of a leadership reshuffle, with fresh concerns that DeepMind's autonomy is being absorbed into Google proper, and the company emphasized Flash demand and its Cyber model rather than a frontier launch. Two days later at Made by Google '26 it pushed Gemini into Pixel 11, Pixel Watch 5 and a new Pixel Tag. Distribution to hundreds of millions of handsets is the hedge no rival lab can copy.

Listen & Watch

Daily broadcasts.

Podcast · Video · Infographic
The Daily Deep Dive

Listen to today's briefing

YouTube · Short

Today's Signal Short

Visual

Intelligence map

The Wire

AI intelligence.

Model Release·Sina Tech

DeepSeek V4 Pro Posts Huge Agentic Jump

DeepSeek formally released V4 Pro 0813 with reported gains concentrated in agentic and coding work: DeepSWE climbing from 12.8% to 62.7% and Terminal-Bench 2.1 reaching 87.9. The jump lands days after DeepSeek warned customers that its...

Read arrow_outward
Model Release·NVIDIA

Nemotron 3.5 Lightning Targets Code Review

NVIDIA released Nemotron 3.5 Lightning as an open model on August 11, a 30B-parameter mixture-of-experts design that activates only 3B parameters per token. It is aimed squarely at operational work rather than chat: code review, tool...

Read arrow_outward
Open Source·NVIDIA

Unsloth Desktop Ships Local Training App

Unsloth released Unsloth Desktop in beta on August 11, a free open-source macOS, Windows and Linux app for running and fine-tuning models on local hardware, covering GGUF, MLX, diffusion image and video, and audio models. It claims 2x...

Read arrow_outward
Consolidation·CNBC

Google Reshuffles DeepMind Leadership Under Pressure

Google moved power in its top AI ranks as it fights to regain model supremacy, with Koray Kavukcuoglu at the center of the change and fresh concerns surfacing that DeepMind's autonomy is being absorbed into Google proper. The company...

Read arrow_outward
Hardware·TechCrunch

Made By Google Loads Pixel 11 With Gemini

At Made by Google '26 on August 12, Google unveiled the Pixel 11 lineup, Pixel Watch 5, a Pixel Tag competitor to Apple's AirTag, and a batch of new Gemini features across the portfolio. The event is Google's main consumer channel for...

Read arrow_outward
The Lab

Tool of the day, field tested.

All Lab reports arrow_forward
Close screenshot Deep Dive 7.4/ 10
Voice Free tier, Pro pricing not published under an hour for a working pipeline, longer if you are migrating

Close

Most CRMs record what your reps did; Close ships an AI agent named Chloe that picks up the phone and does it.

Close is for small and mid-size sales teams who live on the phone and are tired of paying for a CRM plus a dialer plus a sequencer plus an AI SDR bolt-on. If inbound speed-to-lead is your bottleneck, the built-in dialer and Chloe make it worth a trial; if you need deep enterprise customization or a quoted price before you commit, this is a harder sell.

Capability 7.8
Ease 8.0
Value 6.4
Momentum 7.6
The Arena

Competitive intel.

Google DeepMind

Koray Kavukcuoglu takes over as DeepMind CEO; Hassabis moves to chair

OpenAI

Executive exodus continues as Kevin Weil raises $150M for his own startup

Cohere

Cohere Labs ships North Micro Vision Instruct on Hugging Face

THE DOJO · BUILD TODAY

Rebuild your eval suite around agents, not chat

01

Re-benchmark your coding agent this week. Point your harness at DeepSeek V4 Pro 0813 and score it on real repository tasks, then compare the delta against your current model at the new API price before you migrate anything.

02

Test Nemotron 3.5 Lightning on code review. The 30B mixture-of-experts activates only 3B parameters per token, so route pull request summaries and tool-calling steps to it and measure latency and cost against your general-purpose model.

03

Install Unsloth Desktop and fine-tune locally. Grab the free beta on macOS, Windows or Linux, load a GGUF or MLX checkpoint, and run one small fine-tune to see whether the claimed 70% VRAM reduction holds on your own hardware.

The Bottom Line

The benchmark moved from conversation to execution

DeepSeek's quintupled DeepSWE score and NVIDIA's code-review-focused Nemotron 3.5 Lightning point the same direction: the models that matter now are judged on whether they can finish a task in a terminal. That shift makes your evaluation harness the most valuable thing you own, because vendor-reported jumps mean nothing until they survive your repository. Meanwhile Unsloth Desktop pushes capable training down to a laptop, so the gap between hosted frontier models and what you can run yourself is narrowing at both ends. Price your agents carefully, since DeepSeek raised API rates in the same announcement that raised the bar.

Want this every morning?

AI analysis, world news, markets, and tools. One briefing, delivered free.

One email per day. No spam. Unsubscribe anytime.