K Koda Intelligence
Subscribe
Cover image for the 10 October 2026 Signal: Just 31 of 857 Chinese AI models, or 3.6%, came with safety results traceable to a specific model, SemiAnalysis finds.
THE SIGNAL · 10 OCT 2026 · 10 MIN READ

Chinese labs published matched safety tests for just 3.6% of their models

SemiAnalysis matched safety results to only 31 of 857 Chinese models. Meanwhile Tavus, AWS and Microsoft pushed synthetic humans and agents closer to unattended work.

AI SAFETYAGENTSTOOLS
27 of 32 figures cited13 receipts in the register
  1. 48%1NetEase
  2. $3.992Sayonic: a voice assistant living in your Mac's notch
  3. 3.6%3U.S. News & World Report
The register
Lead Story

The day's defining move.

China · U.S. News & World Report
Lead storyChina 3.6%3 U.S. News & World Report · 09 OCT 2026
China·U.S. News & World Report·09 OCT 2026

Chinese Labs Publish Safety Tests for 3.6%3

SemiAnalysis reviewed 857 models released by nine major Chinese AI companies from 2021 through September 15, 2026. Only 31 of them, or 3.6%3, had published safety-evaluation results that could be matched to a specific model. U.S. News reported the findings on October 9.

Why it mattersBefore deploying a Chinese open-weight model, check whether it is among the 31 releases with matched safety results; without one, a risk review has no public eval to cite.

Continue reading
AI

News reported October 9.

Markets

The Nasdaq gained 0.64% and the S&P 500 added 0.59%, while WTI crude rose 0.19%, Bitcoin climbed 0.98% and the crypto Fear & Greed index sat at 64 (Greed).

KODA PROFrom the desk

The Operator Tier is coming.

A weekly operator deep dive, the full Dojo Pro prompt packs, and the complete prompt database. Founding members lock the launch price forever.

Join the founding list →
Listen & Watch

Daily broadcasts.

Podcast · Video · Infographic
YouTube · Short

Today's Signal Short

Visual

Intelligence map

Markets

6 levels at the close.

Market TerminalLAST CLOSE 09 OCT 2026 · READ 2026-10-10 · Yahoo Finance
S&P 500 7,811.547 +0.59%7 6D, window short 7,722.72 low / 7,818.93 high
Nasdaq 27,366.178 +0.64%8 6D, window short 27,190.86 low / 27,599.79 high
Bitcoin $82,475.969 +0.98%9 6D, window short $81,676.34 low / $86,480.30 high
Ethereum $2,486.6410 +0.60%10 6D, window short $2,471.91 low / $2,726.51 high
Crude Oil (WTI) $91.6611 +0.19%11 6D, window short $88.28 low / $91.66 high
Crypto Fear & Greed 64 / 10012 Greed 0 to 100 index 25 fear / 50 neutral / 75 greed · alternative.me
Today's Focus

What else the day turns on.

01

Anthropic Launches Claude Dashboards, Motion

Anthropic launched two Claude features aimed at everyday business output. Claude Dashboards builds auto-refreshing dashboards from a team's existing data sources, and Claude Motion turns ideas, decks and proposals into short explainer videos and animations. Anthropic positions both as complements to current software.

02

Microsoft Adds 30+ Copilot Skills

Microsoft adds 30+ Copilot skills to Dynamics 365, spanning Sales, Service and Customer Insights and working across Chat, Cowork, Autopilot and Code. The companion Copilot Managed Runtime hosts code inside a company's own environment under IT governance, a pitch aimed at keeping control with corporate IT.

03

Apple HomePad Delayed for Siri

Apple opens its fall hardware cycle with an Oct. 13 smart-home event in New York, teased as "Welcome home," where it is expected to show the long-rumored HomePad display hub, a new HomePod mini and an updated Apple TV. Apple reportedly delayed the HomePad until its Siri AI was ready, tying the device's timing to the assistant.

The Lab

Tool of the day, researched.

All Lab reports
Sayonic home page, captured for the 10 October 2026 Lab report Tool of the day 6.213/ 10
Voice

Sayonic

Sayonic parks a small voice pill in your Mac's notch, and when you talk to it, it does the task itself, down to handing work to the coding agent already on your machine.

Sayonic's pricing page lists a $3.992 per Mac BYOK plan with every feature, so a voice layer that opens apps and tabs, files notes and calendar events, and briefs your coding agent costs little to try if you already pay for a model. Mac users who live in Finder searches and constant app switching will get the most from it, and developers can skip Cloud credits entirely. Give it a week of your own commands before committing to a yearly plan.

Capability 6.5
Ease 7.0
Value 7.0
Momentum 4.0

An editorial verdict calls Sayonic a practical, privacy-conscious way to control a Mac by voice or text from a subtle on-screen pill, rating it 8 out of 10.Paraphrased · zekaiwork.com

Also on the radar
Creativity

Spoke: transcribe, summarize, and edit video calls as text

Spoke transcribes and summarizes video conversations in real time. You edit the video by editing its transcript, then share a subtitled summary clip. It uses a freemium model, so you can try it before paying. One practical use: cut a 60-minute client call down to the three decisions that matter and send that clip in place of meeting notes.

Try it
Build

Rocket.new: turn an app idea into a working build without code

Rocket.new is an AI app builder that turns a plain-language idea into a production-ready application, with no coding required. Start with a narrow internal tool, such as an intake form that feeds a simple dashboard, and write out the data fields and user roles before you prompt. Then judge the output on two things: whether you can edit it after generation, and how you export or host it.

Try it
Productivity

IrisGo: show a workflow once, then let AI repeat it

IrisGo is built for solopreneurs. It watches you complete a process one time, then reruns it as an AI-powered workflow. It fits lead management, customer follow-up, and repetitive admin work. Record a task you do at least weekly, such as logging new leads and sending a first-touch email, then check the first few automated runs before you let it run unattended.

Try it
Coding

Rill Browser: a web browser shared with your coding agents

Rill Browser is designed so coding agents such as Claude Code and Codex can work in the browser next to you. That cuts the copy-paste between research tabs and your terminal. Use it for tasks that mix browsing and building, for example letting an agent read API docs and draft integration code while you review each page it visits.

Try it
The Arena

3 labs moved this week.

  1. Google DeepMindGoogle opens a private enterprise preview of a universal Gemini agent that runs multi-step work across apps and devices mlq.ai · 09 OCT 202601
  2. MistralArtificial Analysis ranks Mistral Large 4 the strongest model built outside the U.S. and China tomshardware.com · 09 OCT 202602
  3. Meta AIMeta joins nine other AI developers in data-protection commitments secured by the U.K. Information Commissioner's Office ghacks.net · 09 OCT 202603
Builder Radar

1 receipt from the repos, one paper.

Primary sources, pulled by index
Receipts
  1. Behind: Microsoft Launches Fast Routing and Classification Model DocsRouting

    Training Module Route Work Efficiently with Unified Routing in Dynamics 365 Customer Service - Training Transform how customer requests reach the right people by configuring intelligent classification, queue routing, and AI-assisted work...

    learn.microsoft.com
Paper of the day

Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

Lasse B. Strand, Robert Jakob, Kevin O'Sullivan et al. · arXiv 2026-10-06

Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation.

In the Accuracy experiment, which isolates the search from cost, the agent is sample-efficient: its best configuration within the first 10 trials already matches or beats every statistical baseline’s full 30-trial held-out judge accuracy, at 0.817 against...
Read the paper

Receipts come from the Firecrawl Developer Index (READMEs, docs, issues, merged PRs), first-party sources only. The paper comes from the Firecrawl Research Index: arXiv, indexed in the last 7 days and submitted this month or last, picked by today's research theme. Quotes are verbatim.

THE DOJO · BUILD TODAY

Put guardrails around agents before they run alone

  1. 01

    Run one agent inside Strands Box. Take an agent you already let run unattended and move it into AWS's open-source sandbox, then note which actions it tried to take outside its permissions.

  2. 02

    Add an out-of-band check to video approvals. Since Griffin-Lite passed as human for 48% of testers, require a callback to a known number before acting on any payment or access change requested on a call.

  3. 03

    Try Sayonic on one coding task. Install the $3.99 BYOK plan with your existing model key and use voice to brief your coding agent for a day, then keep it only if it saves you typing.

The Bottom Line

Trust now has to be checked, because models and faces no longer vouch for themselves.

SemiAnalysis shows most Chinese releases arrive without matched safety results, and Tavus shows a video caller can pass as human about half the time, so the evidence builders once took for granted is thin at both ends. AWS's Strands Box and Microsoft's sandboxed tools give you a place to contain what an agent does when nobody is watching. Move your riskiest agent into a sandbox this week. Then require a second channel for any request that moves money.

Receipts

Every number, and where it came from.

27 of 32 figures cited, 13 receipts: 3 verified, 6 reported, 4 estimated
  1. 0148%Tavus Griffin-Lite Passes as Human for 48%.NetEase09 OCT 202603B search corroborationV
  2. 02$3.99Plans start at $3.99/mo per Mac, with Pro and Max tiers above that.Sayonic: a voice assistant living in your Mac's notch10 OCT 202603B search corroborationE
  3. 033.6%Chinese Labs Publish Safety Tests for 3.6%.U.S. News & World Report09 OCT 202603B search corroborationV
  4. 042.4%Tavus says its previous real-time video system fooled 2.4% of testers into thinking it was human; Griffin-Lite...NetEase09 OCT 202603B search corroborationV
  5. 058x...and ChatGPT Work, claiming near-Astra intelligence at up to 8x the speed of Sol StandardOpenAI, via TLDR09 OCT 2026as publishedE
  6. 0675%Anthropic releases Claude Haiku 5.5, roughly 75% cheaper to run than Haiku 4.5Anthropic, via TLDR09 OCT 2026as publishedE
  7. 077,811.54S&P 500 close, +0.59% on the sessionYahoo Finance09 OCT 2026yfinance closeR
  8. 0827,366.17Nasdaq close, +0.64% on the sessionYahoo Finance09 OCT 2026yfinance closeR
  9. 09$82,475.96Bitcoin close, +0.98% on the sessionYahoo Finance10 OCT 2026yfinance closeR
  10. 10$2,486.64Ethereum close, +0.60% on the sessionYahoo Finance10 OCT 2026yfinance closeR
  11. 11$91.66Crude Oil (WTI) close, +0.19% on the sessionYahoo Finance09 OCT 2026yfinance closeR
  12. 1264 / 100Crypto Fear & Greed close, Greed on the sessionalternative.me10 OCT 2026as publishedR
  13. 136.2 / 10Sayonic, Koda score across four dimensions04R dossier10 OCT 202604R dossier scoringE
V
Verified: the stat gate corroborated this figure against an independent search before publication.
R
Reported: the figure is carried as published by the linked source and was not independently corroborated.
E
Estimated: the figure is approximate or hedged, either in the source or by the stat gate.

Method names the computation path behind the row, from a closed vocabulary: as published, yfinance close, 03B search corroboration, 04R dossier scoring, not recorded.

A figure is money, a percentage, a spelled magnitude, a multiple, a rate carrying its unit, a score over its denominator, a thousands separated number or an index close. Bare integers and years are not counted, so a phrase like "over 100 integrations" carries no mark. A chart's own scale line is the axis of the level above it and is counted once, with that level.

Get the morning Signal

197 editions so far, one a day. Unsubscribe anytime.

Forward this to one operator you work with. Your referral link is in every email; milestones at koda.community/refer.