K Koda Intelligence
The Lab
THE LAB · STANDARD METHOD

How a tool gets scored.

Every Lab report opens with an intake slip carrying a number out of ten and a stamped disposition. This page is where that number comes from: what each dimension means, what it is weighted, how the arithmetic rounds, and what the research behind it does and does not cover.

METHOD SHEET GENERATED 2026-08-30 FROM THE RENDERER'S OWN CONSTANTS

// 01 / SCOPE

What a report is

One tool a day, researched properly, published whether the verdict flatters the product or not. A report is a desk investigation: the product's own claims, its published pricing, a live screenshot, and what practitioners have written about using it, pulled together and judged against the same four dimensions every other tool is judged against.

It is not a hands-on long-term review, and it does not pretend to be. The slip on every report prints exactly how much research went into it: pages scraped, search passes run, community sources read. If that line is thin, weigh the verdict accordingly.

// 02 / METHOD

How the research runs

01SELECT

One tool a day is drawn from the tools that surfaced in that morning's Signal, with a fallback queue behind it. Nothing reviewed in the previous 60 days is eligible, enforced twice: a permanent ledger of everything published, and an independent scan of the report filenames on disk so a stale ledger cannot cause a repeat.

02RESOLVE

The product URL is resolved and then checked. A live domain containing the tool's name proves nothing, so the scraped page has to match the product that was described; a language model gate confirms the identity and the run retries once against a different host if it does not.

03SCRAPE

The homepage is scraped for structured facts, branding and a live screenshot, then a pricing or docs subpage is scraped for the tiers. The screenshot on every report is that capture, not an illustration.

04CANVASS

Three community searches (reviews, forum threads, alternatives) plus one live search pass for sentiment. Quotes that reach the report are paraphrased faithfully and linked to where they were found.

05SYNTHESISE

The research is handed to Claude Opus, which returns a strict JSON dossier: the four dimension scores with a rationale, the verdict, the plays, the pricing tiers, the limitations and the setup path. The dossier is schema-validated and retried on failure.

06RENDER

A deterministic Python template draws the page from that dossier. The model never writes markup, and it does not get to write the headline number either: the Koda Score is computed from its four dimensions by the formula below.

// 03 / RUBRIC

The four dimensions

Each dimension is scored 0 to 10 against these anchors. They are deliberately blunt: a score is a band with a description attached, not a feeling with a decimal point.

CAPABILITYweight x0.35What it can actually do
9 to 10

Does something no mainstream alternative does, and does it reliably enough to depend on.

7 to 8

Covers the whole job it advertises. Edges fray under load or scale, but the core holds.

5 to 6

Solves a real slice of the job. You will still reach for something else to finish.

3 to 4

Demo-grade. Impressive in the happy path, brittle the moment inputs get messy.

0 to 2

The advertised capability could not be evidenced from the product's own material or from users.

EASE OF USEweight x0.20Zero to productive
9 to 10

Productive inside ten minutes with no install, no config file and no support ticket.

7 to 8

A short setup with clear docs. One or two steps need a second read.

5 to 6

Real onboarding cost: keys, schemas or a mental model you have to build first.

3 to 4

Fights you. Undocumented steps, unclear errors, or an interface that hides the main action.

0 to 2

Blocked before first value without vendor help.

VALUEweight x0.25What you get per dollar
9 to 10

Free or priced far under what the work it replaces costs, with no metering trap.

7 to 8

Priced fairly against the alternatives. A usable free tier or an honest trial.

5 to 6

Defensible for the right user, expensive for everyone else, or the free tier is a demo.

3 to 4

Credits, seats or usage limits that make the real bill hard to predict before you commit.

0 to 2

Priced above what the output is worth, or pricing is not published at all.

MOMENTUMweight x0.20Shipping pace and traction
9 to 10

Shipping visibly and often, with an active community and a public roadmap or changelog.

7 to 8

Steady releases and answered questions. Clearly staffed and moving.

5 to 6

Alive but quiet. Updates land in months rather than weeks.

3 to 4

Little public evidence of work since launch. Community threads go unanswered.

0 to 2

Signs of abandonment: dead links, stale docs, unanswered outages.

// 04 / ARITHMETIC

The weights

Koda Score = weighted blend: capability 35, ease 20, value 25, momentum 20. Capability carries the most weight because a tool that cannot do the job is not rescued by being cheap and pleasant. Value carries the second most because the Lab exists to answer whether something is worth paying for.

SPECIMEN 0207 · Ugic AISCOREWEIGHTSHARE
Capability7.0x0.352.450
Ease of use7.5x0.201.500
Value7.5x0.251.875
Momentum5.5x0.201.100
SUM 6.925, ROUNDED HALF UP TO ONE DECIMAL6.9 / 10

Worked on a real report so the arithmetic is checkable: open specimen 0207 and add it up. Rounding is half up, not Python's default banker's rounding, so a genuine 6.95 publishes as 7.0 rather than 6.9 and the printed number always reconciles with the printed weights.

// 05 / DISPOSITION

The stamp

The stamp on the intake slip is a direct read of the Koda Score. It is not a separate opinion and there is no editorial override.

STANDOUT8.5 to 10.0
RECOMMENDED7.5 to 8.4
CONDITIONAL6.5 to 7.4
NARROW USE5.0 to 6.4
NOT RECOMMENDEDbelow 5.0
// 06 / COVERAGE

What carries a score

REPORTS FILED
207

Dated report pages in the archive, counted on disk.

CARRY A KODA SCORE
35

Reports whose four dimensions are recorded and reconcile with the weights above.

LISTED UNSCORED
172

Filed before 2026-07-19, when the rubric started. We do not backfill scores we did not compute.

// 07 / LIMITS

What this is not

  • This is desk research, not a long-term deployment. Reports are built from the product's own material, its published pricing, and what practitioners have written. They are not the output of running the tool in production for a quarter.
  • Pricing is captured on the intake date printed on the slip and moves without notice. Every report says so next to its tiers.
  • Community quotes are paraphrased and linked. The praise, mixed and critique labels are the desk's reading of the source, not the author's own label.
  • Tools are picked by the pipeline from the day's news. No vendor pays for placement, sees a report before it publishes, or reviews its own score.
Back to The Lab

One tool a day, tested properly.

The Lab lands in the morning brief. Unsubscribe anytime.