Every Lab report opens with an intake slip carrying a number out of ten and a stamped disposition. This page is where that number comes from: what each dimension means, what it is weighted, how the arithmetic rounds, and what the research behind it does and does not cover.
METHOD SHEET GENERATED 2026-08-30 FROM THE RENDERER'S OWN CONSTANTS
One tool a day, researched properly, published whether the verdict flatters the product or not. A report is a desk investigation: the product's own claims, its published pricing, a live screenshot, and what practitioners have written about using it, pulled together and judged against the same four dimensions every other tool is judged against.
It is not a hands-on long-term review, and it does not pretend to be. The slip on every report prints exactly how much research went into it: pages scraped, search passes run, community sources read. If that line is thin, weigh the verdict accordingly.
One tool a day is drawn from the tools that surfaced in that morning's Signal, with a fallback queue behind it. Nothing reviewed in the previous 60 days is eligible, enforced twice: a permanent ledger of everything published, and an independent scan of the report filenames on disk so a stale ledger cannot cause a repeat.
The product URL is resolved and then checked. A live domain containing the tool's name proves nothing, so the scraped page has to match the product that was described; a language model gate confirms the identity and the run retries once against a different host if it does not.
The homepage is scraped for structured facts, branding and a live screenshot, then a pricing or docs subpage is scraped for the tiers. The screenshot on every report is that capture, not an illustration.
Three community searches (reviews, forum threads, alternatives) plus one live search pass for sentiment. Quotes that reach the report are paraphrased faithfully and linked to where they were found.
The research is handed to Claude Opus, which returns a strict JSON dossier: the four dimension scores with a rationale, the verdict, the plays, the pricing tiers, the limitations and the setup path. The dossier is schema-validated and retried on failure.
A deterministic Python template draws the page from that dossier. The model never writes markup, and it does not get to write the headline number either: the Koda Score is computed from its four dimensions by the formula below.
Each dimension is scored 0 to 10 against these anchors. They are deliberately blunt: a score is a band with a description attached, not a feeling with a decimal point.
Does something no mainstream alternative does, and does it reliably enough to depend on.
Covers the whole job it advertises. Edges fray under load or scale, but the core holds.
Solves a real slice of the job. You will still reach for something else to finish.
Demo-grade. Impressive in the happy path, brittle the moment inputs get messy.
The advertised capability could not be evidenced from the product's own material or from users.
Productive inside ten minutes with no install, no config file and no support ticket.
A short setup with clear docs. One or two steps need a second read.
Real onboarding cost: keys, schemas or a mental model you have to build first.
Fights you. Undocumented steps, unclear errors, or an interface that hides the main action.
Blocked before first value without vendor help.
Free or priced far under what the work it replaces costs, with no metering trap.
Priced fairly against the alternatives. A usable free tier or an honest trial.
Defensible for the right user, expensive for everyone else, or the free tier is a demo.
Credits, seats or usage limits that make the real bill hard to predict before you commit.
Priced above what the output is worth, or pricing is not published at all.
Shipping visibly and often, with an active community and a public roadmap or changelog.
Steady releases and answered questions. Clearly staffed and moving.
Alive but quiet. Updates land in months rather than weeks.
Little public evidence of work since launch. Community threads go unanswered.
Signs of abandonment: dead links, stale docs, unanswered outages.
Koda Score = weighted blend: capability 35, ease 20, value 25, momentum 20. Capability carries the most weight because a tool that cannot do the job is not rescued by being cheap and pleasant. Value carries the second most because the Lab exists to answer whether something is worth paying for.
Worked on a real report so the arithmetic is checkable: open specimen 0207 and add it up. Rounding is half up, not Python's default banker's rounding, so a genuine 6.95 publishes as 7.0 rather than 6.9 and the printed number always reconciles with the printed weights.
The stamp on the intake slip is a direct read of the Koda Score. It is not a separate opinion and there is no editorial override.
Dated report pages in the archive, counted on disk.
Reports whose four dimensions are recorded and reconcile with the weights above.
Filed before 2026-07-19, when the rubric started. We do not backfill scores we did not compute.
The Lab lands in the morning brief. Unsubscribe anytime.