Everything this lab has looked into, with its status on the row rather than in a footnote. Published means a result with a source and a date. In progress means it is running and has produced nothing you can cite. Illustrative means arithmetic on public prices or a dated snapshot — informative, but not a measurement of ours.
Our Truth Auditor gated agent output for three months and reported 230 successful runs at 75/100. It was failing open: the same score came back on failed runs, and 150 runs were never verified at all. We paused the workforce and published the whole finding.
Dated 2026-08-08 · Source: founder_self_reported — internal audit of the workforce run log
Every reviewer decision on Keenoble is logged with its outcome — accepted first try, accepted after revision, or escalated to a human. The live totals render on the homepage from the public API, and the block shows nothing at all rather than a placeholder when that API is unreachable.
Dated 2026-08-08 · Source: ai.keenlabs.pro/api/public/reviewers/outcomes (live, 60s cache)
The measurement that turns our routing claim into a published result, and the one this site spent months saying it did not have. Every model in the v1 roster, across eleven task domains: one task set, identical conditions, graded by a deterministic check wherever one applies and a blinded jury only where none can. Pass rate, cost per task, latency and failure modes are published per model, each next to the number of tasks actually behind it — because the depth is uneven and that weakens the comparison. Directional, not statistical.
Dated 2026-08-17 · Source: internal qualification sweep, suite qualification-v1 — data/findings-public.json, exported verbatim from the eval harness
Three models in the sweep begin their visible answer with an internal reasoning channel on more than nine attempts in ten. They hold the three lowest pass rates in the benchmark, and that is a format failure rather than a capability one — the answer is frequently present, wrapped in thinking our grader is strict about. Anything consuming these models has to strip the channel before parsing.
Dated 2026-08-17 · Source: internal qualification sweep — data/findings-public.json, findings.reasoning_channel_leak
Two models returned 200 with an empty body and a refusal stop reason, carrying no refusal text at all, on benign work including "fix this function". Reproduced directly against the provider API, so it is provider behaviour rather than an artifact of our harness. Both are models we route to, and we are publishing it about them anyway: a caller that reads an empty 200 as an empty answer will mis-handle these.
Dated 2026-08-17 · Source: internal qualification sweep — data/findings-public.json, findings.silent_refusals
145 real answers out of the production cost log, 2026-08-01 to 2026-08-18: task class, the model the router actually chose, measured latency, and the Sparks the user was actually charged. The row shown per class is the median by latency, and the number of rows behind it is published beside it because the depth is uneven. Still illustrative rather than a benchmark — every row is a different request, so this shows the method working, not models compared.
Dated 2026-08-18 · Source: production command_llm_cost_log join ai_runtime_events on request_id, status=ok, excluding estimated / test-account / internal rows; median-latency row per task class
Every priced model in the qualification sweep ran one task set, so the cost per task can be laid side by side without a list price anywhere in it. The library spans three orders of magnitude end to end, and nearly two when the comparison is narrowed to only the models that ran the full task set. Price does not predict the score in either direction: the cheapest model to clear a 90% pass rate did it on the full suite, and the dearest model that ran the full suite holds the lowest pass rate in the whole sweep. That gap, per prompt, is what the router is for. Depth is uneven across the runs and is published beside every figure.
Dated 2026-08-17 · Source: internal qualification sweep, suite qualification-v1 — data/findings-public.json, models[].usd_per_task
Lay the providers’ own published list prices side by side and the shape is unarguable — a greeting routed to a frontier model costs about a hundred times what a greeting is worth. This is arithmetic on public pricing, not a measurement of ours, and it is no longer the argument: the measured version of the same spread now leads /pricing. What this table still does is convert the gap into Sparks, the unit a customer is actually billed in, which the sweep cannot do because the sweep measured models rather than our billing.
Dated 2026-08-19 · Source: the providers’ own public list prices — labelled illustrative wherever it is shown
The snapshot below is the illustrative row above, rendered: what the router chose on a set of real prompts, dated, with what each answer cost.
MEASURED · 2026-08-01 to 2026-08-18
| Task | Model | Latency | Charged |
|---|
| Code n=20 | llama-3.3-70b-versatile | 376ms | 0 sp |
| Everyday question n=72 | deepseek-v4-flash | 936ms | 0.01 sp |
| Quick question n=48 | deepseek-v4-flash | 998ms | 0 sp |
| Research n=1 | llama-3.3-70b-versatile | 1.5s | 0.5 sp |
| Long analysis n=2 | llama-3.3-70b-versatile | 2.6s | 0.2 sp |
| Regulated guidance n=2 | gpt-4o-mini | 3.4s | 0.8 sp |
145 real answers from the production cost log, 2026-08-01 to 2026-08-18. One median-latency row per class, with n beside it — one row is an anecdote with a receipt, not a result. Rows charged 0 sp were genuinely free. Illustrative rather than a benchmark: every row is a different request, and the same-test comparison is above.