Lab status: workforce rebuilding on measured foundations — qualification program runningread the audit →
KEEN LABS

Research

How we measure. What we've found.

The first qualification sweep has published. 54 models, 11 task domains, 2,746 scored tasks — same tasks, same conditions, same grader, measured 2026-08-17. Pass rates, cost per task, latency and the failure modes are all below, with the limits of the measurement stated at the same weight as the numbers.

What you will not find here is our routing configuration. We publish findings and method; the config is the part we keep. That line has not moved because the benchmark landed.

Smart fusion routing — the method

01

Catalogue

Every model we can reach gets added to the catalogue — 335 of them, counted, not rounded, as of 2026-08-08. Growing weekly.

02

Measure

Real tasks, not synthetic benchmarks: quality, cost per answer, speed, reliability. Every answer logs model, cost, latency, outcome.

03

Route

The router serves whatever wins today — 12+ models in active daily rotation, promoted and demoted by data, not by brand.

04

Re-measure

The loop repeats. Today humans run it and gate every change; the paused workforce is being rebuilt to run it continuously, with approval gates.

See the dated method example below, or the same table in context on the homepage's routing snapshot. That table is production traffic, not a benchmark — the head-to-head is the results section below.

Measured results · same test, every model

The head-to-head benchmark, published.

One task set, 11 domains, every model run under identical conditions and graded the same way. 54 models, 2,746 scored tasks. Costs are computed from measured token counts at published provider rates. Latency is wall-clock, from our infrastructure.

This is a benchmark of 54 models, not of the whole 335-model catalogue. The catalogue figure counts models we can reach and price; this table counts models we have run the same tasks against. The two are not the same claim and the gap is the work still to do.

Full suite

17 models

Ran the whole task set — around 120 scored tasks each. These are the rows with enough behind them to compare.

ModelTasks scoredPass$ / taskLatency p50p95Reasoning overheadTruncated
gpt-5-miniOpenAI12494.4%$0.002967893ms69.0s31%0%
gpt-5.6-lunaOpenAI12491.9%$0.0008862033ms15.6s8%0%
gemini-3.1-flash-liteGoogle12491.1%$0.0007141034ms3982ms0%4%
gemini-flash-lite-latestGoogle12488.7%$0.00127859ms5028ms0%4%
gpt-5-nanoOpenAI12488.7%$0.0010414.4s61.6s62%0%
gemini-3.5-flash-liteGoogle12487.1%$0.00125931ms4958ms0%5%
gpt-4.1-miniOpenAI12487.1%$0.0008001565ms10.8s0%1%
gpt-5.4-nanoOpenAI12485.5%$0.0006661142ms11.4s0%4%
deepseek-v4-flashDeepSeek12480.6%$0.000527360ms751ms19%7%
deepseek-v4-proDeepSeek12380.5%$0.00183371ms841ms27%4%
gpt-4o-miniOpenAI12479.8%$0.0002931988ms9395ms0%0%
gpt-4.1-nanoOpenAI12479.0%$0.0001911296ms6125ms0%0%
llama-3.3-70b-versatileGroq9575.8%$0.0004981024ms32.6s0%1%
llama-3.1-8b-instantGroq11464.0%$0.000043562ms14.3s0%0%
gemma-4-31b-itGoogle11842.4%free tier18.0s56.9s0%11%
gemma-4-26b-a4b-itGoogle11731.6%free tier15.9s43.4s0%23%
qwen/qwen3.6-27bGroq11515.7%$0.003322601ms15.8s0%36%

Partial suite

1 model

Stopped short of the full set for a reason on our side, not the model's — an account output cap in this case. Read the pass rate with the observation count.

ModelTasks scoredPass$ / taskLatency p50p95Reasoning overheadTruncated
openai/gpt-oss-120bGroq4671.7%$0.0004992198ms13.4s0%0%

Reduced set

26 models

Around 21 scored tasks each — every domain covered, little repetition. Directional.

ModelTasks scoredPass$ / taskLatency p50p95Reasoning overheadTruncated
o4-miniOpenAI21100.0%$0.004854823ms16.9s14%0%
gpt-4.1OpenAI2195.2%$0.004291211ms8092ms0%0%
gpt-5OpenAI2195.2%$0.021316.6s102.2s43%0%
gpt-5.4-miniOpenAI2195.2%$0.002341170ms10.4s0%0%
kimi-k3Moonshot2095.0%$0.021212.1s350.2s29%0%
claude-sonnet-4-6Anthropic2190.5%$0.009874702ms25.1s0%0%
gpt-5.2OpenAI2190.5%$0.007881771ms23.2s0%0%
gpt-5.4OpenAI2190.5%$0.007771498ms15.0s0%0%
gpt-5.6-terraOpenAI2190.5%$0.01361675ms33.6s14%0%
grok-4.3xAI2190.5%$0.001873598ms15.8s0%0%
grok-4.5xAI2190.5%$0.004576625ms25.7s0%0%
kimi-k2.7-code-highspeedMoonshot2190.5%$0.007642943ms8712ms14%5%
claude-sonnet-4-5-20250929Anthropic2185.7%$0.007793890ms21.6s0%0%
gpt-5.1OpenAI2185.7%$0.005361989ms16.6s0%0%
grok-4.6xAI2185.7%$0.004418219ms49.2s0%5%
kimi-k2.7-codeMoonshot2185.7%$0.0042815.5s177.9s14%5%
o3OpenAI2185.7%$0.008353958ms17.3s19%0%
kimi-k2.6Moonshot2085.0%$0.0066928.8s290.0s38%5%
claude-haiku-4-5-20251001Anthropic2181.0%$0.002731832ms9530ms0%0%
claude-sonnet-5Anthropic2181.0%$0.008153342ms16.2s0%0%
sonar-proPerplexity2176.2%$0.009092987ms15.0s0%0%
gemini-3.5-flashGoogle2161.9%$0.003603492ms8782ms0%29%
gemini-3.6-flashGoogle2161.9%$0.003064442ms11.4s0%24%
gemini-flash-latestGoogle2161.9%$0.003662205ms7286ms0%24%
gemini-3.1-pro-previewGoogle2147.6%$0.003727415ms17.6s0%24%
gemini-pro-latestGoogle2147.6%$0.003827784ms16.9s0%33%

Core spine only

10 models

One byte-identical task per domain, 11 in total. Enough to say a model can do a thing at all; nowhere near enough to rank it. A perfect score here is 11 for 11, and that is all it is.

ModelTasks scoredPass$ / taskLatency p50p95Reasoning overheadTruncated
chat-latestOpenAI11100.0%$0.01611466ms16.6s0%0%
claude-opus-4-7Anthropic11100.0%$0.02013276ms29.4s0%0%
gpt-5.6-solOpenAI11100.0%$0.03913881ms102.2s9%0%
claude-opus-4-5-20251101Anthropic1190.9%$0.01653587ms26.9s0%0%
claude-opus-4-6Anthropic1190.9%$0.02065042ms47.1s0%9%
claude-opus-4-8Anthropic1190.9%$0.02253959ms32.0s0%9%
gpt-5.5OpenAI1190.9%$0.02182931ms31.5s9%0%
gpt-5-search-apiOpenAI1181.8%$0.01683446ms83.0s0%0%
claude-fable-5Anthropic1172.7%$0.04364623ms32.0s0%0%
claude-opus-5Anthropic1163.6%$0.03255019ms62.2s9%9%

Grading. Deterministic checks decide a task wherever one applies and are authoritative: executed code, parsed JSON, matched values, detected language, enforced length. Only open-ended reasoning reaches a jury.

Reasoning overhead is the share of attempts where a model spent its entire output budget on hidden reasoning and returned nothing visible, and had to be re-asked with a larger budget. The wasted attempt is billed, so it is a real per-call cost rather than a quality score. Truncated is the share of attempts cut off by our own output cap — a limit of ours, published so the pass rate can be read with it in mind.

Snapshot 2026-08-17 · suite qualification-v1 · export schema v1 · source data/findings-public.json, committed verbatim as the eval harness emitted it.

Failure findings

The most useful things we found are the ways models fail.

A pass rate tells you how often something worked. These tell you what happens when it does not — which is the part you have to engineer around. Two of the three are about models we route to, and the third is a limitation of our own harness.

Three models leak their reasoning channel into visible output

Three models emit their internal reasoning channel into visible output, so their low scores are a FORMAT failure rather than a capability failure. Any consumer of these models needs to strip the channel before parsing.

gemma-4-26b-a4b-itleaked on 95% of attemptsmeasured pass rate 31.6%
gemma-4-31b-itleaked on 95% of attemptsmeasured pass rate 42.4%
qwen/qwen3.6-27bleaked on 93% of attemptsmeasured pass rate 15.7%

Read their rows in the table above with this in mind. Those three pass rates are the lowest in the sweep and they are not a statement about whether the model can reason — the answer is often in the output, with the thinking wrapped around it, and our grader is strict about format on purpose.

Two models return an empty 200 and no refusal text

Two models returned HTTP 200 with an empty body and a refusal stop reason, carrying no refusal text at all, on benign tasks including "fix this function". Reproduced directly against the provider API, so it is provider behaviour and not a harness artifact. A caller that treats an empty 200 as an empty answer will mis-handle these.

claude-fable-5on 3 domainsD2_repair_fn · D3_doc_analysis · D6_long_context
claude-opus-5on 2 domainsD3_doc_analysis · D6_long_context

Both are models we route to, and we are publishing this about them anyway. It was reproduced directly against the provider API rather than inferred from our own logs, which is the only reason it is stated as provider behaviour.

Reasoning models can spend an entire output budget and return nothing

Several models spend an entire small output budget on hidden reasoning and return no visible text. Given four times the budget most then answer normally. Budget this as a real per-call cost: the wasted attempt is billed.

gpt-5-nano62% of attempts
gpt-543% of attempts
kimi-k2.638% of attempts
gpt-5-mini31% of attempts
kimi-k329% of attempts
deepseek-v4-pro27% of attempts
o319% of attempts
deepseek-v4-flash19% of attempts

Every failure, classified

602attempts failed and were classified. The taxonomy was written before the run, not fitted to it, and three of these classes are our fault rather than the model's — those are marked, and they are excluded from the quality figures rather than counted against the model.

412Format break — output unparseable, or cut off by our cap
75Tool failure — wrong tool, wrong order, or wrong arguments
64Rate-limit collapse — provider tokens-per-minute exhausted (ours) · ours
23Empty content — a 200 response with nothing in it
15Language mismatch — answered in the wrong language
5Refusal — the model declined the task
3Timeout
3Judge failure — the blinded grader errored (ours) · ours
2Config error — a bad request of our own making (ours) · ours

Findings catalogue

6 published. 2 that aren't, listed anyway.

Everything this lab has looked into, with its status on the row rather than in a footnote. Published means a result with a source and a date. In progress means it is running and has produced nothing you can cite. Illustrative means arithmetic on public prices or a dated snapshot — informative, but not a measurement of ours.

  • Our Truth Auditor gated agent output for three months and reported 230 successful runs at 75/100. It was failing open: the same score came back on failed runs, and 150 runs were never verified at all. We paused the workforce and published the whole finding.

    Dated 2026-08-08 · Source: founder_self_reported — internal audit of the workforce run log

  • Every reviewer decision on Keenoble is logged with its outcome — accepted first try, accepted after revision, or escalated to a human. The live totals render on the homepage from the public API, and the block shows nothing at all rather than a placeholder when that API is unreachable.

    Dated 2026-08-08 · Source: ai.keenlabs.pro/api/public/reviewers/outcomes (live, 60s cache)

  • The measurement that turns our routing claim into a published result, and the one this site spent months saying it did not have. Every model in the v1 roster, across eleven task domains: one task set, identical conditions, graded by a deterministic check wherever one applies and a blinded jury only where none can. Pass rate, cost per task, latency and failure modes are published per model, each next to the number of tasks actually behind it — because the depth is uneven and that weakens the comparison. Directional, not statistical.

    Dated 2026-08-17 · Source: internal qualification sweep, suite qualification-v1 — data/findings-public.json, exported verbatim from the eval harness

  • Three models in the sweep begin their visible answer with an internal reasoning channel on more than nine attempts in ten. They hold the three lowest pass rates in the benchmark, and that is a format failure rather than a capability one — the answer is frequently present, wrapped in thinking our grader is strict about. Anything consuming these models has to strip the channel before parsing.

    Dated 2026-08-17 · Source: internal qualification sweep — data/findings-public.json, findings.reasoning_channel_leak

  • Two models returned 200 with an empty body and a refusal stop reason, carrying no refusal text at all, on benign work including "fix this function". Reproduced directly against the provider API, so it is provider behaviour rather than an artifact of our harness. Both are models we route to, and we are publishing it about them anyway: a caller that reads an empty 200 as an empty answer will mis-handle these.

    Dated 2026-08-17 · Source: internal qualification sweep — data/findings-public.json, findings.silent_refusals

  • 145 real answers out of the production cost log, 2026-08-01 to 2026-08-18: task class, the model the router actually chose, measured latency, and the Sparks the user was actually charged. The row shown per class is the median by latency, and the number of rows behind it is published beside it because the depth is uneven. Still illustrative rather than a benchmark — every row is a different request, so this shows the method working, not models compared.

    Dated 2026-08-18 · Source: production command_llm_cost_log join ai_runtime_events on request_id, status=ok, excluding estimated / test-account / internal rows; median-latency row per task class

  • Every priced model in the qualification sweep ran one task set, so the cost per task can be laid side by side without a list price anywhere in it. The library spans three orders of magnitude end to end, and nearly two when the comparison is narrowed to only the models that ran the full task set. Price does not predict the score in either direction: the cheapest model to clear a 90% pass rate did it on the full suite, and the dearest model that ran the full suite holds the lowest pass rate in the whole sweep. That gap, per prompt, is what the router is for. Depth is uneven across the runs and is published beside every figure.

    Dated 2026-08-17 · Source: internal qualification sweep, suite qualification-v1 — data/findings-public.json, models[].usd_per_task

  • Lay the providers’ own published list prices side by side and the shape is unarguable — a greeting routed to a frontier model costs about a hundred times what a greeting is worth. This is arithmetic on public pricing, not a measurement of ours, and it is no longer the argument: the measured version of the same spread now leads /pricing. What this table still does is convert the gap into Sparks, the unit a customer is actually billed in, which the sweep cannot do because the sweep measured models rather than our billing.

    Dated 2026-08-19 · Source: the providers’ own public list prices — labelled illustrative wherever it is shown

The snapshot below is the illustrative row above, rendered: what the router chose on a set of real prompts, dated, with what each answer cost.

MEASURED · 2026-08-01 to 2026-08-18

TaskModelLatencyCharged
Code n=20llama-3.3-70b-versatile376ms0 sp
Everyday question n=72deepseek-v4-flash936ms0.01 sp
Quick question n=48deepseek-v4-flash998ms0 sp
Research n=1llama-3.3-70b-versatile1.5s0.5 sp
Long analysis n=2llama-3.3-70b-versatile2.6s0.2 sp
Regulated guidance n=2gpt-4o-mini3.4s0.8 sp

145 real answers from the production cost log, 2026-08-01 to 2026-08-18. One median-latency row per class, with n beside it — one row is an anecdote with a receipt, not a result. Rows charged 0 sp were genuinely free. Illustrative rather than a benchmark: every row is a different request, and the same-test comparison is above.

Skill cards

11 skill cards, one per task domain.

A card is one task domain with every measured model run against the same task set. 11 cards, published 2026-08-17.

No card names a winner, and that is a finding rather than an omission. On 10 of the 11 domains the median model simply passes, with between ten and forty-four models tied on a perfect score and at most twelve tasks behind any of those scores. Naming one of forty-four tied models the best at a skill would be a ranking the measurement cannot support. What each card gives instead is the shape of the field.

The useful conclusion is the one the saturation points at: for most of this work, the question is not which model can do it. It is what each one costs, how long it takes, and how it fails — which is what the results table and the failure findings are for.

D1_build_pageseparated the field

Build page

Build a working single-file web page from a plain-language brief.

models measured
54
tasks scored
277
median pass rate
50%
passed everything given
10 of 54
deepest perfect run
2 tasks
D2_repair_fn

Repair fn

Fix a defective function; the fix is executed against tests the model never saw.

models measured
54
tasks scored
279
median pass rate
100%
passed everything given
41 of 54
deepest perfect run
12 tasks
D3_doc_analysis

Doc analysis

Answer quantitative questions from a long operational document.

models measured
54
tasks scored
261
median pass rate
100%
passed everything given
31 of 54
deepest perfect run
11 tasks
D4_tool_multistep

Tool multistep

Use tools over several turns and answer from what they returned.

models measured
54
tasks scored
221
median pass rate
100%
passed everything given
36 of 54
deepest perfect run
9 tasks
D5_multilingual

Multilingual

Answer in the language the user wrote in.

models measured
54
tasks scored
267
median pass rate
100%
passed everything given
35 of 54
deepest perfect run
12 tasks
D6_long_context

Long context

Find and combine two facts far apart in a very long document.

models measured
50
tasks scored
151
median pass rate
100%
passed everything given
40 of 50
deepest perfect run
9 tasks
D7_json_format

Json format

Return exactly the requested JSON structure and nothing else.

models measured
54
tasks scored
267
median pass rate
100%
passed everything given
41 of 54
deepest perfect run
12 tasks
D8_math

Math

Arithmetic word problems with a single checkable answer.

models measured
53
tasks scored
266
median pass rate
100%
passed everything given
44 of 53
deepest perfect run
12 tasks
D9_summarize

Summarize

Faithful summary under a hard length limit.

models measured
54
tasks scored
240
median pass rate
100%
passed everything given
28 of 54
deepest perfect run
11 tasks
D10_reasoning

Reasoning

Open-ended business reasoning, graded by a blinded jury.

models measured
53
tasks scored
252
median pass rate
100%
passed everything given
35 of 53
deepest perfect run
12 tasks
D11_speed

Speed

Short factual answers where latency is the point.

models measured
53
tasks scored
265
median pass rate
100%
passed everything given
41 of 53
deepest perfect run
12 tasks

Grader agreement

One domain is graded by judgement. Here is how much the judges agreed.

Ten of the eleven domains are decided by a deterministic check — code that runs, JSON that parses, a number that matches, a language that is detected. Only open-ended reasoning prose reaches a blinded jury, and a jury is only worth what its consistency is worth.

75.9%exact agreementboth graders gave the identical score
89.7%within one pointscores differed by no more than 1
0.483mean absolute differenceaverage gap between the two scores
29double-graded itemsthe whole sample these three rest on

Scope. Measured on open-ended reasoning prose only — the one domain with no deterministic check. It does NOT generalise to the other ten domains, which never reach a jury. Graders are blinded to model identity and are never from the graded model's own provider family.

These are raw agreement rates, not Cohen's κ. κ additionally discounts the agreement two graders would reach by chance and is always the lower number; we have not computed it, so we do not quote it. And 29 double-graded items is a small sample — small enough that these figures indicate the grading was not arbitrary, and not enough to put a confidence interval on.

Limits of this measurement

What this benchmark does not establish.

This benchmark is directional, not statistical. It is one run, with uneven depth, on tasks we wrote. It is enough to tell you what a model costs, roughly how fast it is, and how it fails. It is not enough to certify one model as better than another, and nothing on this page should be read as though it were.

  1. Depth is uneven. Some models ran the full suite and others a fixed core subset, so a difference between a high-observation and a low-observation model is weaker evidence than the two pass rates alone suggest. Observation counts are published per model and per domain so this is checkable rather than hidden.
  2. Output caps truncated some answers. A truncated answer can never be scored as a pass, so a model that writes long is penalised by our limit, not only by its own ability. The truncation rate is published per model.
  3. Six tasks were withdrawn from scoring entirely because their expected values had no settled reading, and are excluded from every figure here.
  4. Calls that failed for reasons on our side — rate limits, timeouts, our own configuration — are excluded from quality figures rather than counted against the model.
  5. Costs are computed at published provider rates from measured token counts. They are not an invoice and they do not include negotiated or free-tier pricing.
  6. One benchmark run is a snapshot. Model behaviour changes without notice, and two of the findings above are provider behaviours that a version bump could alter either way.

6 stated limits, rendered verbatim from the export that produced the numbers above — not rewritten for this page.

The other figures on this site keep the labels they had. The pricing page is still arithmetic on public list prices, and the routing snapshot below is still a dated snapshot — publishing a benchmark does not retroactively promote anything else.

The audit

Audit every claim on this site.

For three months, our autonomous workforce reported success — 230 completed runs, score 75/100, every time. Then we checked: the verifier returned the same score on runs that had failed, and 150 of those runs had never been verified at all. We paused the workforce ourselves and published the full finding.

Every number, date, and capability statement on keenlabs.pro is listed below with its source. We apply our own product's Truth Auditor principle to ourselves — if a claim is not verifiable, it does not belong here.

Last full audit: 2026-08-08

Workforce metrics

0 agents active — workforce paused deliberately after our own audit caught the Truth Auditor verifier failing open

Source: ai.keenlabs.pro/api/public/workforce/now, active_agent_count

Verified: 2026-08-08

✓ verified

66 total brain nodes, 0 added in the last 7 days while paused

Source: ai.keenlabs.pro/api/public/brain/stats

Verified: 2026-08-08

✓ verified

2 agent-drafted PRs merged after review (24h, pre-pause)

Source: git_log: PR #81 + #82

Verified: 2026-05-21

✓ verified

12+ brain notes auto-committed (24h, pre-pause)

Source: git_log: commits matching "brain: add"

Verified: 2026-05-21

✓ verified

12+ models in active daily rotation

Source: src/lib/models config, internal estimate

Verified: 2026-05-21

⚠ self-reported, not independently audited

335 models reachable and measured by our catalogue pipeline, counted not rounded, growing weekly; the router promotes whatever the evidence ranks best

Source: founder_self_reported — catalogue pipeline

Verified: 2026-08-08

⚠ self-reported, not independently audited

5 registered sources

Source: command_sources table after Sprint O1 Phase 2

Verified: 2026-05-21

⚠ self-reported, not independently audited

7 sprints shipped (pre-pause)

Source: Derived from SPRINT_LOG: A, B, C, E2, E3, W1, K6

Verified: 2026-08-08

⚠ self-reported, not independently audited

Projects

Pensionskollen active since May 2026

Source: founder_self_reported

Verified: 2026-05-21

⚠ self-reported, not independently audited

Pensionskollen has 0 paying customers

Source: founder_self_reported

Verified: 2026-05-21

✓ verified

Keenoble active since May 2026

Source: founder_self_reported

Verified: 2026-05-21

⚠ self-reported, not independently audited

Product capabilities

Smart routing across a 335-model library, 12+ in active rotation

Source: src/lib/models config + founder_self_reported catalogue pipeline

Verified: 2026-08-08

⚠ self-reported, not independently audited

56 reviewer decisions logged: 25 accepted first try, 30 accepted on revision, 1 flagged for human

Source: ai.keenlabs.pro/api/public/reviewers/outcomes

Verified: 2026-08-08

✓ verified

Sprint C shipped the Truth Auditor verifier; our own audit later found it failing open (approving output it should have blocked) — the reason the workforce is paused

Source: founder_self_reported

Verified: 2026-08-08

⚠ self-reported, not independently audited

For three months the workforce reported success on 230 completed runs, scored 75/100 every time

Source: founder_self_reported

Verified: 2026-08-08

⚠ self-reported, not independently audited

The verifier returned a 75/100 score even on runs that had failed

Source: founder_self_reported

Verified: 2026-08-08

⚠ self-reported, not independently audited

150 of those 230 runs were never verified at all

Source: founder_self_reported

Verified: 2026-08-08

⚠ self-reported, not independently audited

What we DON'T claim

We do not claim revenue (we have none yet)

We do not claim independent benchmarks (none exist)

We do not claim customer count beyond zero

We do not claim Truth Auditor catches every hallucination — only the ones it sees

We do not claim cost or speed figures without cited measurement methodology

If you spot something we should remove or qualify, email hello@keenlabs.pro.

Standing rule

We publish our findings and method — never our current routing configuration. The findings are the pitch; the config is the moat.

This page is the archive. New findings enter the catalogue with a status, and measured same-test results land in skill cards with their row counts.