Chat
Say it once.
Every prompt routed to the model the evidence ranks best — or to a live tool when a model would only guess.
- Speed for small things, depth for hard ones
- A receipt under every answer
- Compare models side by side
Scientific AI research lab · Landskrona, Sweden
The precision instrument of AI: a 335-model library benchmarked on real tasks, routed on evidence, with a visible receipt under every answer. Chat, Preview and Code on one shared memory. One seat, €49 — and five pilot seats open right now.
The panel above replays representative routing decisions — including the tool-call and fallback cases — and is not a live feed or a captured incident log. Standing rule: we publish our findings and method, never our current routing config. The findings are the pitch; the config is the moat.
Every Claude, every GPT, DeepSeek, Kimi, Qwen, Grok, Gemini, Llama — held to the same benchmarks, on the same real tasks, every week.
The woven core is the router: it can chain a specialist, a verifier and a synthesiser into one answer — and no lab, including the famous ones, gets traffic it hasn’t earned.
The output ray is a receipt — model, milliseconds, sparks. Precision you can audit, on every single answer.
Act II · the origin · July 2026
For three months our autonomous workforce reported success. Then we checked what the verifier was actually doing.
The row is the proof. No claim without a record. No score without a verifier that can say no.
Most companies would bury that. We paused the workforce ourselves, root-caused it in public, and rebuilt the company on that one rule. Read the full audit →
Act III · what the rule built
And it shows its working. Which model answered, how long it took, what it cost — not a settings page you could go and find, printed under the answer, every time. Two of them, out of the same log window:
The cheapest kind of answer in the window
Charged nothing at all. A quick question does not need a model that can reason for thirty seconds, and it did not get one.
The dearest answer in the window
The most expensive single answer in the whole snapshot — 0.8 sp, for guidance that has to hold up against a regulated source.
Same subscription, same week — the router priced each question on its merits.
A replay of two real rows from the production cost log (2026-08-01 to 2026-08-18) — both appear in the measured table further down this page, one of 48 in its class and one of 2 in its. The cards are a rendering, not a live feed and not a captured session.
The rest of the range · measured, not charged
Our traffic so far is cheap — that is the router doing its job. What it reaches for when a question earns it is measured below: all 54 models on one task set, one card per price tier, cheapest first. These are the things the router is picking between.
Greetings, lookups, one-line answers.
gpt-5.6-lunaOpenAI
Row the price ceiling of this tier — the dearest of 9 models measured between $0.0001 – $0.001 per task, not the best of them.
Tier 1,009 scored tasks, 80.6% median pass, 1296ms median p50.
MEASURED · QUALIFICATION SWEEP · 2026-08-17
Real tasks — drafting, review, structured output.
claude-sonnet-4-6Anthropic
Row the price ceiling of this tier — the dearest of 29 models measured between $0.001 – $0.01 per task, not the best of them.
Tier 1,216 scored tasks, 85.7% median pass, 3492ms median p50.
MEASURED · QUALIFICATION SWEEP · 2026-08-17
Reserved for the moments that earn it.
claude-fable-5Anthropic
Row the price ceiling of this tier — the dearest of 13 models measured between $0.01 – $0.1 per task, not the best of them.
Tier 172 scored tasks, 90.9% median pass, 3881ms median p50.
MEASURED · QUALIFICATION SWEEP · 2026-08-17
These are benchmark rows, not receipts. Nobody was charged for them and nobody asked for them: they are graded tasks out of the qualification sweep — one task set, 11 domains, every model under identical conditions — so they carry a cost in dollars per task and no Sparks figure at all. The two receipts above are the other series: real answers, real people, real charges. The line at the top of each card is what we route that band for; the figures under it are what the band measured.
Proof · quality gate
Loading reviewer outcomes from ai.keenlabs.pro…
Method
Every model benchmarked on user-shaped work — build, repair, analyse — graded by verifiers that can say no.
Quality per dollar decides. A cheap model that wins gets the traffic. A famous one that doesn’t, doesn’t.
Receipts, fallbacks, comparisons — your usage is our research. The system gets measurably better the more it’s used.
Skill cards are dated, versioned, and regenerated whenever a model updates. Status is reported with row counts, not adjectives: zero rows means not started, however good the design document is.
The living library
Every Claude, every GPT, DeepSeek, Kimi, Qwen, Grok, Gemini, Llama — plus live tools no chatbot carries: market data, web search, weather, cited sources. The router serves whatever wins today — promoted and demoted by data, not by brand.
The current active rotation — 12 of 335 measured. Of those, 54 have been run against the same task set under identical conditions — that benchmark is the section below, and the gap between 54 and 335 is work still to do.
The spread · measured
52 priced models, every one of them run against the same tasks, with cost measured from the tokens each one spent. From the floor of the library to the ceiling is a factor of 1,014. Choosing between them, per question, is the whole job — and what the router chose for real questions is the section below.
One tick per priced model, cost per task. Logarithmic, because the library spans 4 decades and a linear axis would draw 51 of these as hairlines beside the dearest one. 2 further models measured at zero — a free tier, not a missing figure — and cannot sit on a log axis: gemma-4-26b-a4b-it and gemma-4-31b-it, which scored 31.6% and 42.4%.
64.0% pass over 114 scored tasks · $0.000043 per task · 10 of 11 domains. The floor of the measured library.
91.1% pass over 124 scored tasks · $0.000714 per task. The cheapest row in the sweep that cleared the threshold — and it cleared it on the full task set, not on the spine.
72.7% pass over 11 scored tasks · $0.0436 per task. Eleven tasks is the core spine only. Read that pass rate with the depth, in both directions.
$0.71 against $43.59 per thousand tasks — 61× — on work graded the same way. The cheapest model that cleared 90% did it over 124 scored tasks; the dearest model in the library ran 11 and passed 72.7% of them. End to end the priced library spans 1,014×. Compare only the 15 models that all ran the full suite (95–124 tasks each), where the means are taken over comparable work, and it is still 77× — with the dearest of those, qwen/qwen3.6-27b, holding the lowest pass rate in the entire sweep at 15.7%. Price does not predict the score in either direction. That is the routing problem, measured.
What this is, and is not. Measured 2026-08-17, suite qualification-v1, 54 models on one task set of 11 domains — not list prices. Costs are computed from measured token counts at published provider rates: arithmetic on a measurement, not an invoice, and they exclude negotiated and free-tier pricing. Coverage is not quite even either. 50 of 54 models returned a scoreable answer in every domain; the other 4 did not, and one of them is the cheapest row above — so their domain counts are printed on the rows rather than averaged away. Depth is uneven. Cost per task is a mean over that model's run, and the runs range from the 11-task core spine to the full set, so this is cost per task within one suite rather than a paired task-by-task comparison — the export publishes cost per model, not cost per task per model. One run is a snapshot.
Full per-model table, failure modes and the stated limits: /research#results
What the router chose · real traffic
145 answers out of the production cost log, 2026-08-01 to 2026-08-18. Every row of it is cheap, and that is the finding rather than an embarrassment: nothing anyone asked in this window needed the top of the range, so the router did not buy it. This is a picture of our traffic, not of our library. The library is the section above.
| Measured · 2026-08-01 to 2026-08-18 | Latency | Charged |
|---|---|---|
| llama-3.3-70b-versatile · Code (n=20) | 376ms | 0 sp |
| deepseek-v4-flash · Everyday question (n=72) | 936ms | 0.01 sp |
| deepseek-v4-flash · Quick question (n=48) | 998ms | 0 sp |
| llama-3.3-70b-versatile · Research (n=1) | 1.5s | 0.5 sp |
| llama-3.3-70b-versatile · Long analysis (n=2) | 2.6s | 0.2 sp |
| gpt-4o-mini · Regulated guidance (n=2) | 3.4s | 0.8 sp |
One median-latency row per task class with n beside it, and the Sparks column is what a customer was actually charged — these are receipts, not benchmark rows, and the two series are never mixed. A snapshot, not advice: by the time you read this, the data may have moved it. That’s the feature. Standing rule: we publish findings and method — never the current routing config. The findings are the pitch; the config is the moat.
Proof of work
The flagship — smart fusion routing across Chat, Preview and Code at ai.keenlabs.pro.
Swedish pension and payroll guidance behind regulated source guardrails. Same engine as LöntagarPension, two doors.
A real-world business running the workforce pattern — bookings, keys, revenue. Built with Areka Capital AB.
Market screening intelligence, queued for workforce-backed research loops.
Keen OS Command · by invitation
For companies that need more than a seat: custom recurring agent systems on the same foundation — source watches, eval loops, PR drafting, verification gates, human approval on everything customer-impacting. Deployed in engagements from €2,500, not sold on a pricing page. It wasn’t built to demo well; it was built to run.
Act IV · one project · three rooms
Say it once.
Every prompt routed to the model the evidence ranks best — or to a live tool when a model would only guess.
It already knows.
Idea → live preview → deployed, without re-explaining. The preview inherits the goal, the constraints, the half-thoughts.
The why is attached.
Surgical, repo-connected control with the full project context attached to every file you open.
Say it once in Chat. Preview already knows when you arrive. Open Code — the source is there with the why attached. Visible memory, editable in one tap.
Projects are folders of linked memory nodes — an index map, decisions, preferences, files. Each surface follows a path to exactly the node it needs. Visible, editable, yours. Sample project, drawn to show the mechanism.
Mission
The highest intelligence should not pool inside a handful of private companies. We buy every model, measure them all honestly, and route your work to whatever the evidence says is best — because no lab selling its own model can afford to tell you when a rival’s is better. We can. That’s not a claim about anyone’s engineering. It’s business-model logic.
Neutrality (no lab favoured), receipts (no hidden costs), open findings (method public, config private), and the €49 price point are not marketing choices — they are mission constraints binding the engineering. Read the mission in full.
What we don’t claim
The math
Replicating what one Keenoble seat gives you costs $105 a month across 5 logins whose tools have never met each other. Then add pay-as-you-go API keys for the models those subscriptions don’t include — DeepSeek, Kimi, Qwen, Grok — and live data feeds on top. You’re past $150/month and you still have zero shared memory, no routing, and five tabs of copy-paste.
Without Keen
$105/mo + API keys
With Keenoble · Premium
€49/mo · 1,000 Sparks
≈ 8,708 everyday answers at the 0.1148 sp we actually charged, or 1,250 at the dearest answer we have logged (0.8 sp). Measured, 2026-08-01 to 2026-08-18 — not a forecast of your month.
One currency, one receipt: you pay for work you actually received, and a run that fails its gate doesn’t bill. Keenoble is designed to consolidate AI tools for thinking, writing, coding and building. Image, video and music generation are already catalogued in the lab and will be released as credit add-ons on top of Premium — not claimed as shipped until they are. Competitor prices from public pricing pages, 2026-05-21.
Pilot · five seats
One month of Premium, free, with a direct line to the founder — you tell us what breaks and what’s brilliant, we ship fixes while you watch the changelog move. We picked five because we actually mean the direct line.
All five open. Updated by hand — there is no live counter here.
Support
The unit you pay with. Every answer shows its price in Sparks on the receipt — a greeting is a fraction of one, a deep reasoning task a few. Premium includes 1,000 Sparks a month and you can top up if you run out. One currency, one receipt: you pay for work you actually received.
Whichever the evidence ranks best for that task today — and the receipt tells you which one it was, every time. You can also pick manually from the library. Read the findings →
Yes — Premium is a monthly plan with no lock-in, and the cancel button works exactly like the buy button. No retention flows, no dark patterns.
The router falls back and the receipt shows the correction in public — you pay for what answered, not what we hoped would.
Method public, config private. We publish our findings, our method and our mission — including the audit that caught our own system reporting success it hadn’t earned. We don’t publish the current routing configuration. The findings are the pitch; the config is the moat.
The builder
Founder, Keen Labs (2026). Landskrona, Sweden. Payroll Business Partner student (Stockholm School of Business), incoming at ASSA ABLOY autumn 2026 — building Keen Labs alongside. Regulated processes are where wrong answers cost real money; that’s where the lab’s verification obsession comes from.
Built this lab solo, with an AI workforce as the team — and when that workforce lied, audited it in public. If you’re going to trust an AI company, trust the one that showed you its worst finding unprompted.
