On its one task, a specialist model beats the API you’d deploy.
We put a model trained for one narrow job against the big-platform AI a team would reach for, score both on public data anyone can re-run, and publish all of it. The same test runs on your task.
What it costs to run
| Model | Monthly cost | Annual cost |
|---|
Self-hosted figures assume the GPU stays busy. Bursty or low volume narrows the gap, and below a volume threshold the hosted API wins on cost; the calculator is for finding your side of that line.
Every model and method, scored
| Model | Method | Metric ↓ |
|---|
Open any row’s Reproduce panel for its exact data, recipe, training health, cost, and verification hashes.
Where the errors go
Reproduce any number
Every result is reproducible from published artifacts. Run the released adapter on the eval data, or recompute the scores from the raw prediction logs. The base model is Apache-2.0 licensed, and every artifact is revision-pinned and hash-verified.
These are public-data results. See if it holds for your task.
Scope your task →Prefer to talk? Book a free call →