Benchmark White Paper · Public Edition
The Cost of a Correct Answer
Seven AI providers. 25 analytical questions. One live 381,523-row manufacturing database. We measured what a correct answer actually costs, and the ranking is not what token pricing suggests.
25 / 25
The only provider in the field to answer every question correctly
$0.0046
Cost per correct answer from the Oriona agent
17.5x
What the most expensive frontier model charged for the same correct answer
A preview of the results
Price did not buy accuracy
Among the raw models, cost per correct answer spans roughly 90x between the cheapest and the most expensive, yet their accuracy varies by only 8 points. The whole Oriona run cost $0.114.
Provider
Correct
Cost / correct
vs Oriona
Oriona Agent
25 / 25
$0.0046
1.0x
GPT-5.5 (raw)
24 / 25
$0.0441
9.7x
Claude Opus 4.8 (raw)
24 / 25
$0.0489
10.7x
Claude Sonnet 5 (raw)
23 / 25
$0.0797
17.5x
Kimi K2.7 (raw)
23 / 25
$0.0089
1.9x
DeepSeek V4 Pro (raw)
22 / 25
$0.0020
0.4x
DeepSeek V4 Flash (raw)
22 / 25
$0.0009
0.2x
Single run of 175 graded answers, July 7 2026. Two raw models undercut Oriona on price and neither clears 95% accuracy. The white paper states all three limits of this benchmark in full.
What is in the white paper
The full seven-provider results table: accuracy, judge score, total cost, cost per correct answer, and latency.
Benchmark design: the 381,523-row database, the 25 questions, and how every provider was given the same direct database access.
Both cost-accuracy charts, including the frontier that shows why only one provider clears 95% accuracy under 1 cent.
A breakdown of all 12 failed answers and the three scaffolding features that prevent each failure pattern.
An honest reading: the two raw models that beat us on price, and the three limits of this benchmark stated in full.