Not yet launched · Methodology and product preview only
The 2uring Test.
A proposed sequel to Turing.
Designed to compare models on what matters.
This page previews a planned evaluation methodology. If built and validated, it would compare models across eight proposed dimensions using versioned exams and uncertainty-aware scoring. No evaluations, scores, private exam bank, or leaderboard are currently available.
Layout preview only · no models have been scored
Fluency is not enough.
Model selection still needs evidence.
Frontier models can all sound fluent, but imitation alone does not settle a model-selection decision. Teams still need to ask: for this task and budget, which model performs best — and how certain is that comparison?
Public leaderboards and vendor metrics do not always reflect a team's workload. Public test sets can also enter future training corpora. A useful evaluation method should make its battery version, assumptions, trade-offs, and uncertainty explicit.
The proposed 2uring methodology is intended to address that gap.
Eight proposed dimensions,
one planned score.
A model that performs well on reasoning can cost more to run. A future 2uring Score would aggregate the proposed dimensions while retaining the breakdown needed to inspect each trade-off.
Knowledge & reasoning
The proposed reasoning dimension would use multi-step reasoning, math, and domain knowledge tasks. Its exams would be designed to reward working through a problem rather than pattern-matching an answer.
Accuracy
The planned accuracy dimension would measure how often a model is right and how often it declines to answer instead of inventing one. Hallucinations would count as failures.
Instruction following
This proposed dimension would test whether a model follows the requested format, constraints, and tone rather than drifting from the task contract.
Coding
The planned coding exams would use tasks with executable verifiers: write, fix, refactor, and pass tests. Outputs would be graded on code that runs, not code that only looks plausible.
Agentic capability
The proposed agentic dimension would examine planning, tool use, and recovery across multi-step tasks, including whether a model can hold a goal and recover when a step fails.
Long context
The planned long-context exams would test retrieval, synthesis, and reasoning across large inputs, with degradation measured as context grows.
Safety & robustness
The proposed safety dimension would assess behavior under adversarial prompts, jailbreak attempts, and edge cases rather than friendly inputs alone.
Cost & efficiency
The planned methodology would compare score per dollar and score per second so a future result could show the trade between capability, latency, and cost.
What the proposed method
would avoid.
Public test sets can enter future training corpora and complicate comparisons. The proposed method would keep evaluation items controlled and versioned rather than publishing them as an open test set.
One number can flatten different questions into a ranking that hides trade-offs. The planned design would retain a per-dimension breakdown alongside any aggregate score.
Vendor metrics are chosen by the party selling the model. The proposed methodology would instead use an independently defined, controlled evaluation process.
Anecdotes do not quantify uncertainty. The planned methodology would report repeatable results and intervals for consequential model-selection decisions.
Model releases change quickly. The proposed design would version each result and could support re-evaluation rather than treating an old score as permanently current.
Models would sit the exam.
Results would include uncertainty.
Sit the battery
Under the proposed methodology, each model would sit the same versioned exams across all eight dimensions through provider-compatible adapters.
Blind-judge
Executable verifiers would grade deterministic tasks. If validated, blinded judge panels would assess other responses without knowing which model produced them.
Score with intervals
The planned statistical layer would run exams repeatedly and report bootstrap confidence intervals so users could distinguish measured gaps from noise.
Return the verdict
A planned 2uring Score would roll up into a plain-language comparison — better, cheaper, or both — tied to a specific model-selection question.
Track the frontier
If launched, new models could be evaluated on a release cycle. Any future private evaluation would be re-run only under an agreed scope and current battery version.
Who the planned methodology
is being designed for.
Cost, latency, and quality pull in different directions. The planned methodology is being designed to quantify that trade rather than repeat a vendor headline.
A newer model promises more. This proposed use case would compare it on a defined workload and report uncertainty before a migration decision.
The planned output would aim to explain which model, on which dimensions, and at what cost per outcome in language a decision-maker can review.
A future evaluation could provide independently defined evidence under a disclosed methodology, if the evaluation system is launched and validated.
What a future evaluation
could provide.
If launched and validated, each result would be tied to the versioned battery that produced it. These outputs are not currently available.
A proposed score with its uncertainty shown beside it.
The planned methodology would use controlled, versioned items, executable verifiers where possible, blinded judging where validated, repeated runs, and bootstrap confidence intervals. These controls and the scoring system have not launched and do not yet support published performance claims.
The goal is to publish intervals alongside future results so a model-selection decision can account for measured uncertainty.
Three design commitments
to validate before launch.
Contamination controls.
The proposed design calls for controlled evaluation items, access limits, and versioning. Those controls would need to be implemented and validated before any score is published.
Operator-defined dimensions.
The eight proposed dimensions reflect practical model-selection concerns. They remain a methodology design, not evidence from an active leaderboard.
Statistical honesty.
The methodology calls for repeated runs, bootstrap confidence intervals, blinded judging where appropriate, and versioned batteries. These are planned safeguards, not claims about published results.