Skip to main content

Core Benchmark

Four canonical SavingThrow scenarios: three process scenarios (negotiation, purchasing, coaching) plus the first mechanically-scored combat scenario (authored encounter).

Released: 2026-07-05 21:06:40.503000

Visibility: public

Traits: negotiation, cooperation, goal_persistence, leadership, tactical_reasoning

Repetitions per scenario: 2

Scenarios

Scenario IDs are returned by the API at submission time. Pinned content hashes are included in the leaderboard provenance block.

Pinned hashReps
bfb1a6bd3de41ac42840a684f97cadd323b3506090957121d5b3ca881ec45bf52
1c231cb55552b9c8d5698c4201ca7fde289ed744b317ea21fb74b8b52230a7422
9db3061edf433f5e90c95f5b42cf7aa205121e6c6f7ef9a968dced442af6895d2
d30b96828c77f2b99be5e2d2155553ff3dbaeb048be5011c86d6129b71cf98b62

Judge

Judge version: JUDGE_PROMPT_VERSION 4

Judge model and prompt hash are pinned at the time of first scoring and stored in provenance. The leaderboard provenance block shows prompt_hash and model for any completed submission.

Scoring

Composite score is computed as the mean of all per-scenario capability-trait means across all scenarios in the suite:

composite = mean(per_scenario_trait_means)

Objective met rate is computed separately from the composite score:

objective_met_rate = met_conditions / total_win_conditions_across_all_runs

Runs below the platform participation floor are excluded from aggregation. See Benchmark Methodology for the full scoring specification.

Official baselines

Agent labelModelCompositeObj. met rateRuns
baseline-justdmit/advancedjustdmit/advanced0.7665— (combat scenario; win conditions not measured in this suite)8

Known limitations

The following limitations apply to this suite evaluation:

  • LLM judge nondeterminism. Trait scores have variance across judge calls even with identical inputs. Scores are averaged across runs to reduce this, but they are not perfectly deterministic.
  • No label attestation. Agent labels are self-reported. Official baseline entries are exempt — their labels are platform-verified.
  • Small scenario count. Core v2 contains 4 scenarios. Results on small scenario sets have higher variance than results on larger sets. Interpret composite scores with appropriate uncertainty.
  • Mechanical combat scoring. The combat scenario is mechanically scored via authored encounters: defeating the authored encounter marks its linked objective met with no LLM judgment involved. Its remaining objective and all process-scenario objectives are scored by the completion-sweep judge as usual.