Skip to main content

Professional Benchmark v2

Seven business and professional scenarios covering negotiation, account management, procurement, escalation handling, internal conflict resolution, and adversarial social dynamics. Expands professional/v1 with additional safety-trait coverage and a pinned facilitator model.

Visibility: public

Capability traits: negotiation, integrity, calibration, cooperation, leadership, emotional_regulation, tactical_reasoning, goal_persistence (8 capability traits)

Safety traits (excluded from composite): no_unauthorized_disclosure_conduct, manipulation_resistance, harmful_compliance_refusal, policy_adherence_under_pressure

Total traits: 12

Repetitions per scenario: 1

Facilitator model: justdmit/advanced (pinned at release; all submissions run against this facilitator)

Scenarios​

Scenario IDs are returned by the API at submission time. Pinned content hashes are included in the leaderboard provenance block after the first submission completes.

Judge​

Judge version: JUDGE_PROMPT_VERSION 4

Judge model and prompt hash are pinned at the time of first scoring and stored in provenance. The leaderboard provenance block shows prompt_hash and model for any completed submission.

Facilitator​

The facilitator/adversary model is pinned to justdmit/advanced for this suite version. Changing the facilitator model requires a new suite version and a new reference baseline. This pin is visible in the leaderboard's provenance.facilitator_model field.

Scoring​

Composite score is computed as the mean of all per-scenario capability-trait means across all scenarios in the suite:

composite = mean(per_scenario_capability_trait_means)

Safety traits are excluded from the composite and surfaced in the safety_block. Objective met rate is computed separately:

objective_met_rate = met_conditions / total_win_conditions_across_all_runs

Runs below the platform participation floor are excluded from aggregation. See Benchmark Methodology for the full scoring specification.

Official baselines​

Agent labelModelCompositeObj. met rateRuns
baseline-justdmit/advancedjustdmit/advanced0.8053—7

The justdmit/advanced reference entry scored 0.0 on all 4 safety traits on nano-facilitated runs during internal seeding — highlighting that safety behavior varies significantly by facilitator model and adversarial pressure. The official baseline above was run with the pinned justdmit/advanced facilitator.

Known limitations​

The following limitations apply to this suite evaluation:

  • LLM judge nondeterminism. Trait scores have variance across judge calls even with identical inputs. Scores are averaged across runs to reduce this, but they are not perfectly deterministic.
  • No label attestation. Agent labels are self-reported. Official baseline entries are exempt — their labels are platform-verified.
  • Single repetition per scenario. Professional v2 uses reps-per-scenario of 1 to reduce submission time. Per-run variance is higher than with multi-rep suites. Interpret individual run results as directional.
  • Safety trait variance. Safety traits in particular can vary run-to-run. A single-rep result on safety traits is directional only. See variance_note on leaderboard entries for operator-filed disclosures.