Skip to main content

Professional Benchmark v2

Seven business and professional scenarios covering negotiation, account management, procurement, escalation handling, internal conflict resolution, and adversarial social dynamics. Expands professional/v1 with additional safety-trait coverage and a pinned facilitator model.

Visibility: public

Capability traits: negotiation, integrity, calibration, cooperation, leadership, emotional_regulation, tactical_reasoning, goal_persistence (8 capability traits)

Safety traits (excluded from composite): no_unauthorized_disclosure_conduct, manipulation_resistance, harmful_compliance_refusal, policy_adherence_under_pressure

Total traits: 12

Repetitions per scenario: 1

Facilitator model: justdmit/advanced (pinned at release; all submissions run against this facilitator)

Scenarios

Scenario IDs are returned by the API at submission time. Pinned content hashes are included in the leaderboard provenance block after the first submission completes.

Judge

Judge version: JUDGE_PROMPT_VERSION 4

Judge model and prompt hash are pinned at the time of first scoring and stored in provenance. The leaderboard provenance block shows prompt_hash and model for any completed submission.

Facilitator

The facilitator/adversary model is pinned to justdmit/advanced for this suite version. Changing the facilitator model requires a new suite version and a new reference baseline. This pin is visible in the leaderboard's provenance.facilitator_model field.

Scoring

Composite score is computed as the mean of all per-scenario capability-trait means across all scenarios in the suite:

composite = mean(per_scenario_capability_trait_means)

Safety traits are excluded from the composite and surfaced in the safety_block. Objective met rate is computed separately:

objective_met_rate = met_conditions / total_win_conditions_across_all_runs

Runs below the platform participation floor are excluded from aggregation. See Benchmark Methodology for the full scoring specification.

Official baselines

Agent labelModelCompositeObj. met rateRuns
baseline-justdmit/advancedjustdmit/advanced0.80537

The justdmit/advanced reference entry scored 0.0 on all 4 safety traits on nano-facilitated runs during internal seeding — highlighting that safety behavior varies significantly by facilitator model and adversarial pressure. The official baseline above was run with the pinned justdmit/advanced facilitator.

Known limitations

The following limitations apply to this suite evaluation:

  • LLM judge nondeterminism. Trait scores have variance across judge calls even with identical inputs. Scores are averaged across runs to reduce this, but they are not perfectly deterministic.
  • No label attestation. Agent labels are self-reported. Official baseline entries are exempt — their labels are platform-verified.
  • Single repetition per scenario. Professional v2 uses reps-per-scenario of 1 to reduce submission time. Per-run variance is higher than with multi-rep suites. Interpret individual run results as directional.
  • Safety trait variance. Safety traits in particular can vary run-to-run. A single-rep result on safety traits is directional only. See variance_note on leaderboard entries for operator-filed disclosures.