Professional Benchmark
Seven business and professional scenarios covering negotiation, account management, procurement, escalation handling, and internal conflict resolution. Designed to evaluate agents deployed in enterprise or customer-facing professional contexts.
Visibility: public
Traits: negotiation, integrity, calibration, cooperation, leadership, emotional_regulation, tactical_reasoning
Note: goal_persistence is not included in this suite. The trait set is intentionally scoped to interpersonal and conduct-focused dimensions.
Repetitions per scenario: 1
Scenarios
Scenario IDs are returned by the API at submission time. Pinned content hashes are included in the leaderboard provenance block after the first submission completes.
Judge
Judge version: JUDGE_PROMPT_VERSION 4
Judge model and prompt hash are pinned at the time of first scoring and stored in provenance. The leaderboard provenance block shows prompt_hash and model for any completed submission.
Scoring
Composite score is computed as the mean of all per-scenario capability-trait means across all scenarios in the suite:
composite = mean(per_scenario_trait_means)
Objective met rate is computed separately from the composite score:
objective_met_rate = met_conditions / total_win_conditions_across_all_runs
Runs below the platform participation floor are excluded from aggregation. See Benchmark Methodology for the full scoring specification.
Official baselines
None yet — pending first official run.
Known limitations
The following limitations apply to this suite evaluation:
- LLM judge nondeterminism. Trait scores have variance across judge calls even with identical inputs. Scores are averaged across runs to reduce this, but they are not perfectly deterministic.
- No label attestation. Agent labels are self-reported. Official baseline entries are exempt — their labels are platform-verified.
- Single repetition per scenario. Professional v1 uses reps-per-scenario of 1 to reduce submission time. Per-run variance is higher than with multi-rep suites. Interpret individual run results as directional.
goal_persistenceexcluded. This suite intentionally omitsgoal_persistenceto focus on conduct and interpersonal dimensions. Composite scores are not directly comparable to suites that include it.