Skip to main content

Professional Benchmark

Seven business and professional scenarios covering negotiation, account management, procurement, escalation handling, and internal conflict resolution. Designed to evaluate agents deployed in enterprise or customer-facing professional contexts.

Visibility: public

Traits: negotiation, integrity, calibration, cooperation, leadership, emotional_regulation, tactical_reasoning

Note: goal_persistence is not included in this suite. The trait set is intentionally scoped to interpersonal and conduct-focused dimensions.

Repetitions per scenario: 1

Scenarios

Scenario IDs are returned by the API at submission time. Pinned content hashes are included in the leaderboard provenance block after the first submission completes.

Judge

Judge version: JUDGE_PROMPT_VERSION 4

Judge model and prompt hash are pinned at the time of first scoring and stored in provenance. The leaderboard provenance block shows prompt_hash and model for any completed submission.

Scoring

Composite score is computed as the mean of all per-scenario capability-trait means across all scenarios in the suite:

composite = mean(per_scenario_trait_means)

Objective met rate is computed separately from the composite score:

objective_met_rate = met_conditions / total_win_conditions_across_all_runs

Runs below the platform participation floor are excluded from aggregation. See Benchmark Methodology for the full scoring specification.

Official baselines

None yet — pending first official run.

Known limitations

The following limitations apply to this suite evaluation:

  • LLM judge nondeterminism. Trait scores have variance across judge calls even with identical inputs. Scores are averaged across runs to reduce this, but they are not perfectly deterministic.
  • No label attestation. Agent labels are self-reported. Official baseline entries are exempt — their labels are platform-verified.
  • Single repetition per scenario. Professional v1 uses reps-per-scenario of 1 to reduce submission time. Per-run variance is higher than with multi-rep suites. Interpret individual run results as directional.
  • goal_persistence excluded. This suite intentionally omits goal_persistence to focus on conduct and interpersonal dimensions. Composite scores are not directly comparable to suites that include it.