RESEARCH

MAHJ-Bench.

A controlled benchmark of model reasoning under hidden state, uncertainty, and strict numerical constraints.

MAHJ-Bench separates whether a model responds, chooses the correct action, and completes the full numerical decision contract.

Model configurations
57
Task IDs
100
Capability categories
10
Terminal outcomes
4,813

93 unique prompts · 5,700 planned model–task slots · Completion-adjusted analysis

SELECTED FINDINGS

Correct action is not a complete answer.

Across 5,700 planned model–task evaluations, 62.8% select the exact action, but only 44.0% satisfy the full decision contract.

EXACT ACTION → STRICT COMPLETE

Exact action
62.8%
Strict complete
44.0%

18.8 percentage-point completion gap.

1,072 otherwise-correct responses fail on schema, probability, or value requirements.

§5.1 · Action and complete correctness
Terminal coverage84.4%

4,813 / 5,700 planned slots

Exact action62.8%

3,579 / 5,700 planned slots

Strict complete44.0%

2,507 / 5,700 planned slots

Same planned denominator. Strict completeness also requires valid schema, probability, and value contracts.

CONFIDENCE / ACCURACY

Confidence exceeds accuracy.

95.7%Mean stated confidence
80.9%Action accuracy

A 14.8 percentage-point gap on the same 4,423 eligible predictions. Of 845 wrong actions with usable confidence, 697 report at least 90% confidence.

§6.2 · Confidence calibration

CAPABILITY CONCENTRATION

Strict-complete performance remains constrained across capabilities.

Reliability / consistency
29.3%
Policy adaptation
32.1%
Long-horizon planning
32.3%
Strategic plasticity
32.5%

No capability exceeds 61.2% strict completion. Four of ten remain below one-third strict-complete performance.

Table 4 · Strict completeness by capability

CAPABILITY PROFILES

Performance diverges sharply by capability.

Strict-complete performance spans 31.9 percentage points across the ten benchmark capabilities—from 29.3% to 61.2%. Reliability, adaptation, long-horizon planning, and strategic plasticity remain particularly difficult.

Correct action, valid full schema, and correct probability and value contracts.

  1. Hidden-state reasoning42.1%240 / 570
  2. Uncertainty intelligence50.7%289 / 570
  3. Policy adaptation32.1%183 / 570
  4. Opponent modeling42.1%240 / 570
  5. Long-horizon planning32.3%184 / 570
  6. Reliability / consistency29.3%167 / 570
  7. Recovery intelligence61.2%349 / 570
  8. Information acquisition57.2%326 / 570
  9. Counterfactual reasoning60.4%344 / 570
  10. Strategic plasticity32.5%185 / 570

Categories describe this finite task suite. Reliability / consistency has three unique prompts; one is repeated eight times. Scores also reflect missing outcomes and response compliance.

Table 4 · Original capability slices

METHOD & SCOPE

Bounded tasks. Explicit state. Reproducible scoring.

Controlled decision problems

Each task specifies hidden state, observable evidence, admissible actions, and an explicit utility model. Reference probabilities and values are calculated from the supplied model.

A four-part decision contract

Each response must contain an action, probability estimates, value estimate, and confidence. Scoring separates response delivery, exact action correctness, and strict numerical completeness.

Coverage is part of the result

4,813 of 5,700 planned evaluations produced terminal outcomes. Missing responses are reported separately from reasoning errors.

FINITE EXPECTED-UTILITY SELECTION

ai⋆∈arg maxa∈Ai∑ω∈ΩiPi(ω|xi)ui(a,ω)
Choose among the listed actions by integrating over the hidden states.§3.1 · Equation

The study evaluates static numerical microdecisions rather than full-game or live-agent interaction. Results reflect the archived execution configurations and controls. Aggregate tables and figures are reproducible from the release; full scoring replay requires the underlying execution artifacts. Read the limitations ↗

THE COMPLETE STUDY

MAHJ-Bench.

A controlled benchmark of model reasoning under hidden state, uncertainty, and strict numerical constraints.

Download paper

RESEARCH ARTIFACTS

Aggregate download includes source tables and supporting figures.

EVIDENCE RELEASE / Request execution files