MAHJ-Bench — Selected website aggregate results

Source: Airende Ojeomogha, MAHJ-Bench: A Multidimensional Benchmark of Language
Model Reasoning Under Uncertainty, September 2026, 22 pages.
Evidence release: MAHJ-AsExecuted-v1.1-20260924.

This package reproduces selected aggregate tables and supporting figures for
the MAHJ Labs Research page. It is a curated extraction from the paper, not the complete
frozen evidence release. No benchmark prompts, oracle instances, per-item
responses, credentials, or account records are included.

All capability and flow results use the completion-adjusted view. Flow shares
use 5,700 planned slots; each capability uses 570. Confidence metrics use their
own eligible rows, as recorded in confidence-results.csv. Percentages and
percentage-point gaps are calculated from counts before rounding to one decimal.
The 1,234 terminal outcomes without exact-action success include parsing and
validation failures, and must not be treated as pure wrong-choice counts.

Files:
- aggregate-data.json: source counts used by the website and this exporter
- evaluation-flow.csv: seven waterfall stages and their denominators
- capability-results.csv: all ten original categories and three endpoints
- confidence-results.csv: confidence summary and the high-confidence error rate
- evaluation-waterfall.svg and capability-completion-gap.svg: vector figures
- MAHJ-Bench.bib: paper citation
- export-research-artifacts.py: figure and table reproduction script

Reproduce with Python 3 and Matplotlib:
  python export-research-artifacts.py aggregate-data.json reproduced

Full methods, limitations, and all reported model comparisons remain in the paper.
Execution artifacts may be requested from research@mahjlabs.com.
