How AccountingBench is designed, executed, and scored.
Questions are drawn from professional examination materials and university coursework, aligned with real-world accounting competency standards across educational and professional practice levels.
Each model answers each question 3 times. The majority answer is used as the final answer, reducing random variance.
Structured answers (SC/MC) are scored by formula. Open-text responses are evaluated by GPT-5-mini as an LLM judge, rating 0–100 on correctness and completeness following specific grading criteria.
Scores are averaged per category and framework, then aggregated to an overall figure. All runs are timestamped and fully reproducible from logged configuration metadata.
Single-choice and multiple-choice questions use a penalized scoring formula. Correct selections earn points, incorrect selections subtract points.
Open-text responses are scored by GPT-5-mini (0–100). Each question runs 3 times — SC/MC uses majority vote, open-text averages the three judge scores. If a model fails to produce a parseable answer, it scores 0, but is not included in the benchmarking results.
| Range | Badge | Meaning |
|---|---|---|
| ≥ 90% | Excellent | Near-expert level |
| 75–90% | Strong | Professional-level |
| 60–75% | Moderate | Requires verification |
| 50–60% | Weak | Below professional threshold |
| < 50% | Poor | Not suitable for professional use |
@misc{kaburek2026accountingbench,
author = {Manuel Kaburek and Ewald Aschauer and Alexander Hofer and Markus Isack},
title = {AccountingBench: A Structured Benchmark for Systematic Evaluation of Large Language Models in Accounting Education and Professional Tasks},
year = {2026},
institution = {Vienna University of Economics and Business}
}