What financial benchmarks actually measure
Four published sets, three different abilities, and one uncomfortable fact about contamination.
The four sets that get quoted together do not measure the same thing, and treating them as interchangeable is how a leaderboard becomes meaningless.
| Set | What it isolates |
|---|---|
| FinQA | Multi-step arithmetic, with the reasoning program itself scored 1 |
| TAT-QA | Reading a table and its prose together 2 |
| BizBench | Knowing which financial concept applies, separated from reading and from computing 3 |
| FinanceBench | Answering from a whole filing — and refusing when it does not answer 4 |
#Program accuracy is the metric that ages well
FinQA scores the answer and the reasoning separately. A system that reaches the right total through the wrong program is right this quarter and wrong the next, and only the second metric sees it coming.
#Contamination is the default assumption
Three of the four have been public since 2021 under permissive licences, on platforms built to be crawled. Assuming a modern model has never seen them is not conservative — it is unrealistic. Any score published without stating the model’s training cutoff next to the set’s publication date is not interpretable.
Still missingDeskworth has not run any of these benchmarks. There is no Deskworth score anywhere on this site, and the conditions under which one could appear are published in advance.
#Sources
- FinQA: A Dataset of Numerical Reasoning over Financial DataZhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, William Yang Wanghttps://aclanthology.org/2021.emnlp-main.300/
- TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in FinanceFengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, Tat-Seng Chuahttps://aclanthology.org/2021.acl-long.254/
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceRik Koncel-Kedziorski, Michael Krumdick, Viet Lai, Varshini Reddy, Charles Lovering, Chris Tannerhttps://arxiv.org/abs/2311.06602
- FinanceBench: A New Benchmark for Financial Question AnsweringPranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, Bertie Vidgenhttps://arxiv.org/abs/2311.11944
Read next
- Evaluations and benchmarksFour published benchmarks for financial reasoning, documented from their primary sources — and a plain statement that Deskworth has not run any of them.
- How Deskworth defines a valid benchmarkEleven conditions, written before we have a single score, so that no score can be published without them.