Research

Finance intelligence research

What makes a financial document hard for a language model, and what a desk has to add before an answer is worth acting on.

Primary sources, read and citedFinexia · published · updated · 6 sources

A financial answer is not a sentence. It is a number, the window it covers, the document it came from, and the date that document was read. Drop any one of those four and the sentence survives — which is exactly the problem. This page collects what the published work says about that gap, and what we build against it.

#Why filings resist a language model

Three properties of financial documents make them harder than they look, and each one is measured in the literature rather than asserted here.

The scale of the gap is the part worth sitting with. FinanceBench evaluated sixteen model configurations on questions its authors describe as clear-cut, and found that a strong model paired with a retrieval system answered incorrectly or declined on 81% of the sampled cases 1. That result is from 2023 and the models have moved; the structural point has not. A retrieval system that returns the wrong page produces a fluent answer about the wrong page.

#Point-in-time truth

A figure is only true inside a window. Revenue for FY2025 is not revenue for the trailing twelve months, and a restated figure is not the figure that was published. A desk that answers without carrying the window is not wrong once — it is wrong in a way nobody can detect afterwards.

The consequence for a product is uncomfortable: when the data cannot answer the question at the date asked, the honest output is a refusal, not an approximation. In Finexia that refusal is a first-class state, not an error path.

#Facts, assumptions and calculations are three different things

A model that produces a number blends the three by default. A spreadsheet built by an analyst separates them, because that is what makes it auditable a year later. Finexia takes the analyst’s side: the workbook contract in the core keeps inputs, model, output, sources and checks on separate sheets, and a fact without a declared source is refused before the file is written 5.

EvidenceThat contract is written and covered end to end by tests — the file is composed, written, reopened by a path other than the one that wrote it, and its formulas found again. Version 0.0.23 extends it to a deck and a memo rendered from the same dossier, so the three cannot diverge. None of it has yet been exercised by a real agent run, and the core says so in those words 6.

#Why a specialist beats a generalist here

The argument for named roles is not that a specialist model is smarter. It is that a mandate is a written limit, and a limit is what makes a wrong answer visible. A role that may read filings and may not reach the open web cannot quietly substitute a blog post for a 10-K. A role that owns the workbook format is the only one that can produce one, so a deck request cannot silently return a spreadsheet.

This is a capability argument, not a benchmark claim. We have not run a head-to-head evaluation of one general assistant against a roster of mandates, and until we do, this page will not pretend otherwise.

#What this means for Finexia

LimitNothing on this page is a measured claim about Finexia’s own accuracy. We have not run any of the benchmarks below on our own stack. The evaluation page states the protocol we intend to follow and says plainly that it has not been run.

#Sources

  1. FinanceBench: A New Benchmark for Financial Question AnsweringPranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, Bertie Vidgenpeer-reviewed paper · arXiv:2311.11944v1 · published 2023-11-20 · read 2026-09-10 · CC BY-NC-ND 4.0https://arxiv.org/abs/2311.11944
  2. FinQA: A Dataset of Numerical Reasoning over Financial DataZhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, William Yang Wangpeer-reviewed paper · EMNLP 2021 · anthology 2021.emnlp-main.300 · published 2021-11-01 · read 2026-09-10 · CC BY 4.0https://aclanthology.org/2021.emnlp-main.300/
  3. TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in FinanceFengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, Tat-Seng Chuapeer-reviewed paper · ACL-IJCNLP 2021 · doi 10.18653/v1/2021.acl-long.254 · published 2021-08-01 · read 2026-09-10 · CC BY 4.0https://aclanthology.org/2021.acl-long.254/
  4. BizBench: A Quantitative Reasoning Benchmark for Business and FinanceRik Koncel-Kedziorski, Michael Krumdick, Viet Lai, Varshini Reddy, Charles Lovering, Chris Tannerpeer-reviewed paper · arXiv:2311.06602v2 · revised 2024-03-12 · published 2023-11-11 · read 2026-09-10 · CC BY-NC-SA 4.0https://arxiv.org/abs/2311.06602
  5. Finexia OS — handover report 0.0.22Finexiainternal handover report · Finexia OS 0.0.22 · commit d59e30e · published 2026-09-10 · read 2026-09-10Internal document, not published
  6. Finexia OS — product state at 0.0.23Finexiainternal handover report · Finexia OS 0.0.23 · commit produit 5968ef0 · lu à bcae42d · published 2026-09-10 · read 2026-09-10Internal document, not published

Read next

More in Research