FeaturesPricingSupportFAQ
DE Get in touch
FeaturesPricingSupportFAQ
Get in touch

On this page

  1. Why 99.8 percent is not enough
  2. What a check would have to prove
  3. Confidence scores show how sure the model is, not whether it is right
  4. A cited source shows where a figure is, not whether it is right
  5. Two models can be wrong together
  6. XBRL tags do not always fit your model
  7. A total that adds up proves less than you think
  8. What checks are good for
  9. Why the human look cannot be skipped

Why no automatic check can replace review

Julian from jumoca FLOW · Published 27 September 2026

Confidence scores, cited sources, a second model, XBRL data, totals that add up: each of these safeguards sounds as if AI figures could go into a model unchecked. Why that holds for none of them, what each one really measures, and what that means for the word “automatic”.

← Articles

Why 99.8 percent is not enough

The best AI models read financial statements remarkably well, and every generation gets better. So why not just wave the figures through?

Because “automatic” is a promise without exceptions. If a value goes into your model unseen, you need a rule that always holds, not one that almost always does. Suppose a model reads 99.8 percent of all figures correctly. With 150 figures per model, roughly one run in four then contains at least one wrong figure. Nobody knows which one. It looks just like the other 149.

And it is not always the same figure. One time it is revenue, in the next report a provision. So there is no short list of fields to keep an eye on. To find the one wrong figure, you have to look at all of them, and that is exactly the review you wanted to save.

High accuracy even makes it worse. After two hundred correct values, nobody looks closely at the two hundred and first. A tool that is almost always right teaches you to trust it exactly where you should not.

What a check would have to prove

For a check to release a value, it would have to know the right figure independently. That is exactly what it cannot do. Every automatic check compares the answer with something from the same document, from a second AI run, or with the judgement of the very AI that produced the wrong value.

So a check can show that something is wrong, but never that something is right. If two routes arrive at different figures, at least one of them is wrong. If they arrive at the same figure, both can still be wrong.

Many of these checks do measure something, just not what they are credited with. The following sections show this for several of them.

Confidence scores show how sure the model is, not whether it is right

A confidence score says how sure the model is of its answer. It comes out of the same step that produces the error. If the model takes the prior-year figure, it does so because that figure looks right to it at that moment. Had it been in doubt, it would have taken a different one.

That is why a model does not report 50 percent for an error it has not noticed. It reports as much as for a correct figure, often 99 or 100 percent. Part of the reason is how a language model writes a number: not in one go, but in small pieces called tokens, for 4,011 for example “4”, “,” and “011”. There is usually a real choice only at the first or second token, where it is decided which figure is meant at all. Once that start is written, the rest follows almost inevitably, and every further token comes with around 99 percent certainty. If the confidence score is an average over all tokens, as it often is, the few close decisions disappear among the many certain ones.

After all, the prior-year figure is a perfectly plausible figure, sitting in the right table under the right label. 100 percent only means that no other answer came into question for the model, and that is exactly true of a plausible figure from the wrong column. Nor does a low score mean the figure is wrong. It only tells you how likely the sequence of tokens was with which the model chose this figure.

A language model consists entirely of probabilities. 99 percent means: in similar cases from training, this token was the right one 99 times out of 100. That says nothing about the individual case. The score does not tell you whether this figure is one of the 99 or the one that is wrong. And whether the 99 percent still holds for a report the model has never seen in this form is an open question.

A cited source shows where a figure is, not whether it is right

Many tools name a source for every figure, and some even check automatically that the figure actually appears there. This check shows exactly one thing: the figure occurs in the document. It does not show whether it is the right figure for the field.

The typical errors therefore go unnoticed. The prior-year figure, the revenue of a single segment and the figure from the notes are all in the report. The check confirms each of them, and it is even right to, because each of them is there.

Two models can be wrong together

Another approach: a second model reads the same report, or the same model is asked twice, and only what agrees is accepted. Agreement shows that both read the same thing, not that they read it correctly.

Models learn from similar data and read the same layouts in similar ways. A table with the prior year on the left, a unit stated only once above the table, two lines with almost the same label: what misleads one model usually misleads the other. If both pick the same wrong figure, the agreement even makes the error look safer.

Models trained specifically on financial reports do not change this. Such training makes a model more accurate, and it often gives its answers even higher confidence scores. But it still works with probabilities: it picks the figure that, based on everything it has learned, most likely fits. The higher hit rate applies to its answers overall, not to the individual figure. And because such models learn from similar reports, they also share their typical errors, for example with unusually structured tables or units.

XBRL tags do not always fit your model

Listed companies also publish their financial statements in machine-readable form, with XBRL tags. It is tempting to check the AI’s figures against them.

An XBRL tag assigns a figure from the financial statements to a fixed item of the XBRL taxonomy, such as revenue. If these items match the fields in your model exactly, you do not need AI at all, because you can take the figures straight from the tags. If they do not, for example because your model defines EBITDA or net debt differently or the figure only appears in the notes, someone has to decide which tag belongs to which field. That is the same problem as reading the figures from the report.

This is why XBRL tags cannot be used reliably as a check either.

A total that adds up proves less than you think

The most convincing check is arithmetic. If the line items the AI read add up exactly to the total it read, surely they must be right?

A total that adds up only means the figures used fit together. It does not mean each figure sits in the right line. The total adds up perfectly with all of these errors:

  • All values, the total included, come from the prior-year column. The prior year adds up too.
  • All values are in thousands instead of millions. The total is just as consistent, only a thousand times too small.
  • Two lines are swapped, say cost of sales and distribution costs. The total does not change.
  • Two errors of the same size cancel each other out.

If you confirm the total yourself, the first two cases are ruled out, the last two are not. A sum check tests whether the right figures are present, not whether each one is in its place.

And a swapped line is not harmless just because the total is right. Cost of sales and distribution costs determine the gross margin, current and non-current liabilities every liquidity ratio. Every ratio built on a single line is wrong, and the check still reports: correct.

What checks are good for

None of this makes checks useless. They work as a warning, not as a release. If a check fires, you know that something is off and where to look, and that is worth a lot. If it does not fire, you only know that nothing contradicts, not that everything is right.

CheckWhat it showsWhat it cannot show
Confidence scoreHow sure the model isWhether the figure matches the report
Figure appears at the sourceThe figure occurs in the documentWhether it is the right figure for the field
Two models agreeBoth read the sameWhether both read it correctly
Comparison with XBRLThe AI read the same figure the company taggedWhether the tag means what your field means
Total adds upThe figures fit togetherWhether each figure is in the right line

As a warning, each of them helps. As a licence to put a value into the model unchecked, each of them fails, and precisely in the cases that matter.

Why the human look cannot be skipped

Each of these checks replaces a person looking with a comparison against something that can be wrong in the same way. A figure is only right once someone has seen it in the report, looked at the field in the model and decided that the two belong together.

None of this speaks against AI. Modern models read in seconds what takes a person hours, and they are mostly right. That advantage is worth using. The only point is that no value goes into the model unchecked. That is why jumoca FLOW is built around the review, not the extraction.

Software should not abolish that look, but make sure it takes as little time as possible. FLOW shows every value on the report page it came from, with the figure highlighted. If the value is right, one key is enough, if it is wrong, one click on the right figure. Nothing goes into the workbook automatically, not even when every check agrees.

The mistakes unchecked AI tools have already caused are covered in Why faster is not always better.

Auditable data extraction for Excel models.

Resources

Articles Help Provider docs

Legal

Imprint Privacy Terms

Language

Deutsch

A product of jumoca · © 2026