Research · updated August 2026

What we find when we look

Every engagement starts with a data diagnostic, which means we have now examined the financial data of a substantial number of mid-market companies under similar conditions. The findings are remarkably consistent, and almost none of them were known to the company beforehand.

Want yours examined?

A diagnostic takes a week from read-only access. The findings are yours whether or not you continue.

1 / 3
n = 74 diagnosticsMeasured, not surveyedFindings recur across industries
8.4%median duplicate rate in vendor masters
31%of companies had a period that did not balance
64%had at least one account untied for over a year
74diagnostics run between 2024 and 2026
Extractread source, no writesMapaccounts, customers, vendorsLoadinto a staged tenantReconciletrial balance, per periodgate · must tieReviewyour controller signsCut oversource goes read-onlygate · must tievariance → back to mapping, never waivednothing advances past a gate until the trial balance agrees to the penny

What a diagnostic examines and in what order. Extraction is proven before anything is assessed, because an assessment of an incomplete extract is worse than none.

What we found

What recurs, across industries and company sizes.

These findings appear at similar rates regardless of revenue, industry, or which accounting system the company runs, which suggests they are properties of how finance data accumulates.

Duplicate master records

Median 8.4% in vendors and 6.1% in customers. Every company believes theirs is cleaner and three of seventy-four were below three percent.

Missing dimension data

A median of 22% of transactions carry no department, location, or project. Historical periods are worse, because the dimension was frequently introduced part-way through.

Periods that do not balance

31% had at least one closed period where the trial balance did not tie, usually by a small amount and usually unnoticed for years.

Untied accounts

64% had at least one balance sheet account with no reconciliation in over twelve months, and in most cases nobody could say which accounts those were without checking.

Chart of accounts bloat

Median 340 accounts of which 38% had no activity in three years. The dormant ones are mostly dimension-shaped accounts created for a purpose that has passed.

Orphaned and inconsistent references

Transactions pointing at deleted classes, items referencing closed accounts, and customers merged then recreated. Individually harmless, collectively why reporting cannot be trusted.

None of this stops the system working

That is the central point and the reason these problems persist. A company with an 11% vendor duplicate rate and four untied accounts closes its books every month, files its taxes, and passes its audits. The system works.

What does not work is anything built on top of it. Reporting is approximately right, automation escalates constantly because the patterns are inconsistent, and a migration carries every one of these problems into the new system where they are then blessed by a project.

None of this stops the books closing. It stops everything you would want to build on top of the books.

Why the duplicate rate is always a surprise

Duplicates do not look like duplicates from inside. Each record was created by somebody with a reason — a slightly different legal name, a new contact, a subsidiary, a supplier who changed their trading name. From the ledger they are separate vendors and each is individually correct.

They only become visible when matched on tax identifier and bank account, which is not a comparison anyone runs in normal operation. The bank account match is the one that most often surfaces something genuinely worth investigating.

The unbalanced period

Thirty-one percent is higher than we expected when we started counting. The amounts are usually small — under a thousand dollars — and they date from a system migration, a bulk import, or a period that was reopened and reclosed.

They matter for two reasons. Any consolidation or reporting built on that period inherits the discrepancy, and an auditor who finds it will ask questions that are difficult to answer years later. Finding them is a query that takes minutes.

Dimensions and the retrospective problem

The 22% of transactions with no dimension data is not evenly distributed. It concentrates in older periods, because the dimension was introduced part-way through and nobody back-populated.

That is why year-over-year comparison is frequently impossible in companies that believe they have dimensional reporting. The current year is dimensioned and the comparison year is not, so the comparison is either wrong or missing. Deriving dimensions retrospectively from transaction evidence is possible and it is a real piece of work.

What we do with the findings

Report them, in writing, before you have committed to anything further. A meaningful share of diagnostics conclude that the data work should precede any system decision, and several companies have taken the report and fixed the findings themselves.

That is a legitimate outcome. A company that fixes its duplicate rate and its untied accounts is in a materially better position to evaluate any system, including ours.

How we measured this

Seventy-four data diagnostics run between January 2024 and June 2026 on US mid-market companies between roughly $8M and $180M in revenue, across QuickBooks, Xero, NetSuite, Sage Intacct, Dynamics, and Acumatica.

All figures are measured from extracted data rather than reported. Duplicate detection uses fuzzy matching on name, tax identifier, billing address, and bank account, with a conservative threshold — the true rates are likely somewhat higher than reported here.

The sample is self-selecting toward companies that engaged us, which usually means they had a problem they were aware of. Whether that correlates with data quality is unknown, and we would guess it does somewhat.

No company is identifiable in the aggregates. Individual diagnostic reports are confidential to the company that commissioned them.

Questions

Common follow-ups.

How long does a diagnostic take?
About a week from read-only access, with roughly six hours of your team’s time. The findings are written and yours regardless of what you decide.
Do you need write access?
No. Read-only throughout, and nothing changes in any of your systems.
Is our data really that bad?
Probably about like everyone else’s. Three of seventy-four had duplicate rates below three percent, and every one of the other seventy-one believed theirs was cleaner.
Can we fix these ourselves?
Most of them, yes, and several companies have done exactly that with the report. It puts you in a better position to evaluate any system, including ours.
Why do these problems persist?
Because none of them stop the books closing. They only stop what you would want to build on top of the books, and that is a less urgent kind of broken.

Look before you migrate.

Every one of these findings gets carried into a new system and blessed by the project that carried it.