Duplicate master records
Median 8.4% in vendors and 6.1% in customers. Every company believes theirs is cleaner and three of seventy-four were below three percent.
Research · updated August 2026
Every engagement starts with a data diagnostic, which means we have now examined the financial data of a substantial number of mid-market companies under similar conditions. The findings are remarkably consistent, and almost none of them were known to the company beforehand.
A diagnostic takes a week from read-only access. The findings are yours whether or not you continue.
What a diagnostic examines and in what order. Extraction is proven before anything is assessed, because an assessment of an incomplete extract is worse than none.
What we found
These findings appear at similar rates regardless of revenue, industry, or which accounting system the company runs, which suggests they are properties of how finance data accumulates.
Median 8.4% in vendors and 6.1% in customers. Every company believes theirs is cleaner and three of seventy-four were below three percent.
A median of 22% of transactions carry no department, location, or project. Historical periods are worse, because the dimension was frequently introduced part-way through.
31% had at least one closed period where the trial balance did not tie, usually by a small amount and usually unnoticed for years.
64% had at least one balance sheet account with no reconciliation in over twelve months, and in most cases nobody could say which accounts those were without checking.
Median 340 accounts of which 38% had no activity in three years. The dormant ones are mostly dimension-shaped accounts created for a purpose that has passed.
Transactions pointing at deleted classes, items referencing closed accounts, and customers merged then recreated. Individually harmless, collectively why reporting cannot be trusted.
That is the central point and the reason these problems persist. A company with an 11% vendor duplicate rate and four untied accounts closes its books every month, files its taxes, and passes its audits. The system works.
What does not work is anything built on top of it. Reporting is approximately right, automation escalates constantly because the patterns are inconsistent, and a migration carries every one of these problems into the new system where they are then blessed by a project.
Duplicates do not look like duplicates from inside. Each record was created by somebody with a reason — a slightly different legal name, a new contact, a subsidiary, a supplier who changed their trading name. From the ledger they are separate vendors and each is individually correct.
They only become visible when matched on tax identifier and bank account, which is not a comparison anyone runs in normal operation. The bank account match is the one that most often surfaces something genuinely worth investigating.
Thirty-one percent is higher than we expected when we started counting. The amounts are usually small — under a thousand dollars — and they date from a system migration, a bulk import, or a period that was reopened and reclosed.
They matter for two reasons. Any consolidation or reporting built on that period inherits the discrepancy, and an auditor who finds it will ask questions that are difficult to answer years later. Finding them is a query that takes minutes.
The 22% of transactions with no dimension data is not evenly distributed. It concentrates in older periods, because the dimension was introduced part-way through and nobody back-populated.
That is why year-over-year comparison is frequently impossible in companies that believe they have dimensional reporting. The current year is dimensioned and the comparison year is not, so the comparison is either wrong or missing. Deriving dimensions retrospectively from transaction evidence is possible and it is a real piece of work.
Report them, in writing, before you have committed to anything further. A meaningful share of diagnostics conclude that the data work should precede any system decision, and several companies have taken the report and fixed the findings themselves.
That is a legitimate outcome. A company that fixes its duplicate rate and its untied accounts is in a materially better position to evaluate any system, including ours.
Seventy-four data diagnostics run between January 2024 and June 2026 on US mid-market companies between roughly $8M and $180M in revenue, across QuickBooks, Xero, NetSuite, Sage Intacct, Dynamics, and Acumatica.
All figures are measured from extracted data rather than reported. Duplicate detection uses fuzzy matching on name, tax identifier, billing address, and bank account, with a conservative threshold — the true rates are likely somewhat higher than reported here.
The sample is self-selecting toward companies that engaged us, which usually means they had a problem they were aware of. Whether that correlates with data quality is unknown, and we would guess it does somewhat.
No company is identifiable in the aggregates. Individual diagnostic reports are confidential to the company that commissioned them.
Questions
Every one of these findings gets carried into a new system and blessed by the project that carried it.