How to Read a Data Provider's Accuracy Number
Every data provider claims an accuracy number, but what that number describes is not necessarily the most important information to an underwriter. Two providers can report high accuracy on the same field while measuring different things.
Over the past year we have run several head-to-head benchmarks between ResiQuant’s Engineering AI and incumbent data sources in insurance (see our reports on geocoding, historic designation, roof age and submission intake). Confusion on the definition of accuracy has came up every time. The worst case is when the reported accuracy metric measures something other than what a user assumes. This white paper is intended to explain to underwriters and consumers of data how to interpret and question the “accuracy” statements of the different services in the market so you can make more informed decisions about what licenses are worth your budget.
What counts as correct?
On a field like roof age, the definition of correct has to be decided before anything can be scored. A roof replaced in December may not appear in imagery for a few months, so some tolerance is reasonable. How much tolerance is acceptable is the main choice that can move the result significantly.
In our roof age study, ResiQuant named the exact year a roof was replaced on 66% of properties, against 50% for SpatialKey and 46% for BuildFax. Allowing one year of tolerance raised and compressed those figures to 75%, 64% and 64%. Allowing two raised them to 78%, 69% and 68%. Every score rose, and the distance between them narrowed from sixteen points to nine.
A wider tolerance lifts every provider and compresses the differences between them. A figure reported without its tolerance cannot be compared with any other.

How often the same answer is right?
Some attributes are rare. Fewer than one building in three in our historic designation study was listed on the National Historical Registry. On a field like that, a system that answered no to every address would score 72% accuracy without finding a single listing.
MapRisk scored 91% accuracy on that set. Its recall, the share of genuinely listed buildings it actually found, was 74%. Both figures are accurate. Most of the 91% comes from buildings that were never historic in the first place, and recall isolates the part that depends on the product. ResiQuant's recall on the same buildings was 90%.
For a rare attribute, recall is the important number for an underwriter. Accuracy alone will flatter any system working against an unbalanced ground truth.

Counted out of what?
There are two ways to count how many buildings an extraction system missed. One way is to count across every building in a test set, agnostic to which submission it was in. The other way is to count how many submissions came back with at least one building missing.
In our submission intake benchmark, ResiQuant Submission AI returned 98.7% of all buildings. Claude Opus 5 returned 90.0% and GPT-5.6 returned 79.6%.
However, an underwriter works through one submission at a time, and a missing building isn’t flagged. When a system drops a building, nothing in its output shows that anything was skipped, or which building it was. If a submission is missing even one, the only safe response is to check every building in it against the source documents, however well the rest were extracted.
That makes the second count the more important one. ResiQuant Submission AI returned a complete schedule on 93.0% of submissions. Claude Opus 5 dropped at least one building on 12.7% of submissions, and GPT-5.6 on 25.5%. Each of those is a file someone has to reconcile line by line.

Who wrote the answer key?
Every accuracy figure is a comparison against a source of truth, and that data set does as much to set the score as the system being tested. A key assembled from the outputs of the systems under evaluation will tend to agree with whichever of them it drew on most.
The strongest keys are independent of every system being scored. Our historic designation study graded each answer against the National Park Service register itself, with a reference number recorded for every listed building. Our geocoding study checked each returned point against verified parcel boundaries and building imagery. Neither key depended on any provider's answer.
Before relying on a reported figure, it is reasonable to ask what the result was compared against and who assembled it.
Is a label the same as a result?
Many providers attach a confidence label to each answer. A geocoder may classify a point as a rooftop match, the highest tier it offers. A label of that kind describes how the system assessed its own answer.
In our geocoding study, the three leading commercial geocoders each offer a top-tier rooftop result. When every returned point was checked against the parcel it was meant to land on, the most confident of the three turned out to be the least accurate.
Self-reported confidence and measured accuracy are separate quantities. In this study they ran in opposite directions.
Which way does it miss?
Each question above concerns how often a system is right. On some fields, the more important question is what a wrong answer means.
Calling a roof older than it is costs margin, and the error surfaces at the quote. Calling it newer than it is prices a worn roof as a sound one, and the error surfaces at the claim. In our roof age study, ResiQuant never called a roof a decade or more newer than the imagery showed. CoreLogic Spatial did on 17% of properties.
On fields where the two directions of error carry different prices, the direction of the misses decides what being wrong costs. No accuracy figure contains it.
Conclusion
An accuracy figure in a provider’s marketing materials is a summary, and summaries leave things out. Each section above reduces to one question that recovers part of what was left out.
What tolerance was it scored at? A wider band raises every score and narrows the gaps between providers.
How often the same answer is right? For a rare attribute, recall is the more important figure.
Was it counted by building or by submission? A single missing building sends a whole submission back for review.
What was left blank? A skipped answer carries no error flag and never enters an accuracy figure.
Who built the answer key? The strongest keys are independent of every system being scored.
Is it a measured result or a confidence label? Self-reported confidence and measured accuracy can run in opposite directions.
Which way do the misses run? On some fields, the direction of an error decides what it costs.
Asked of any provider, these questions turn a single marketed percentage into a figure an underwriter can compare and rely on.




