Are Automated CMAs Accurate? Where They Beat Manual Comps and Where They Don't
The question is usually asked as though it has one answer. It has six — one per stage of the process, and automation wins about half of them decisively.
Short answer
Automation is more accurate than a human at retrieval and arithmetic and less accurate at comp selection and condition. Net effect: a reviewed automated CMA usually beats a manual one, because it removes the error class humans are worst at (transcription, arithmetic, omission) while the agent still supplies the judgment. An unreviewed one is worse than either. The 15-minute review isn't a formality — it's the whole difference.
This is the accuracy companion to automated CMA vs. manual comps, which covers the time math. Speed is easy to argue about. Accuracy is the objection that actually stops brokerages from adopting, and it deserves a straight answer.
First: an automated CMA is not an AVM
Most "automated CMAs are inaccurate" arguments are actually about AVMs, and the two get conflated constantly.
| AVM | Automated CMA | |
|---|---|---|
| Output | One estimated number | A comp set, adjustments, context, a range |
| Human in the loop | None | Agent reviews and signs |
| Shows reasoning | No | Yes — that's the point |
| Handles condition | No | Via the agent's walkthrough |
| Accountable party | Nobody | The agent |
An AVM answers "what is it worth." An automated CMA answers "here is the evidence, here is what it implies, and here is where I disagree with it." Zillow's own published accuracy — around 1.74% median error on-market and 7.20% off-market — is a fair benchmark for the first and irrelevant to the second.
Stage by stage
Here's the honest breakdown, and it doesn't go one way.
| Stage | More accurate | Why |
|---|---|---|
| Retrieving candidate sales | Automated | Queries every match; humans stop when they have enough |
| Transcribing figures | Automated | No typos, no transposed digits |
| Computing adjustments | Automated | Arithmetic applied consistently every time |
| Market statistics | Automated | Full dataset, correctly segmented |
| Selecting which comps count | Manual | Boundaries, severances, local knowledge |
| Assessing condition | Manual | Requires having seen the property |
| Thin-comp markets | Manual | Knowing when to stop rather than pad |
Where automation is clearly better
Agents under-rate this half because these errors are invisible. Nobody notices a transposed square footage — the report just quietly carries a wrong adjustment through to the price.
- It doesn't stop searching early. An agent who finds four decent comps at 9pm on a Wednesday stops. A query doesn't get tired at comp five.
- It doesn't fat-finger. Manual comp tables are transcription exercises, and transcription has a nonzero error rate every single time.
- It applies the same logic to every comp. Humans unconsciously adjust harder against comps that disagree with the number they've already formed. This is the most under-discussed accuracy problem in manual CMAs.
- It computes market stats over everything. Segmented absorption and DOM by band, not eyeballed.
The manual error most likely to move your price isn't a bad comp. It's the fourth comp you adjusted a little harder because it didn't fit the number you'd already decided on. Automation is immune to that one, and agents almost never count it.
Where automation is clearly worse
- Boundaries. Distance is a proxy for similarity and a poor one. Software will happily cross a school district line, an arterial or a jurisdiction boundary because the property is 0.4 miles away. This is the most common and most consequential automated error. Screening rules here.
- Below-grade space. If the source record lumps finished basement into total square footage, an automated adjustment inherits the error and multiplies it.
- Condition. Two identical-on-paper homes, one renovated in 2023 and one original 1987, are the same property to a dataset. This is the single largest source of automated error.
- Thin markets. Software asked for six comps tends to return six. Knowing that only two are real — and saying so — is judgment.
- Non-disclosure states. Where sale prices aren't public record, anything leaning on public data has less to work with than it appears to.
The review that closes the gap
Fifteen to twenty minutes, in this order:
- Boundary check. Look at each comp on a map against the boundaries you know. This catches the biggest error class first.
- Substitution test. Would that buyer have plausibly bought your subject instead? Anything failing this goes, regardless of specs.
- Square footage verification. Cross-check against a second source. Note anything below grade.
- Condition pass. Open the listing photos for each surviving comp. Adjust against what you saw on the walkthrough.
- Net adjustment sanity check. Anything over ~15% net is telling you it isn't comparable. Drop it.
- Thin-market honesty. If only three survive, ship three and say so. A padded set is worse than a small one.
Do that and you have the automation's advantages plus your own. Skip it and you've published a dataset's opinion under your name.
Which errors actually cost you
Not all inaccuracy is equal, and this is where the comparison gets interesting.
| Error | Typical source | Visible to client? | Cost |
|---|---|---|---|
| Comp across a school boundary | Automated | Yes — instantly | High. Local sellers know. |
| Transposed square footage | Manual | No | Moderate. Silent price error. |
| Condition not adjusted | Both | Sometimes | High. Biggest single driver. |
| Comp set stops at four | Manual | No | Moderate. Unknown unknowns. |
| Padded thin market | Automated | Sometimes | High. Range looks false. |
| Adjusting to fit a preconceived number | Manual | No | High and almost never caught. |
Automated errors are more visible, which makes them feel worse — a seller spots a wrong-side-of-the-boundary comp immediately. Manual errors are quieter and can be equally expensive. Visibility isn't the same as magnitude, though it is the same as embarrassment.
What to ask a vendor
- Does comp selection use boundaries, or just radius?
- Is below-grade square footage separated from above-grade?
- What happens in a thin-comp market — does it pad, or report the shortfall?
- Is each comp labelled as a confirmed closed sale, a list price, or an estimate?
- Can I remove a comp and have everything downstream recompute?
That last one matters more than it sounds. If removing a bad comp means rebuilding the report by hand, the review step won't survive contact with a busy Thursday. Fuller list in the buyer's guide.
The bottom line
"Are automated CMAs accurate" is the wrong question, because it assumes accuracy is one property of a whole process. It isn't — it's a property of each stage, and the stages split cleanly.
Automation wins retrieval, transcription, arithmetic and statistics. Humans win selection, condition and knowing when the data is too thin. A reviewed automated CMA combines both and is generally the most accurate option available to a working agent. An unreviewed one is the least.
The fifteen minutes is not overhead. It's the part where the accuracy comes from.
Frequently asked questions
Are automated CMAs accurate?
They are more accurate than a human at retrieval and arithmetic and less accurate at comp selection and condition assessment. A reviewed automated CMA is typically at least as accurate as a manual one because it eliminates transcription and calculation errors while the agent still supplies the local judgment. An unreviewed one is not.
Is an automated CMA the same as an AVM?
No. An AVM produces a single estimated value from a statistical model with no human involvement. An automated CMA assembles comparable sales, adjustments and market context into a report that an agent reviews, adjusts and signs. One is a number, the other is an argument.
What errors do automated CMAs make?
Selecting comps across boundaries that matter — school districts, arterials, jurisdiction lines — because distance is used as a proxy for similarity. Treating below-grade square footage as living area. Missing condition differences entirely. And producing a full comp set in a thin market where the honest answer is that few real comps exist.
How long should reviewing an automated CMA take?
Around 15 to 20 minutes. Check each comp against local knowledge, remove anything that fails the substitution test, verify square footage against a second source, and apply condition adjustments from the walkthrough. Skipping this step is what makes automated reports inaccurate.
Related reading
- Automated CMA vs. Manual Comps: What a 30-Agent Office Actually Saves
- Is AI Actually Good at Real Estate Comps Yet?
- CMA vs. Appraisal: The Difference That Costs Deals
- How to Find Real Estate Comps Without Guessing
Sources
- Zillow published Zestimate accuracy figures — 1.74% median error on-market, 7.20% off-market.
- Review timings and error-frequency characterisations are practitioner estimates based on the stages described, stated as such, not survey data.