← Back to Case Studies

Pricing Property In Seconds Instead Of Days

ML LABS designed and built a PropTech platform's core automated valuation engine, which returns a property valuation through an API in seconds, where a human appraisal took several days. It met the bar set for it: within 10% of the actual closing price in dense metro areas for 90% of cases. The estimates it cannot stand behind go to a human appraiser, which is exactly what lets the rest run without one.

The process it replaced depended on appraisers visiting properties, reviewing comparables by hand, and producing a report over several days, a structural bottleneck against the volume that enterprise buyers wanted to operate at. The engine ingests property photos, satellite imagery, geospatial features, and market data. Producing the point estimate turned out to be the straightforward half, and calibrating the confidence band around it was the half that decided whether the system could run without an appraiser behind every answer.

A property analyst with a closed tablet looks up at a stone terraced house while two colleagues talk on the street
The confidence band decided the product.

Seconds, Not Days, Within 10%

The system hit its target: valuations within 10% of the closing price in dense metro areas for 90% of cases, returned in seconds, not the days an appraisal took.

The 90% is a commercial number before it is a statistical one, and the confidence band is what converts it. Enterprise buyers would consume a probabilistic valuation only if the system told them where it was confident and where it was not. Properties outside the high-confidence band are routed to a slower workflow with a human appraiser; the rest pass through the automated path. The buyer is buying the confident slice, and the system earns more trust by admitting the boundary of that slice than by claiming a wider one.

"We needed a valuation engine that could price properties within 10% of closing price in dense markets, in seconds instead of days. Omar built it, and it lifted our real estate insights and client engagement." — Innovation Director, JLL

Signals That Cover Each Other's Gaps

Transactional comparables alone hit their ceiling fast in heterogeneous urban housing stock. Two properties on the same street differ in condition, layout quality, and outlook in ways a comparables sheet does not record, so the error distribution goes fat-tailed exactly where enterprise buyers care most: high-value properties and unusual configurations.

The fusion architecture attacks that tail directly. Computer vision over listing photos reads condition and finish quality. Satellite imagery reads what no comparables panel contains: proximity to construction, traffic exposure, street tree coverage. Geospatial features add walkability and amenity density around the address. Each source carries its own independent error pattern, and the ensemble weights them against the specific property's configuration instead of trusting a single channel whose errors all lean the same way.

Every Estimate Arrives Explained

The engine was built as a production service for enterprise integration, and each stage exists because a downstream consumer needed something from it.

Flowchart: property photos, satellite imagery, and geospatial and market data feed feature extraction, then an ensemble valuation model, which produces a price estimate with a confidence interval served through an enterprise API.
  • Feature extraction. Vision and geospatial models run in parallel over photos, satellite imagery and transactional data, each emitting structured features rather than free-form opinions.
  • Ensemble prediction. Visual, spatial, and transactional signals combined into one valuation with a calibrated confidence bound and feature-importance scores.
  • API layer. Valuations returned to enterprise clients with full feature attribution, so a downstream system can audit and explain any estimate that it acted on.

Every Price Says How Far To Trust It

The failure modes worth naming are the ones that move the estimate without announcing themselves. Comparable scarcity pushes weight onto the visual channel and exposes the model to staged photos. A renovation the comparables panel has not caught up with produces an under-valuation the photo channel is supposed to correct, but only if the agent photographed the renovation. Each mode has a measurable signal: the spread between channel-specific estimates widens before the ensemble estimate becomes unreliable, which is the quantity the confidence bound was calibrated against.

Standard prediction intervals from tree ensembles run over-confident in dense neighborhoods and under-confident in sparse ones, the reverse of what the routing logic needs. Research on spatially weighted conformal prediction for automated valuation models (Hjort et al., 2023) documents exactly this, and only spatially weighted calibration restores consistent coverage. Residuals were calibrated on local neighborhoods rather than a global pool, and claimed confidence was validated against held-out closings rather than believed.

Confidence is what the buyer purchases. The point estimate is what they talk about.

Accuracy Holds Where Data Is Dense

The bar was achievable because three properties of the data environment did most of the work. Dense metros have high transaction velocity, so the comparables panel is recent. Listing photos were uniformly available and curated by the source platform. Satellite revisit cadence was stable enough that spatial features tracked conditions on the ground. Weaken any of the three and the error widens, which is why suburban and rural expansion was treated as a separate model regime rather than a bigger training set.

Three Steps To A Trusted Estimate

  1. Start in the densest market you have. Prove accuracy against actual closing prices where the comparables panel is thickest, before anyone discusses expansion.
  2. Verify each channel earns its place. A new data source must reduce the error independently, not add correlated noise that the ensemble then has to unlearn.
  3. Calibrate the band locally, then check its coverage. Validate the interval against held-out closings inside each neighborhood, not against a national pool.

Confidence That Ships With Every Price

Accuracy and calibration are two problems with two test harnesses: one for point-estimate error against closing prices, one for interval coverage against the same ground truth. Every new data source, model change, and geography expansion passes both or it does not ship.

A demo returns a confident number for the property in a thin postcode. A product returns a number and a band wide enough to send it to a human, and that is a design decision made at build time. Taking one workflow that far, from ingestion and calibration to routing, is what an engineering partnership is for, kept running by the person who built it. Most AI projects fail on exactly this seam, between a model that works and a system that can be trusted.

References

  1. Hjort, A., Hermansen, G. H., Pensar, J., & Williams, J. P. Uncertainty quantification in automated valuation models with spatially weighted conformal prediction. arXiv preprint, 2023.

NEXT · TO PRODUCTION

Could this work for you?

Two minutes. Find out where you stand.

15 minutes · no charge · with Omar