Restoration estimate
Output · R1Cascade Property Group · IM-84213 · fire & water
An AI copilot that turns a finished 3D scan into a signed property record. The model drafts every number. A human signs it before it ships.
The record workspace: a dark evidence twin on the left, a signable legal record on the right. Blue is machine-drafted, black is human-signed.
4:40pm. Marcus, a restoration estimator, opens a delivered scan of a fire-and-water loss. The AI has already drafted the whole estimate: rooms, measurements, damage, a dollar total. It looks finished.
But he cannot tell what the software guessed from what it actually saw. So he opens the twin, re-measures three rooms by hand, and the hour the scan was meant to save is gone.
The tool had answered every question except the only one that mattered: which of these numbers can I actually stake my name on?
Restoration and insurance do not buy 3D models. They buy a number that survives a dispute. The value, and the risk, is in the numbers derived from the scan.
A flat table of extracted fields: no confidence, no provenance, everything equally certain. It looked like automation and behaved like a liability.
Distrust is not about how often the AI is wrong. It is about not being able to tell when.
They don't distrust AI for being wrong. They distrust it because they cannot tell when it is wrong, so they re-check all of it.
With no link back to the scan, fixing a value meant re-measuring by hand. The AI made work instead of removing it.
Without a record of what the AI proposed versus what a human finalized, a disputed number has nothing behind it.
A raw number is only a guess. Verity also carries where it came from, and who signed it.
The same measurement is a guess, a location, and a signed fact. The raw export only ever showed the first.
I watched estimators scope real losses on site, then re-key every measurement into Xactimate back at the desk. The scan had already saved that hour. They spent it again by hand.
Restoration estimators, plus 3 carrier adjusters, on how a scoped number gets disputed
Scoping a live loss on-site, then re-keying it into Xactimate back at the desk
Where estimators trust an auto-measurement and where they still re-check by tape
Exported scan tables audited line-by-line against the sketches they replaced
They do not need fewer AI errors. They need to see which value might be one.
A confident wrong number costs more than a blank one, so drafted values stay visibly drafted until signed.
An adjuster disputes it months later, so the audit trail is built at draft time, not reconstructed under challenge.
Re-keying erases the savings, so the verified record has to be the deliverable, not a source to copy from.
Design-target archetypes, drawn from the interviews, ride-alongs and the problem space.
Verity is not an extractor. It is a trust instrument. Its job is to make uncertainty legible and correction nearly free, so a human's attention goes only to the values that actually need judgment.
Two rules fell out of that, and I never broke them: confidence is always shown, and every value links to the region that produced it.
The model will be wrong, and its output feeds a legally disputable document. Five principles fell out of that box.
Hiding uncertainty to look clean is a lie the user pays for later.
An unlinked number cannot be corrected fast, or defended at all.
A correction flow that fights the analyst's speed gets its defaults slammed.
One action serves the dispute file and the calibration loop. Nothing is thrown away.
Earned per field type against a real reliability curve, not a threshold that flatters the demo.
Judgment shows in what you refuse. Each of these was tempting, and each would have quietly broken the thesis.
An immersive 3D walkthrough as the main surface. It demos beautifully.
Why I killed it: The job is verifying numbers, not flying through space. The 3D became a focused evidence pane, not the product.
Let the model write the estimate and skip the human. The cleanest leverage story.
Why I killed it: A disputed number nobody signed is indefensible. The human stays on the values that carry money and risk.
Type a question, get a scope. Modern and effortless.
Why I killed it: Nothing to verify against. Q&A became a cited layer on the record, not a replacement for it.
The audit trail is an object, not a log line. Every value knows where it came from.
Scroll sideways to read the full diagram →
Scroll sideways to read the full diagram →
The machine drafts every value in blue. Marcus signs only what he will defend in black. These are the two paths he walks at his desk.
Clearing the record one value at a time.
Anything below the line, or unbacked by the scan, re-enters the loop until he can sign it.
Turning confirmed damage into a number a carrier cannot wave off.
No AI-estimated dollar reaches the carrier without being verified first.
Two scales, and they are never allowed to collide: who authored a value, and how sure the model is.
Geometry, glow and the probe live here. Nothing signable does.
The legal document. Every value here is signable and disputable.
One question, asked of every value on screen: has a human signed this?
A separate scale, spent only on how sure the model is. Each tier names an action.
Scroll sideways to see every state →
Color answers who authored this and how sure the model is. It never answers what stage the job is in. The moment color starts carrying workflow status, both of the scales above stop meaning anything.
The record workspace carries the whole trust idea, so it took the most iteration: five paper screens each pinned to one open question, then a greyscale pass that locked confidence and provenance before color could flatter a weak decision.
Not about layout. About the one question every screen had to answer: where does a value's evidence live, and how does a human sign it?
Greyscale forced the hard calls before color could flatter them: a ring on every value, focus that flies to the source, and a machine-versus-human split.
Dark evidence left, light legal record right, machine-blue settling to signed black. The color law does the trust work the wireframe only promised.
Sixty drafted values on one screen. The design's whole job is making any one of them a glance to check and a keystroke to fix.
Damage extent is a judgment call, not a measurement. The model proposes, the estimator adjudicates, and only a confirmed finding is allowed to become money.
This is the document a carrier actually receives, so every dollar on it has to trace back to the finding that justifies it.
A confidence score is worthless if it is not calibrated. This layer keeps the score honest, then hands the dial to the human.
Answers you can check. It cites the regions it read, and refuses outright when the record cannot support an answer.
Every principle above was forged by a failure below. This is where the design actually came from.
A vaulted ceiling, the value the model is least sure of, rendered exactly like a floor area it nails. The interface had laundered a guess into a number.
Confidence has to be visible AND calibrated. Visibility alone adds anxiety, calibration alone stays invisible.
Eleven seconds an edit. Analysts batch corrections, then skip them and accept the defaults to move on.
A correction path that fights the analyst's speed loses, and takes the audit trail down with it.
One line over-accepted vaulted ceilings, the model's weakest field, while holding back floor areas it nails.
Auto-accept is a per-field decision against real reliability, and the human should own how aggressive it gets.
I sat a working restoration estimator through the prototype on a real delivered scan. Every time he reached for his tape instead of the screen, the design was wrong. Three of those moved it.
A low-confidence value stopped meaning re-measure and started meaning a ten-second look at the source. That is the hour the scan was always supposed to save.
None of these calls were mine alone, and none of them were free.
Co-defined the damage taxonomy and what counts as a finding.
Set where confidence comes from and what the score is allowed to claim.
Cut MVP scope with me, desktop review first.
Compute and storage. Every value stores a link back to its source region.
A number that corrects in seconds and survives a dispute.
Throughput. More fields route to a human than the model strictly needs.
Fewer shipped errors. In restoration, an error becomes a dispute.
The field companion app.
Focus on the desk, where the estimate is signed and sent.
Design targets, not shipped results. Each pairs a mechanism I can point to in the interface with the outcome it is meant to move.
A glance and one key, instead of a manual re-measure.
Needs-review sorts to the top, calibrated-high collapses into quiet rows.
The model takes the fields it is provably good at, nothing else.
A disputed estimate becomes defensible in minutes, not an argument.
Every edit feeds calibration. The tool gets less wrong the more it is used.
Measured from telemetry: correction rate, calibration drift and time-per-record, in a phased rollout.
A well-displayed but miscalibrated score is the real danger. First test: does showing confidence calibrate trust, or just add anxiety?
From enterprise teams to growing startups.