InsideMaps · Property Intelligence

Turning a finished scan into a record you can sign

Verity is the AI copilot that turns a finished 3D scan into a verified property record: room-by-room measurements, materials, damage findings and a restoration estimate, where every number is drafted by the model and signed by a human before it ships.

PlatformWeb workspace · restoration desk
SessionsLong focused desk reviews, one estimator
At a glance
The problemEstimators can't tell which AI-drafted measurements to trust, so they re-measure the whole property by hand and lose the hour the scan was meant to save.
The outcomeDesigned end-to-end and delivered to engineering for build. Target: a signable, verified record in under 15 minutes instead of 60–90.
My roleSole product & interaction designer, embedded with a PM, ML lead, and restoration SME.
TimelineInsideMaps · 2021
ApproachA trust system: machine-drafted vs. human-signed authorship, per-field confidence, and calibrated auto-accept.
Verity
62% verified
$18,240 of $29,410 confirmed7 need you
DollhouseFloor planMeasure
Scan IM-842132,140 sq ft · 4 rooms
LivingBed 1KitchenMaster Bath118.5 ft
1 · WATER_STAIN · master bath ceiling · 71%

The record workspace: a dark evidence twin on the left, a signable legal record on the right. Blue is machine-drafted, black is human-signed.

The number nobody would sign

A record full of numbers, and not one he trusted

4:40pm. Marcus, a restoration estimator, opens a delivered scan of a fire-and-water loss for Cascade Property Group. The AI has already drafted the whole estimate: rooms, measurements, damage, a dollar total. It looks finished.

But the number he'll defend to a carrier rests on measurements he didn't take, and he can't tell what the software guessed from what it actually saw. So he does what he always does: opens the twin, re-measures three rooms by hand, and the hour the AI was supposed to save is gone.

The tool had answered every question except the only one that mattered: which of these numbers can I actually stake my name on?

Business stakes

A claim is not paid on a 3D model, it is paid on a defensible number

InsideMaps captures beautiful spatial data. But restoration and insurance don't buy models, they buy measurements and scopes that survive a dispute. The value, and the risk, lives in the numbers derived from the scan, not the scan itself.

45–90 minManual measurement per property, before structured data
~12%Estimate leakage one wrong measurement can introduce
100%Of AI fields re-checked, because none flagged which were wrong
~2×Cost when a bad number forces a recapture or disputed re-do
Before Verity

The naive AI export launched guesses as facts

The first attempt at property intelligence was a flat table of extracted fields: no confidence, no provenance, everything presented as equally certain. It looked like automation and behaved like a liability.

Before · the manual way
IM-84213_scan.model - InsideMaps Viewer 2.1
FileEditViewMeasureHelp
feet
23.8 ft
1512sf
Ready · Units: feet · Snap: OFFSel: Floor_Living
re-typed
by hand
XEstimate_IM84213_v3.xls - Excel
FileEditViewInsertFormatToolsData
100%
D11fx=SUM(D2:D10)
A
B
C
D
1
Room
Area
Line item
Amount
2
Living
512
Prime + paint, walls/clg
$1,840
3
Kitchen
486
LVP flooring + baseboard
$6,540
4
Bed 1
420
Drywall R&R, texture
$2,760
5
Master Bath
224
Water mitigation
$4,120
6
Masterr Bath
224
Dehumidifier + antimicrob
2830
7
Master Bath
224
Drywall R&R, prime + paint
$3,180
8
Hall
210
Smoke damage cleaning
$1,980
9
Exterior
288 lf
Baseboard R&R + paint
$3,340
10
Fixtures
Detach & reset
$2,820
11
TOTAL
est. re-keyed 06/14
$29,410
12
areas fromIM-84213_scan.model2← hand-entered
Sheet1EstimateSheet3
~45-90 min / property · 2 tools · every number re-keyed by hand
Two disconnected tools: a viewer with a tape-measure tool, and a separate spreadsheet where every number was re-keyed by hand. 45 to 90 minutes per property.
Before · the naive AI export
extraction_output.csvGET /api/v1/extract200 OK2023-11-04T09:12:37Z|15 fields|0.4 kb
#fieldvalue
1job_idIM-84213
2clientCascade Property Group
3scan_date2023-11-02
4total_area_sqft2140
5room_count4
6floor_area342
7ceiling_height9.0
8flooringhardwood
9water_damagetrue
10molddetected
11smoke_damagetrue
12drywall_replace_sqft214
13antimicrobial_sqft486
14dehumidifier_days3
15estimate_total29410
rows 1-15 of 15|confidence: n/a|source: n/a|schema: flat
The first pass at automation: a flat dump where a guess wore the same face as a fact. No confidence, no provenance, nothing to sign.
Six ways the raw export failed the person who had to sign it
  • 1Every value looked equally certain, so a guess wore the same face as a fact.
  • 2No value linked back to where in the scan it came from.
  • 3Correcting a wrong number meant re-measuring from scratch.
  • 4No record of who changed what, so nothing was defensible.
  • 5The model knew its own confidence, then threw it away before the screen.
  • 6The estimate lived in a separate spreadsheet, cut off from the evidence.
The real problem

The problem was not extract more, it was make every value verifiable

Field-shadowing an estimator surfaced the real dynamic: distrust isn't about how often the AI is wrong, it's about not being able to tell when. Three problems compounded.

  • 01

    Uniform-looking output forced total re-verification

    Analysts don't distrust AI for being wrong. They distrust it because they can't tell WHEN it's wrong, so they re-check everything and the saved time evaporates.

  • 02

    A wrong value had no visible source

    With no link between a number and the scan region behind it, fixing a value meant re-measuring by hand in the viewer. The AI made work instead of removing it.

  • 03

    The number could not be defended

    When a carrier or owner disputes an estimate, there's nothing to stand behind without a record of what the AI proposed versus what a human finalized. Accuracy without an audit trail isn't enough.

IM-84213 · Living room · ceiling height

One measurement, three forms

A raw number is only a guess. Verity also carries where it came from, and who signed it.

01Raw model output
model.rawunverified
predict() →
9.0ft
float · p=0.63
no room·no source·not signed
Locate
02Source region
floor plan
IM-84213linked · reconstructed mesh
Living1z · 9.0 ft
Verify
03Human-corrected final
signed record signed
@marcus · on site
8.6ft
corrected 9.0 → 8.699%
human-ink·tape check·locked
a guessa locationa signed factOnly Verity keeps all three.

The same measurement is a guess, a location, and a signed fact at once. The raw export only ever showed the first. Verity's job is to show all three, and make getting from the first to the third nearly free.

01 · RESEARCH & DISCOVERY

The scan saved the hour. Distrust gave it right back.

Before designing a screen, I sat with estimators scoping real water and fire losses, then watched them re-key every measurement into Xactimate at the desk. The 3D scan had already saved the hour of measuring. They spent it again by hand, because a flat AI table gives you no way to tell a fact from a confident guess.

interviews
11 interviews

Restoration estimators, plus 3 carrier adjusters, on how a scoped number gets disputed

shadowing
6 ride-alongs

Scoping a live loss on-site, then re-keying it into Xactimate back at the desk

survey
n=41

Where estimators trust an auto-measurement and where they still re-check by tape

analytics
40 estimates

Exported scan tables audited line-by-line against the sketches they replaced

What discovery surfaced
01

Distrust is about when, not how often

Estimators don't need fewer AI errors, they need to see which value might be one, so calibrated confidence rides on every number instead of the file as a whole.

02

A guess wearing the face of a fact

A flat table makes a confident wrong number cost more than a blank one, so machine-drafted values stay visibly drafted in blue until a human signs them in black.

03

Every number has to survive a carrier

The value gets disputed by an adjuster months later, so provenance and an audit trail have to be built at draft time, not reconstructed under challenge.

04

The saved hour leaks back out

Re-keying the scan into Xactimate erases the scan's time savings, so the verified record has to be the deliverable, not a source to copy from.

Affinity map · from raw notes to five themes
Verity · Discovery synthesisEdited 2h ago
GARK+3
Share
Trust & confidence
“I can't tell when it's wrong”
The table looks the same whether right or wrong
Re-measures every wall the scan already gave
“One bad number and I distrust the whole file”
Trusts the scan for big rooms, not closets
Show confidence per value, not per file
Insight

Estimators can't tell a fact from a confident guess, so confidence has to be calibrated and shown per value, not per file.

Defending the number
“I have to defend this to the adjuster”
No way to prove where a number came from
Screenshots the 3D view to back a line
Disputes drag when there's no evidence
“It's my name and license on the estimate”
Link every line item to its scan region
Insight

Every value gets challenged by a carrier eventually, so each one links to its source region and lands in an audit trail from the start.

The re-keying tax
“Then I type it all into Xactimate again”
Sketches on-site, re-keys at the desk
45 to 90 minutes per property, re-entered
Copy-paste means transcription errors creep in
Rounds measurements up to feel safe
Draft straight into estimate line items
Insight

The scan's saved hour leaks back out at the keyboard, so drafting has to land directly in estimate line items.

Reading the damage
“Is that whole wall wet or just the base?”
Walks the room twice to catch missed damage
Auto-tools miss a closet or a bay window
“Carpet or vinyl changes the whole line”
Pins photos to each affected area by hand
Glowing regions plus defect pins to confirm
Insight

Damage extent and material are judgment calls, so the AI surfaces candidate regions and defect pins for a human to confirm, never decide.

The auto-accept line
“Just don't make me check the obvious ones”
Speed-signs big rooms, slows on damage calls
Reviewing every value is as slow as re-doing it
“I trust it more when it flags its doubts”
Move the line: auto-accept above, review below
Insight

Estimators want to skim safe values and dwell on risky ones, so a movable auto-accept line lets them tune what still needs a signature.

Gazi
you
AFFINITY MAP · 5 THEMES
Legend
Voice of user
Observed behavior
Pain point
Opportunity
11 interviews · 6 ride-alongs · 29 notes clustered
100%+
Design-target personas
ES
The estimator
scopes the loss · signs the number
What they need

To see, per value, where the number came from and how sure the machine is, before I put my name on it.

Goals
  • Scope a loss without re-measuring every wall by hand
  • Sign only numbers I can defend to a carrier
  • Turn the scan into an estimate, not a re-typing job
Frustrations
  • A flat AI table hides which numbers to trust
  • One wrong square-foot and the carrier kicks it back
CA
The claims adjuster
represents the carrier · disputes the number
What they need

A record where every measurement and finding links to the scan region it came from, so I can accept or challenge on evidence.

Goals
  • Verify a scope without a second site visit
  • Trace any line item back to what it is based on
  • Close the claim without over- or under-paying
Frustrations
  • Contractor estimates arrive with no evidence attached
  • Disputes drag on when nobody can show the source
PM
The restoration PM
owns throughput · owns liability
What they need

A signed, auditable record per property that speeds the estimator up without putting the firm on the hook.

Goals
  • Get estimators through more properties per week
  • Ship estimates that don't bounce back in review
  • Keep a defensible record if a claim is challenged
Frustrations
  • Re-keying into Xactimate eats billable hours
  • A bad number becomes the company's liability

Design-target archetypes, drawn from the interviews, ride-alongs and the problem space.

The reframe

Verity is not an extractor. It is a trust instrument. Its job is to make uncertainty legible and correction nearly free, so a human's attention goes only to the values that actually need judgment.

That reframe set two non-negotiables. Calibrated confidence must be shown, never hidden. And provenance must link every value to the region that produced it. Everything else, down to the color law, serves the moment an estimator trusts the record enough to stop re-measuring.

Constraints & principles

The box I designed inside

AI is probabilistic and will be wrong. Its output feeds a regulated, disputable document. One estimator uses this in long focused sessions, working from a single existing capture. Five principles fell out of that box.

01

Show confidence, never hide it

A calibrated per-value score is always on screen. Hiding uncertainty to look clean is a lie the user pays for later.

02

Provenance, or it does not ship

Every derived value links to the region that produced it. An unlinked number can't be corrected fast or defended at all.

03

Correction is one gesture, not a form

The edit path is keyboard-first and inline. Friction here is fatal: a correction flow that fights the analyst's speed gets its defaults slammed.

04

The human's edit is both the audit record and the training signal

One action logs old-to-new for the dispute file and feeds the calibration loop. The correction is never thrown away.

05

Automate acceptance only above a calibrated line

Auto-accept is earned per field type against a real reliability curve, not granted by a hardcoded threshold that flatters the demo.

Rejected directions

Four versions I did not build

Judgment shows in what you refuse. Each of these was tempting, and each would have quietly broken the thesis.

01

The cinematic fly-through-first canvas

A full immersive 3D walkthrough as the primary surface. Tempting: it demos beautifully, it's unmistakably spatial, and InsideMaps already has the captures.

Why I killed it: The job is verifying numbers, not flying through space. I demoted the 3D to a focused evidence pane that answers one question: where did this value come from? Spatial data became a signed document, not a video game.

02

Full auto-accept, no review

Let the model write the whole estimate and skip the human. Tempting: it's the cleanest leverage story, and it gets an AI project funded.

Why I killed it: It destroys the one thing a restoration estimate rests on: when an auto-accepted number is disputed, no one can stand behind it. I kept the human on the values that carry money and risk, and automated only the ones the model earned.

03

A chat-only, ask-the-AI-for-the-estimate box

A clean conversational UI over the scan: type a question, get a scope. Tempting: it feels modern and effortless.

Why I killed it: There's nothing to verify against. Without the spatial record beside the answer you're trusting prose, with no way to correct a value or sign it. I made Q&A a grounded, cited layer on the verified record, not a replacement for it.

04

A confidence-free, clean UI

Hide the messy uncertainty so the product feels calm and trustworthy. Tempting: uncertainty looks like a flaw in a demo.

Why I killed it: It felt more trustworthy while being less so: a pretty lie. I killed it on ethics and calibration. The honest move is to show the score, then earn trust by making it accurate.

Information architecture

The system before the screens

Verity is grounded in a real object model, so the audit trail is a first-class object rather than a log line, and every value knows where it came from.

Object model · every derived value carries a provenance link and a verification event
CapturePropertyRecordRoomMeasurementDamageFindingEstimateLineItem
ExtractedNeeds-reviewConfirmedCorrectedFlagged

A Provenance link points every Measurement, DamageFinding and MaterialDetection at a SourceRegion (a 3D bounding box or image crop, plus model, version and confidence). A Verification event records who, when, and old-to-new value. That event is the audit trail and the training signal at once.

USER FLOWS · THE ESTIMATOR'S DESK

Every number, signed or sent back.

Verity is human-in-the-loop by construction: the machine drafts every value in blue, Marcus signs only the ones he will defend in black. These are the two paths he actually walks at his desk, one to clear the record value by value, one to turn damage findings into a figure a carrier cannot wave off.

Flow 01Sign it, or send it backThe value-by-value loop that replaces re-measuring the whole property by hand.
Open the property recordVerity drafts every value, each with calibrated confidenceConfidence clears your auto-accept line?Below the line: open the measurement inspectorDrill to the source scan regionDoes the twin back the number?No: correct the value by handSign it: black-locks the valueRoom signed, written to the audit trail
Values below the auto-accept line, or ones the scan region does not back, drop into the inspector and re-enter the loop until Marcus can sign them, one number at a time instead of re-measuring everything.
Flow 02Findings into a defensible figureDamage & material review to signed estimate ledger to white-label verified record.
Enter damage & material reviewGlowing regions and defect pins surfaced on the twinConfirm or reject each defect pinLedger links each line item back to its findingDollars confirmed, or still AI-estimated?AI-estimated: open the underlying findingVerify in the twin, adjust the lineSign the estimate, build the verified shareWhite-label record sent to the carrier
Any dollar still marked AI-estimated loops back to its finding in the evidence twin and must be verified before signing, so no unconfirmed number reaches the carrier's desk.
Design system for trust

Color encodes authorship and certainty, never status

Where PulseOps spent color on operational status, Verity spends it on one question, has a human signed this, plus a separate scarce budget for how sure the model is. The two systems never collide.

valueMachine-ink
AI authored · not yet checked
Human-ink
verified · signed by a person
High confidence
calibrated ≥ 88% · auto-accept eligible
Review
medium confidence · look closer
Needs review
low confidence · verify or recapture
Provenance
links a value to its source region
A VALUE, BEFORE AND AFTER
Ceiling height9.0 ft63%
Ceiling height8.6 ftsigned
CONFIDENCE BUDGET
94%auto-accept eligible
78%look closer
61%verify or recapture
Wireframe evolution

The record workspace, three fidelities deep

The record workspace carries the whole trust idea, so it took the most iteration. It began as a wall of five paper screens, each pinned to one open question; the greyscale pass locked confidence and provenance before color could flatter a weak decision; the shipped screen gave it the dark-evidence, light-record voice, where machine-blue values settle to signed black.

Lo-fi · paper
The first wall wasn't about layout, it was about the open question every screen had to answer: where does a value's evidence live, and how does a human sign it? Five sketches, one margin note each.
Mid-fi · greyscale
1Confidence ring on every value2Focus flies pane to source3Blue drafted, black signed
Greyscale forced the hard calls before color could flatter them: a confidence ring on every value, focus that flies the evidence pane to its source region, and a machine-versus-human authorship split. Annotated for the ML lead and the engineering handoff.
Hi-fi · shipped
Verity
62% verified
$18,240 of $29,410 confirmed7 need you
DollhouseFloor planMeasure
Scan IM-842132,140 sq ft · 4 rooms
LivingBed 1KitchenMaster Bath118.5 ft
1 · WATER_STAIN · master bath ceiling · 71%
The shipped record workspace: a dark evidence twin left, a signable legal record right. Blue is machine-drafted, black is human-signed.
Lo-fi · paper

Five screens, one question each

The first wall wasn't about layout, it was about the open question every screen had to answer: where does a value's evidence live, and how does a human sign it? Five sketches, one margin note each.

Mid-fi · greyscale

Confidence and provenance, before color

Greyscale forced the hard calls before color could flatter them: a confidence ring on every value, focus that flies the evidence pane to its source region, and a machine-versus-human authorship split. Annotated for the ML lead and the engineering handoff.

Hi-fi · shipped

Dark evidence, light record

The shipped workspace gives that structure its voice: a dark evidence twin on the left, a light legal record on the right, machine-blue values settling to signed black as the estimator verifies. The color law does the trust work the wireframe only promised.

Marquee flow

Deep dive: verifying an AI-drafted record

The AI drafts the whole record in seconds. The design's job is the rest: show sixty values without overwhelm, let a human verify one at a glance, and make every correction teach the system. Focusing a value flies the evidence pane to its source; A accepts, E edits, and the value flips from machine-blue to human-black.

Verity
62% verified
$18,240 of $29,410 confirmed7 need you
DollhouseFloor planMeasure
Scan IM-842132,140 sq ft · 4 rooms
LivingBed 1KitchenMaster Bath118.5 ft
1 · WATER_STAIN · master bath ceiling · 71%
Verify to black: each accepted value settles from blue to signed black on the record, while its region in the dark evidence pane settles from glow to a calm outline.
The drill-down
Floor plan Measure
Scan IM-84213source region · living room342 sq ft
LivingKitchen18.5 ft1z · 9.0 ft
Focusing a value flies the evidence pane to its exact source and draws the dimension line on the geometry. The AI read sits beside a manual check, and one key signs it.
The differentiator

Deep dive: damage and material detection review

The domain-deep surface. AI-detected damage glows in the evidence pane as translucent regions and numbered pins, each carrying a defect code, severity, and confidence. The estimator walks the pins with J and K, adjudicates each finding, and a confirmed finding maps straight to estimate line items.

Floor planDamageMaterials
Scan IM-842132,140 sq ft · 4 rooms · floor 1
KITCHEN382 sq ftLIVING521 sq ftBED 1338 sq ftHALL96 sq ftMASTER BATH144 sq ft12 sq ft21.7 ftN123
1 · WATER_STAIN · master bath · ceiling · 71%
AI-detected damage glows in the evidence pane. The estimator walks the pins, adjudicates each finding, and a confirmed finding maps straight to estimate line items.
  • Walk the pins. Each numbered damage region opens a finding card: the highlighted region, a source photo crop, and the AI read (for example, water stain, 12 sq ft, moderate, 71%).
  • Adjudicate in one gesture. Confirm, adjust boundary, reject, or reclassify. The AI proposes, the human decides, and the decision is logged.
  • Findings become money. A confirmed water stain auto-maps to line items (drywall replace plus paint), so spatial evidence, AI, and the estimate move in one motion.
The signed output

Turning verified findings into a defensible scope

The estimate ledger is the outward-facing deliverable. Every line item links back to a measurement or a damage finding, the total splits into dollars a human confirmed versus dollars still AI-estimated, and export carries the full AI-proposed versus human-final audit trail.

Verity
62% verified
$18,240 of $29,410 confirmed3 need you

Restoration estimate

Output · R1

Cascade Property Group · IM-84213 · fire & water

EstimateEST-84213-R1
Issued2026-07-24
Scope2,140 sq ft · 4 rooms
Estimator@marcus
Line itemQtyConfAmount
Water mitigationWTR
Subtotal$3,120
Dehumidifier, LGRWATER_STAIN · Master Bath
4 days signed$980
Air mover, axial fanWATER_STAIN · Kitchen
8 units · 4 d signed$1,410
Antimicrobial applicationMOLD · Master Bath
620 sq ft signed$730
DrywallDRY
Subtotal$8,940
Drywall R&R, 1/2 inCHAR · Living
1,240 sq ft signed$6,820
Texture, knockdown matchderived · wall area
1,240 sq ft91%$2,120
PaintPNT
Subtotal$6,180
Seal, stain-block primeSMOKE_DAMAGE · whole-home
1,860 sq ft signed$2,240
Paint, 2 coatsderived · wall area
1,860 sq ft78%$3,940
FlooringFLR
Subtotal$11,170
LVP flooring, supply + installMISSING_FLOORING · Kitchen
940 sq ft signed$6,060
Baseboard, 5.25 in, R&R + paintDELAMINATION · Bed 1
412 LF66%$5,110
Human-confirmed
$18,240
Still AI-estimated
$11,170
Total estimate
$29,410
Export XactimateExport PDFPDF includes full AI-proposed vs human-final audit trail
Every line item links back to the measurement or finding that justifies it. The total splits into dollars a human confirmed versus dollars still AI-estimated, so nothing half-guessed ships as final.
Verityimmutable · export-ready

Audit trail · IM-84213

38 events · Jul 25 2026 · 3:58p to 5:02p · machine-authored and human-signed entries, append-only

All events38Edits9Flags2Exports1
5:02p
MV@marcus
Record exported to Xactimate (.ESX) · total $29,410 · $18,240 signed
export
4:58p
MV@marcus
Signed Living room measurements · 4 fields (floor area, perimeter, ceiling, flooring)
verify
4:51p
MV@marcus
Flagged Kitchen for recapture · defect CHAR extent unclear from mesh
flag
4:47p
MV@marcus
AI proposed 340 sq ft at 82% · edited to 328 sq ft by @marcus
edit
value correctedmachine draft to human sign340 sq ft328 sq ft
Field
Living room · Floor area
Model
verity-visionv4.2
Source region
LR-04 · 1,284 mesh faces
Confidence before
82%review
Confidence after
human-signed
Actor
MV@marcus · Marcus Vega
Reason

Model included the bay-window recess in floor area. @marcus trimmed to the interior wall face and re-signed the value.

event 0x4f· sha256 3f9a…c21 · sealed 4:47:12p
4:39p
MV@marcus
Accepted ceiling height 8.6 ft · Master Bath, no adjustment
verify
4:33p
AI
Detected MOLD in Master Bath · 12 sq ft at 68%
detect
4:12p
AI
Drafted record from scan · 60 values across 4 rooms · estimate $29,410
draft
3:58p
SYS
Scan IM-84213 ingested · 2,140 sq ft reconstructed · 4 rooms · fire & water loss
scan
hash chain verified · head 0x8c21af · no gaps · WORM storageCopy sealExport log (CSV)
The audit trail is a first-class object, not a log line. Every accept, edit, flag and export is a timestamped, inspectable event of AI-proposed versus human-final. This is what makes a disputed number defensible.
Cascade RestorationFire & water restoration
Public verified recordPowered by Verity
Verified Property Record

1428 Camino Real

San Jose, CA 95126·IM-84213·Cascade Property Group

Every figure below was drafted by AI and signed by a licensed estimator.

AI-drafted Human-signed
Verified estimate total
$29,410
Signed by Marcus Vega
Licensed estimator · @marcus · 07/24/2026

Verified measurements

2,140 sq ft · 4 rooms scanned
Total floor area
2,140sq ft
signed
Rooms scanned
4rooms
signed
Ceiling height
9.0ft
signed
Affected area
128sq ft
signed

Damage findings & scope

6 findings · 3 rooms
WATER_STAINMaster bath ceiling
Drywall R&R · texture · prime + paint
$4,860
SMOKE_DAMAGEKitchen walls
Antimicrobial · prime + paint
$3,220
MISSING_FLOORINGLiving room
LVP flooring · baseboard
$6,540
MOLDMaster bath subfloor
Antimicrobial · dehumidifier · air movers
$2,340
CHARKitchen cabinetry
Cabinetry R&R · delamination repair
$2,180
12 line items total · every figure signed on-siteTotal $29,410
One click turns the signed record into a white-label public deliverable. The tenant's brand reskins the accent, every figure carries its human-signed provenance, and the client gets a defensible record, not a screenshot.
Designing control over the AI

The calibration layer

Verity's answer to a capacity layer, but for model trust. A reliability curve asks: when the model says 90%, is it right 90% of the time? A movable auto-accept line, drawn against that curve, lets a human decide how aggressive the automation is allowed to be.

Verity
Rolling 30 daysMVMarcus Vega@marcus

Model calibration

trust quality · last 5,000 verified fields

ECE 4.1%Brier 0.038
0%25%50%75%100%0%25%50%75%100%OVERCONFIDENT88%PREDICTED CONFIDENCEOBSERVED ACCURACY
Auto-accept threshold
≥ 88%
42%auto-accepted
Review load -38%
Legend
Model
Perfectly calibrated
Overconfidence gap
Auto-accept zone
When the model says 90%, is it right 90% of the time? The reliability curve keeps the score honest, and the auto-accept line is a control the human owns, trading throughput against review load in the open.
Grounded Q&A

Ask this property, and only trust what it can cite

A natural-language panel grounded in the verified record. Before answering, it shows which rooms and measurements it will read. Every answer carries inline citations that fly the camera to the cited region, and it refuses when the evidence isn't there.

Verity Ask
62% verified
$18,240 of $29,410 confirmed7 need you
grounded regions
2 citedFloor planSecond floor
Scan IM-84213Second floor · drywall query
Bed 1Master BathBath 2Hall16.0 ft9.4 ft12
grounded in
1Bed 1 · walls2Master Bath · walls + ceiling
Before answering, it shows which verified fields it will read. Every answer cites the source regions it used and flies the evidence pane to them, and when the record cannot support an answer, it refuses instead of inventing one.
Failures & recovery

Three things I got wrong

Every principle above was forged by a failure below. This is where the design actually came from.

The pretty lie

I built

The first record hid confidence to look clean and premium. Every value rendered the same way: calm and certain.

It broke

The clean look made a low-confidence guess read as fact. A vaulted-ceiling height, the value the model is least sure of, rendered identically to a floor area it nails, so nothing told the analyst to check it. The interface had laundered a guess into a number.

I learned

Confidence has to be visible AND calibrated. Visibility alone adds anxiety; calibration alone is invisible. You need both. This forged the show-confidence principle.

The fix

Always-on confidence rings on every value, a machine-ink to human-ink authorship law, and a dedicated calibration layer so the score is honest, not decorative.

The modal tax

I built

Correction opened a modal form with fields for the new value, a reason, and a note. Structured data, in theory.

It broke

The modal cost about eleven seconds an edit, and a path that slow doesn't survive real use: analysts batch corrections, then skip them and accept defaults to move on. The data would be slower AND worse.

I learned

A correction path that fights the analyst's speed will lose, and take the audit trail down with it. The signal has to ride along with the decision, not gate it.

The fix

Inline one-key edit: E opens a number field with the dimension line drawn on the geometry, the change logs old-to-new automatically, and the record climbs toward verified without a single modal.

The uncalibrated threshold

I built

A single fixed 80% auto-accept line applied to every field type at once. Simple, legible, one number.

It broke

It over-accepted vaulted-ceiling heights, the field the model is weakest on, while needlessly holding back floor areas it nails. One threshold can't fit every field.

I learned

Auto-accept is a per-field-type decision against real reliability, and how aggressive it gets should be a control the human owns.

The fix

Per-field-type calibrated thresholds plus a movable auto-accept line drawn against the reliability curve, trading throughput against review load in the open.

Prototyping & validation

Vetted by one estimator, then scoped for five who've never seen it

Two passes pressure-test whether the trust holds. First I ran the working prototype past domain experts and sat a working restoration estimator through it think-aloud on a real delivered scan; that pass is done and already moved the design. Then I scoped a moderated study to put Verity in front of estimators who have never seen it, protocol below, ready to run.

Track 01 · Expert reviews + estimator think-aloud
Done

Walked the live prototype through design reviews with a restoration domain expert and a claims-side adjuster, then sat a working estimator through the record workspace on a real delivered scan, think-aloud. Every place he reached for his tape instead of the screen became a to-do. Three of them moved the design:

Friction

A low-confidence ring told him to doubt the ceiling height but gave him nowhere to go, so he reached for his tape anyway, the exact hour the scan was meant to save.

Change

Made the ring the door to provenance: clicking a shaky value flies the evidence twin to the lit region and the frame it was read from, so doubt lands somewhere to look.

Friction

Confirming the record felt like signing his name to forty auto-accepted numbers he never opened, so he started opening all of them, which is auto-accept defeating itself.

Change

Split the ink in the audit trail: auto-accepted fields sign under the model on calibration, his signature only ever lands on values he actually touched.

Friction

The glowing damage regions read as certain as the measurements, so a detected water line drifted toward the estimate without a second look, the most expensive kind of wrong.

Change

Held damage findings to the same law as numbers: a defect pin can't post a line item to the ledger until its source frame has been opened once.

The pivot

Wiring provenance onto the confidence ring was the moment it clicked: a low-confidence value stopped meaning re-measure and started meaning a ten-second look at the source, which is the entire time-saving the scan had been promising and never delivering.

Track 02 · Moderated study
Scoped

Will an estimator stop hand-re-measuring the values Verity marks high-confidence, and still catch the ones it marks low, well enough to sign a number he would defend to a carrier?

Recruit
Who
Working restoration estimators who write scopes carriers dispute, three-plus years on the job.
How many
Five, the point where each added estimator stops surfacing new trust breaks.
Mix
Three Xactimate-native, two who already work from 3D scans, to separate scan skeptics from scan natives.
Screen out
Anyone who has seen Verity, so first-touch trust stays uncoached.
Format
Sixty-minute remote moderated session on one real delivered scan, think-aloud, an ease and a would-you-sign question after each task.
Task script
01

Open a delivered scan and reach a number you would defend, re-measuring only what you don't trust.

Re-measures flagged fields only, not the property.
02

A vaulted-ceiling height reads low-confidence. Decide whether to trust it.

Drills to the source region before deciding.
03

The model auto-accepted forty fields. Confirm you're willing to sign them.

Finds the calibration basis, signs without opening all forty.
04

A carrier disputes your wall-area figure. Show what stands behind it.

Opens the audit trail: AI-proposed vs human-final plus source.
05

Clear the next property faster by moving the auto-accept line.

Reads the added review load before committing.
Success bar
Pre-registered · target
Low-confidence values drilled5 / 5
High-confidence fields re-measured≤ 1 / property
Wrong values signed0
Time to a value's source< 10s
Would defend to a carrier≥ 4 / 5
Collaboration & tradeoffs

How I worked, and what I chose not to build

The damage taxonomy was co-defined with estimators; where confidence comes from was negotiated with the ML lead; the MVP scope was cut with the PM.

Provenance for every field

Accepted higher compute and storage cost to link every value to a source region, because an unlinked number can't be corrected quickly or defended at all.

Conservative calibrated auto-accept

Accepted more human review to buy fewer shipped errors. In restoration an error becomes a dispute, so a little slack is far cheaper than a wrong number.

Desktop review workspace first

Sequenced the field mobile companion to v2. The highest-stakes work happens at the desk where the estimate is signed and sent.

Projected / target outcomes

The targets I designed toward, and how I would validate them

These are design targets, not shipped results. Each pairs a leading mechanism I can point to in the interface with the lagging outcome it is meant to move, and every figure is tagged TARGET.

Leading mechanism (in the design)Lagging outcome (the business result)
Provenance-linked verify60–90 min → under 15 min
Time-to-verified-recordTarget

Every value carries its source region, so verifying is a glance and one key rather than a manual re-measure. Measured from record open to fully signed.

Confidence sortingre-check 100% → ~20%
Fields a human must touchTarget

Needs-review sorts to the top and calibrated-high collapses into quiet rows, so attention lands only on the roughly one-fifth of fields that need judgment.

Calibrated auto-acceptfixed line → per-field line
Analyst throughputTarget

The model handles the fields it is provably good at, so humans spend attention where it actually moves the number and the risk.

Audit trailnone → every value
Dispute-resolution timeTarget

AI-proposed versus human-final is logged per field, so a disputed estimate is defensible in minutes instead of a re-measure and an argument.

Correction loggingdiscarded → captured
Model accuracy over timeTarget

Every human edit feeds calibration, so the tool gets less wrong the more it is used. This is the flywheel that makes the product compound.

Validation method: log every AI-proposed versus human-final value and measure correction rate, calibration drift and time-per-record from telemetry, in a phased rollout that controls for confounders.

Reflection

What I would validate first, and where Verity goes next

The honest risk is that a confidence number is only as good as its calibration, so a well-displayed but miscalibrated score is the real product danger. I would validate first whether showing confidence actually calibrates trust or merely adds anxiety.

  • v2: capture-time QA hints so a poor capture is caught at the door, not at the desk.
  • v2: cross-scan change detection for reinspections, diffing a property against its last verified record.
  • v2: an active-learning loop that routes the model's weakest, most-corrected field types back into training.
  • v2: a field mobile companion for the second reviewer standing in the property.

Let's build something people remember

From enterprise teams to growing startups.

Let's talkarifin.yeasin@gmail.com