InsideMaps · Property Intelligence

Turning a finished scan into a record you can sign

An AI copilot that turns a finished 3D scan into a signed property record. The model drafts every number. A human signs it before it ships.

PlatformWeb workspace · restoration desk
SessionsLong focused desk reviews, one estimator
At a glance
The problemEstimators can't tell which AI-drafted measurements to trust, so they re-measure the whole property by hand and lose the hour the scan was meant to save.
The outcomeDesigned end-to-end and delivered to engineering for build. Target: a signable, verified record in under 15 minutes instead of 60–90.
My roleSole product & interaction designer, embedded with a PM, ML lead, and restoration SME.
TimelineInsideMaps · 2021
ApproachA trust system: machine-drafted vs. human-signed authorship, per-field confidence, and calibrated auto-accept.
Verity
62% verified
$18,240 of $29,410 confirmed7 need you
DollhouseFloor planMeasure
Scan IM-842132,140 sq ft · 4 rooms
LivingBed 1KitchenMaster Bath118.5 ft
1 · WATER_STAIN · master bath ceiling · 71%

The record workspace: a dark evidence twin on the left, a signable legal record on the right. Blue is machine-drafted, black is human-signed.

The number nobody would sign

A record full of numbers, and not one he trusted

4:40pm. Marcus, a restoration estimator, opens a delivered scan of a fire-and-water loss. The AI has already drafted the whole estimate: rooms, measurements, damage, a dollar total. It looks finished.

But he cannot tell what the software guessed from what it actually saw. So he opens the twin, re-measures three rooms by hand, and the hour the scan was meant to save is gone.

The tool had answered every question except the only one that mattered: which of these numbers can I actually stake my name on?

Business stakes

A claim is not paid on a 3D model, it is paid on a defensible number

Restoration and insurance do not buy 3D models. They buy a number that survives a dispute. The value, and the risk, is in the numbers derived from the scan.

45–90 minManual measurement per property, before structured data
~12%Estimate leakage one wrong measurement can introduce
100%Of AI fields re-checked, because none flagged which were wrong
~2×Cost when a bad number forces a recapture or disputed re-do
Before Verity

The naive AI export launched guesses as facts

A flat table of extracted fields: no confidence, no provenance, everything equally certain. It looked like automation and behaved like a liability.

Before · the manual way
IM-84213_scan.model - InsideMaps Viewer 2.1
FileEditViewMeasureHelp
feet ▾
23.8 ft
1512sf
Ready · Units: feet · Snap: OFFSel: Floor_Living
re-typed
by hand
XEstimate_IM84213_v3.xls - Excel
FileEditViewInsertFormatToolsData
100%
D11fx=SUM(D2:D10)
A
B
C
D
1
Room
Area
Line item
Amount
2
Living
512
Prime + paint, walls/clg
$1,840
3
Kitchen
486
LVP flooring + baseboard
$6,540
4
Bed 1
420
Drywall R&R, texture
$2,760
5
Master Bath
224
Water mitigation
$4,120
6
Masterr Bath
224
Dehumidifier + antimicrob
2830
7
Master Bath
224
Drywall R&R, prime + paint
$3,180
8
Hall
210
Smoke damage cleaning
$1,980
9
Exterior
288 lf
Baseboard R&R + paint
$3,340
10
Fixtures
Detach & reset
$2,820
11
TOTAL
est. re-keyed 06/14
$29,410
12
areas fromIM-84213_scan.model2← hand-entered
Sheet1EstimateSheet3
~45-90 min / property · 2 tools · every number re-keyed by hand
Two disconnected tools, and every number re-keyed by hand between them.
Before · the naive AI export
extraction_output.csvGET /api/v1/extract200 OK2023-11-04T09:12:37Z|15 fields|0.4 kb
#fieldvalue
1job_idIM-84213
2clientCascade Property Group
3scan_date2023-11-02
4total_area_sqft2140
5room_count4
6floor_area342
7ceiling_height9.0
8flooringhardwood
9water_damagetrue
10molddetected
11smoke_damagetrue
12drywall_replace_sqft214
13antimicrobial_sqft486
14dehumidifier_days3
15estimate_total29410
rows 1-15 of 15|confidence: n/a|source: n/a|schema: flat
The first pass at automation: a flat dump where a guess wore the face of a fact.
Six ways the raw export failed the person who had to sign it
  • 1Every value looked equally certain, so a guess wore the same face as a fact.
  • 2No value linked back to where in the scan it came from.
  • 3Correcting a wrong number meant re-measuring from scratch.
  • 4No record of who changed what, so nothing was defensible.
  • 5The model knew its own confidence, then threw it away before the screen.
  • 6The estimate lived in a separate spreadsheet, cut off from the evidence.
The real problem

The problem was not extract more, it was make every value verifiable

Distrust is not about how often the AI is wrong. It is about not being able to tell when.

  • 01

    Uniform-looking output forced total re-verification

    They don't distrust AI for being wrong. They distrust it because they cannot tell when it is wrong, so they re-check all of it.

  • 02

    A wrong value had no visible source

    With no link back to the scan, fixing a value meant re-measuring by hand. The AI made work instead of removing it.

  • 03

    The number could not be defended

    Without a record of what the AI proposed versus what a human finalized, a disputed number has nothing behind it.

IM-84213 · Living room · ceiling height

One measurement, three forms

A raw number is only a guess. Verity also carries where it came from, and who signed it.

01Raw model output
model.rawunverified
predict() →
9.0ft
float · p=0.63
no room·no source·not signed
Locate
02Source region
floor plan
IM-84213linked · reconstructed mesh
Living1z · 9.0 ft
Verify
03Human-corrected final
signed record signed
@marcus · on site
8.6ft
corrected 9.0 → 8.699%
human-ink·tape check·locked
a guessa locationa signed factOnly Verity keeps all three.

The same measurement is a guess, a location, and a signed fact. The raw export only ever showed the first.

01 · RESEARCH & DISCOVERY

The scan saved the hour. Distrust gave it right back.

I watched estimators scope real losses on site, then re-key every measurement into Xactimate back at the desk. The scan had already saved that hour. They spent it again by hand.

interviews
11 interviews

Restoration estimators, plus 3 carrier adjusters, on how a scoped number gets disputed

shadowing
6 ride-alongs

Scoping a live loss on-site, then re-keying it into Xactimate back at the desk

survey
n=41

Where estimators trust an auto-measurement and where they still re-check by tape

analytics
40 estimates

Exported scan tables audited line-by-line against the sketches they replaced

What discovery surfaced
01

Distrust is about when, not how often

They do not need fewer AI errors. They need to see which value might be one.

02

A guess wearing the face of a fact

A confident wrong number costs more than a blank one, so drafted values stay visibly drafted until signed.

03

Every number has to survive a carrier

An adjuster disputes it months later, so the audit trail is built at draft time, not reconstructed under challenge.

04

The saved hour leaks back out

Re-keying erases the savings, so the verified record has to be the deliverable, not a source to copy from.

Affinity map · from raw notes to five themes
Verity · Discovery synthesisEdited 2h ago
GARK+3
Share
Trust & confidence
“I can't tell when it's wrong”
The table looks the same whether right or wrong
Re-measures every wall the scan already gave
“One bad number and I distrust the whole file”
Trusts the scan for big rooms, not closets
Show confidence per value, not per file
Insight

Estimators can't tell a fact from a confident guess, so confidence has to be calibrated and shown per value, not per file.

Defending the number
“I have to defend this to the adjuster”
No way to prove where a number came from
Screenshots the 3D view to back a line
Disputes drag when there's no evidence
“It's my name and license on the estimate”
Link every line item to its scan region
Insight

Every value gets challenged by a carrier eventually, so each one links to its source region and lands in an audit trail from the start.

The re-keying tax
“Then I type it all into Xactimate again”
Sketches on-site, re-keys at the desk
45 to 90 minutes per property, re-entered
Copy-paste means transcription errors creep in
Rounds measurements up to feel safe
Draft straight into estimate line items
Insight

The scan's saved hour leaks back out at the keyboard, so drafting has to land directly in estimate line items.

Reading the damage
“Is that whole wall wet or just the base?”
Walks the room twice to catch missed damage
Auto-tools miss a closet or a bay window
“Carpet or vinyl changes the whole line”
Pins photos to each affected area by hand
Glowing regions plus defect pins to confirm
Insight

Damage extent and material are judgment calls, so the AI surfaces candidate regions and defect pins for a human to confirm, never decide.

The auto-accept line
“Just don't make me check the obvious ones”
Speed-signs big rooms, slows on damage calls
Reviewing every value is as slow as re-doing it
“I trust it more when it flags its doubts”
Move the line: auto-accept above, review below
Insight

Estimators want to skim safe values and dwell on risky ones, so a movable auto-accept line lets them tune what still needs a signature.

Gazi
you
AFFINITY MAP · 5 THEMES
Legend
Voice of user
Observed behavior
Pain point
Opportunity
11 interviews · 6 ride-alongs · 29 notes clustered
−100%+
Design-target personas
ES
The estimator
scopes the loss · signs the number
What they need

To see, per value, where the number came from and how sure the machine is, before I put my name on it.

Goals
  • Scope a loss without re-measuring every wall by hand
  • Sign only numbers I can defend to a carrier
  • Turn the scan into an estimate, not a re-typing job
Frustrations
  • A flat AI table hides which numbers to trust
  • One wrong square-foot and the carrier kicks it back
CA
The claims adjuster
represents the carrier · disputes the number
What they need

A record where every measurement and finding links to the scan region it came from, so I can accept or challenge on evidence.

Goals
  • Verify a scope without a second site visit
  • Trace any line item back to what it is based on
  • Close the claim without over- or under-paying
Frustrations
  • Contractor estimates arrive with no evidence attached
  • Disputes drag on when nobody can show the source
PM
The restoration PM
owns throughput · owns liability
What they need

A signed, auditable record per property that speeds the estimator up without putting the firm on the hook.

Goals
  • Get estimators through more properties per week
  • Ship estimates that don't bounce back in review
  • Keep a defensible record if a claim is challenged
Frustrations
  • Re-keying into Xactimate eats billable hours
  • A bad number becomes the company's liability

Design-target archetypes, drawn from the interviews, ride-alongs and the problem space.

The reframe

Verity is not an extractor. It is a trust instrument. Its job is to make uncertainty legible and correction nearly free, so a human's attention goes only to the values that actually need judgment.

Two rules fell out of that, and I never broke them: confidence is always shown, and every value links to the region that produced it.

Constraints & principles

The box I designed inside

The model will be wrong, and its output feeds a legally disputable document. Five principles fell out of that box.

01

Show confidence, never hide it

Hiding uncertainty to look clean is a lie the user pays for later.

02

Provenance, or it does not ship

An unlinked number cannot be corrected fast, or defended at all.

03

Correction is one gesture, not a form

A correction flow that fights the analyst's speed gets its defaults slammed.

04

The human's edit is both the audit record and the training signal

One action serves the dispute file and the calibration loop. Nothing is thrown away.

05

Automate acceptance only above a calibrated line

Earned per field type against a real reliability curve, not a threshold that flatters the demo.

Rejected directions

Three versions I did not build

Judgment shows in what you refuse. Each of these was tempting, and each would have quietly broken the thesis.

01

The fly-through-first canvas

no record · nothing to sign

An immersive 3D walkthrough as the main surface. It demos beautifully.

Why I killed it: The job is verifying numbers, not flying through space. The 3D became a focused evidence pane, not the product.

02

Full auto-accept, no review

Floor areaauto
Ceiling heightauto
Wall areaauto
Signed bythe model
nobody to stand behind it

Let the model write the estimate and skip the human. The cleanest leverage story.

Why I killed it: A disputed number nobody signed is indefensible. The human stays on the values that carry money and risk.

03

A chat-only estimate box

What is the scope for this loss?
Replace 284 sf of flooring and 612 sf of drywall, totalling $8,240.
no source · no way to verify a number

Type a question, get a scope. Modern and effortless.

Why I killed it: Nothing to verify against. Q&A became a cited layer on the record, not a replacement for it.

Information architecture

The system before the screens

The audit trail is an object, not a log line. Every value knows where it came from.

01Object modelEvery derived value points at the region it came from, and logs who signed it.
produces 1:1contains 1:nholds 1:nevery derived valueconfirmed
Capture
scan_iduuid
device · captured_atmeta
PropertyRecord
addressstring
statusenum
signed_by · signed_atactor
Room
name · typestring
floorint
Measurement
kindenum
value · unitnumber
confidencefloat
authorshipmachine|human
DamageFinding
perilenum
extentnumber
confidencefloat
authorshipmachine|human
MaterialDetection
materialenum
areanumber
confidencefloat
authorshipmachine|human
EstimateLineItem
code · qtystring
unit_cost · totalmoney
justified_by→ DamageFinding
dollar_stateai|confirmed
SourceRegionProvenance
bbox_3d | image_cropgeometry
frame_ref→ Capture
model · versionstring
confidencefloat
VerificationEventAudit + training
actor · atwho · when
old_value → new_valuedelta
actionaccept|edit|flag
feedscalibration

Scroll sideways to read the full diagram →

02Value lifecycleA value that clears its calibrated line signs itself. Everything else routes to a human.
clears its calibrated linebelow the line
Extractedmachine drafted
Needs-reviewwaiting on a human
Confirmedsigned as drafted
Correctedsigned, value changed
Flaggedescalated, unsignable

Scroll sideways to read the full diagram →

USER FLOWS · THE ESTIMATOR'S DESK

Every number, signed or sent back.

The machine drafts every value in blue. Marcus signs only what he will defend in black. These are the two paths he walks at his desk.

Start / endMarcus actsVerity draftsDecision
Flow 01

Sign it, or send it back

Clearing the record one value at a time.

  1. Open the property record
  2. Every value drafted, with calibrated confidence
  3. Clears the auto-accept line?Below it: open the inspector
  4. Drill to the source region
  5. Does the scan back the number?No: correct it by hand
  6. Sign it, the value locks black
  7. Room signed, written to the audit trail

Anything below the line, or unbacked by the scan, re-enters the loop until he can sign it.

Flow 02

Findings into a defensible figure

Turning confirmed damage into a number a carrier cannot wave off.

  1. Enter damage & material review
  2. Glowing regions and defect pins surfaced
  3. Confirm or reject each pin
  4. Ledger links each line to its finding
  5. Confirmed dollars, or still AI-estimated?AI-estimated: open the finding
  6. Verify in the twin, adjust the line
  7. Sign the estimate
  8. White-label record sent to the carrier

No AI-estimated dollar reaches the carrier without being verified first.

Design system for trust

Color encodes authorship and certainty, never status

Two scales, and they are never allowed to collide: who authored a value, and how sure the model is.

01Two surfaces, two jobs

Evidencedark · the machine shows its work

Geometry, glow and the probe live here. Nothing signable does.

Recordlight · the human signs

The legal document. Every value here is signable and disputable.

02Authorship, the law

One question, asked of every value on screen: has a human signed this?

Machine-ink#1D4ED8Drafted by the model. Never signable in this state.
Human-ink#14181FA person put their name on it. Defensible.

03Confidence, a scarce budget

A separate scale, spent only on how sure the model is. Each tier names an action.

High≥ 88%Auto-accept eligible, if calibration earns it.
Review70 – 87%Look closer before signing.
Low< 70%Verify against the source, or recapture.

04One value, every state it passes

Drafted
9.0 ft
machine authored
→
Needs review
9.0 ft
below the line
→
Editing
8.6|
one keystroke
→
Corrected
8.6 ft
old to new logged
→
Signed
8.6 ft
human authored

Scroll sideways to see every state →

The rule

Color answers who authored this and how sure the model is. It never answers what stage the job is in. The moment color starts carrying workflow status, both of the scales above stop meaning anything.

Wireframe evolution

The record workspace, three fidelities deep

The record workspace carries the whole trust idea, so it took the most iteration: five paper screens each pinned to one open question, then a greyscale pass that locked confidence and provenance before color could flatter a weak decision.

Lo-fi · paper
Not about layout. About the one question every screen had to answer: where does a value's evidence live, and how does a human sign it?
Mid-fi · greyscale
1Confidence ring on every value2Focus flies pane to source3Blue drafted, black signed
Greyscale forced the hard calls before color could flatter them: a ring on every value, focus that flies to the source, and a machine-versus-human split.
Hi-fi · shipped
Verity
62% verified
$18,240 of $29,410 confirmed7 need you
DollhouseFloor planMeasure
Scan IM-842132,140 sq ft · 4 rooms
LivingBed 1KitchenMaster Bath118.5 ft
1 · WATER_STAIN · master bath ceiling · 71%
The shipped record workspace: a dark evidence twin left, a signable legal record right. Blue is machine-drafted, black is human-signed.
Lo-fi · paper

Five screens, one question each

Not about layout. About the one question every screen had to answer: where does a value's evidence live, and how does a human sign it?

Mid-fi · greyscale

Confidence and provenance, before color

Greyscale forced the hard calls before color could flatter them: a ring on every value, focus that flies to the source, and a machine-versus-human split.

Hi-fi · shipped

Dark evidence, light record

Dark evidence left, light legal record right, machine-blue settling to signed black. The color law does the trust work the wireframe only promised.

Marquee flow

Deep dive: verifying an AI-drafted record

Sixty drafted values on one screen. The design's whole job is making any one of them a glance to check and a keystroke to fix.

Verity
62% verified
$18,240 of $29,410 confirmed7 need you
DollhouseFloor planMeasure
Scan IM-842132,140 sq ft · 4 rooms
LivingBed 1KitchenMaster Bath118.5 ft
1 · WATER_STAIN · master bath ceiling · 71%
Verify to black. Blue is what still needs a human. When the record is all black, it is signed.
The drill-down
Floor plan Measure
Scan IM-84213source region · living room342 sq ft
LivingKitchen18.5 ft1z · 9.0 ft
Focus a value and the evidence pane flies to the frame it was read from.
The differentiator

Deep dive: damage and material detection review

Damage extent is a judgment call, not a measurement. The model proposes, the estimator adjudicates, and only a confirmed finding is allowed to become money.

Floor planDamageMaterials
Scan IM-842132,140 sq ft · 4 rooms · floor 1
KITCHEN382 sq ftLIVING521 sq ftBED 1338 sq ftHALL96 sq ftMASTER BATH144 sq ft12 sq ft21.7 ftN123
1 · WATER_STAIN · master bath · ceiling · 71%
Walk the pins, adjudicate each finding. Only a confirmed one becomes a line item.
  • Walk the pins. Each numbered damage region opens a finding card: the highlighted region, a source photo crop, and the AI read (for example, water stain, 12 sq ft, moderate, 71%).
  • Adjudicate in one gesture. Confirm, adjust boundary, reject, or reclassify. The AI proposes, the human decides, and the decision is logged.
  • Findings become money. A confirmed water stain auto-maps to line items (drywall replace plus paint), so spatial evidence, AI, and the estimate move in one motion.
The signed output

Turning verified findings into a defensible scope

This is the document a carrier actually receives, so every dollar on it has to trace back to the finding that justifies it.

Verity
62% verified
$18,240 of $29,410 confirmed3 need you

Restoration estimate

Output · R1

Cascade Property Group · IM-84213 · fire & water

EstimateEST-84213-R1
Issued2026-07-24
Scope2,140 sq ft · 4 rooms
Estimator@marcus
Line itemQtyConfAmount
Water mitigationWTR
Subtotal$3,120
Dehumidifier, LGRWATER_STAIN · Master Bath
4 days signed$980
Air mover, axial fanWATER_STAIN · Kitchen
8 units · 4 d signed$1,410
Antimicrobial applicationMOLD · Master Bath
620 sq ft signed$730
DrywallDRY
Subtotal$8,940
Drywall R&R, 1/2 inCHAR · Living
1,240 sq ft signed$6,820
Texture, knockdown matchderived · wall area
1,240 sq ft91%$2,120
PaintPNT
Subtotal$6,180
Seal, stain-block primeSMOKE_DAMAGE · whole-home
1,860 sq ft signed$2,240
Paint, 2 coatsderived · wall area
1,860 sq ft78%$3,940
FlooringFLR
Subtotal$11,170
LVP flooring, supply + installMISSING_FLOORING · Kitchen
940 sq ft signed$6,060
Baseboard, 5.25 in, R&R + paintDELAMINATION · Bed 1
412 LF66%$5,110
Human-confirmed
$18,240
Still AI-estimated
$11,170
Total estimate
$29,410
Export XactimateExport PDFPDF includes full AI-proposed vs human-final audit trail
The total splits into dollars a human confirmed and dollars still AI-estimated.
Verityimmutable · export-ready

Audit trail · IM-84213

38 events · Jul 25 2026 · 3:58p to 5:02p · machine-authored and human-signed entries, append-only

All events38Edits9Flags2Exports1
5:02p
MV@marcus
Record exported to Xactimate (.ESX) · total $29,410 · $18,240 signed
export
4:58p
MV@marcus
Signed Living room measurements · 4 fields (floor area, perimeter, ceiling, flooring)
verify
4:51p
MV@marcus
Flagged Kitchen for recapture · defect CHAR extent unclear from mesh
flag
4:47p
MV@marcus
AI proposed 340 sq ft at 82% · edited to 328 sq ft by @marcus
edit
value correctedmachine draft to human sign340 sq ft328 sq ft
Field
Living room · Floor area
Model
verity-visionv4.2
Source region
LR-04 · 1,284 mesh faces
Confidence before
82%review
Confidence after
human-signed
Actor
MV@marcus · Marcus Vega
Reason

Model included the bay-window recess in floor area. @marcus trimmed to the interior wall face and re-signed the value.

event 0x4f· sha256 3f9a…c21 · sealed 4:47:12p
4:39p
MV@marcus
Accepted ceiling height 8.6 ft · Master Bath, no adjustment
verify
4:33p
AI
Detected MOLD in Master Bath · 12 sq ft at 68%
detect
4:12p
AI
Drafted record from scan · 60 values across 4 rooms · estimate $29,410
draft
3:58p
SYS
Scan IM-84213 ingested · 2,140 sq ft reconstructed · 4 rooms · fire & water loss
scan
hash chain verified · head 0x8c21af · no gaps · WORM storageCopy sealExport log (CSV)
The audit trail is an object, not a log line: AI-proposed versus human-final, per value.
Cascade RestorationFire & water restoration
Public verified recordPowered by Verity
Verified Property Record

1428 Camino Real

San Jose, CA 95126·IM-84213·Cascade Property Group

Every figure below was drafted by AI and signed by a licensed estimator.

AI-drafted Human-signed
Verified estimate total
$29,410
Signed by Marcus Vega
Licensed estimator · @marcus · 07/24/2026

Verified measurements

2,140 sq ft · 4 rooms scanned
Total floor area
2,140sq ft
signed
Rooms scanned
4rooms
signed
Ceiling height
9.0ft
signed
Affected area
128sq ft
signed

Damage findings & scope

6 findings · 3 rooms
WATER_STAINMaster bath ceiling
Drywall R&R · texture · prime + paint
$4,860
SMOKE_DAMAGEKitchen walls
Antimicrobial · prime + paint
$3,220
MISSING_FLOORINGLiving room
LVP flooring · baseboard
$6,540
MOLDMaster bath subfloor
Antimicrobial · dehumidifier · air movers
$2,340
CHARKitchen cabinetry
Cabinetry R&R · delamination repair
$2,180
12 line items total · every figure signed on-siteTotal $29,410
One click turns the signed record into a white-label deliverable.
Designing control over the AI

The calibration layer

A confidence score is worthless if it is not calibrated. This layer keeps the score honest, then hands the dial to the human.

Verity
Rolling 30 daysMVMarcus Vega@marcus

Model calibration

trust quality · last 5,000 verified fields

ECE 4.1%Brier 0.038
0%25%50%75%100%0%25%50%75%100%OVERCONFIDENT88%PREDICTED CONFIDENCEOBSERVED ACCURACY
Auto-accept threshold
≥ 88%
42%auto-accepted
Review load -38%
Legend
Model
Perfectly calibrated
Overconfidence gap
Auto-accept zone
When the model says 90%, is it right 90% of the time? The line is the human's to move.
Grounded Q&A

Ask this property, and only trust what it can cite

Answers you can check. It cites the regions it read, and refuses outright when the record cannot support an answer.

Verity Ask
62% verified
$18,240 of $29,410 confirmed7 need you
grounded regions
2 citedFloor planSecond floor
Scan IM-84213Second floor · drywall query
Bed 1Master BathBath 2Hall16.0 ft9.4 ft12
grounded in
1Bed 1 · walls2Master Bath · walls + ceiling
It shows what it will read, cites what it used, and refuses when the record cannot support an answer.
Failures & recovery

Three things I got wrong

Every principle above was forged by a failure below. This is where the design actually came from.

The pretty lie

I built
Floor area284 sf
Ceiling height9.0 ft
Wall area612 sf
The fix
Floor area284 sf
Ceiling height9.0 ft
Wall area612 sf
It broke

A vaulted ceiling, the value the model is least sure of, rendered exactly like a floor area it nails. The interface had laundered a guess into a number.

I learned

Confidence has to be visible AND calibrated. Visibility alone adds anxiety, calibration alone stays invisible.

The modal tax

I built
Ceiling height9.0 ft
Wall area612 sf
Correct value
8.6
Reason for change
Save~11s
The fix
Ceiling height8.6|
Eedit, logged, signed
It broke

Eleven seconds an edit. Analysts batch corrections, then skip them and accept the defaults to move on.

I learned

A correction path that fights the analyst's speed loses, and takes the audit trail down with it.

The uncalibrated threshold

I built
Floor area
Wall area
Ceiling ht
Trim length
one 80% line for everything
The fix
Floor area
Wall area
Ceiling ht
Trim length
a line each, earned on reliability
It broke

One line over-accepted vaulted ceilings, the model's weakest field, while holding back floor areas it nails.

I learned

Auto-accept is a per-field decision against real reliability, and the human should own how aggressive it gets.

Prototyping & validation

What broke when a real estimator used it

I sat a working restoration estimator through the prototype on a real delivered scan. Every time he reached for his tape instead of the screen, the design was wrong. Three of those moved it.

Track 01 · Estimator think-aloud, on a real scan
Done
01

A doubt with nowhere to go

Friction
Ceiling height9.0 ft
re-measured by hand

A low ring said doubt this, then offered no way in. He reached for his tape.

Change
Ceiling height9.0 ft
FRAME 41

The ring became the door. It flies the evidence pane to the frame the value was read from.

02

A signature covering work he never saw

Friction
Signature
Marcus Reyes47 values
40 never opened

One signature covered forty auto-accepted values, so he started opening all of them.

Change
Signature
Marcus Reyes7 touched
Model · calibration v4.240 auto

Split the ink. His name lands only on values he touched, the model signs its own.

03

Damage drifting into dollars

Friction
Water damage · living room$2,480
posted · source never opened

A detected water line posted a line item before anyone looked at it.

Change
Water damage · living room$2,480
open the source frame to post

A finding cannot post dollars until its source frame has been opened once.

The pivot

A low-confidence value stopped meaning re-measure and started meaning a ten-second look at the source. That is the hour the scan was always supposed to save.

Track 02 · Moderated study, five estimators
Scoped

Will an estimator stop re-measuring what Verity marks high-confidence, still catch what it marks low, and sign a number he would defend to a carrier?

Five working estimators, none who have seen Verity. Sixty-minute moderated sessions on one real delivered scan.

0102030405
0 min5 tasks, think-aloud60 min
Task script
  1. 01Reach a number you would defend, re-measuring only what you distrust.
  2. 02Decide whether to trust a low-confidence vaulted ceiling.
  3. 03Confirm you would sign forty auto-accepted fields.
  4. 04Show a carrier what stands behind your wall-area figure.
  5. 05Move the auto-accept line to clear the next property faster.
Success bar
Pre-registered
Low-confidence values drilled5 / 5
High-confidence fields re-measured≤ 1 / property
Wrong values signed0
Time to a value's source< 10s
Would defend to a carrier≥ 4 / 5
Collaboration & tradeoffs

Who I decided with, and what each call cost

None of these calls were mine alone, and none of them were free.

Restoration estimators

Co-defined the damage taxonomy and what counts as a finding.

ML lead

Set where confidence comes from and what the score is allowed to claim.

Product manager

Cut MVP scope with me, desktop review first.

Provenance on every field

9.0 ft
Bare value
9.0 ftFrame 41
Carries its source
Gave up

Compute and storage. Every value stores a link back to its source region.

Bought

A number that corrects in seconds and survives a dispute.

Conservative auto-accept

REVIEWAUTO70
Loose line · 70
REVIEWAUTO88
Calibrated · 88
Gave up

Throughput. More fields route to a human than the model strictly needs.

Bought

Fewer shipped errors. In restoration, an error becomes a dispute.

Desktop first, mobile in v2

Field mobile · v2
Desk workspace · v1
Gave up

The field companion app.

Bought

Focus on the desk, where the estimate is signed and sent.

Projected / target outcomes

The targets I designed toward, and how I validate them

Design targets, not shipped results. Each pairs a mechanism I can point to in the interface with the outcome it is meant to move.

Time to a verified recordTarget
60–90 min today
under 15 min
Fields a human must touchTarget
100% re-checked
~20%
Leading mechanism (in the design)Lagging outcome (the business result)
Provenance-linked verify60–90 min → under 15 min
→
Time-to-verified-recordTarget

A glance and one key, instead of a manual re-measure.

Confidence sortingre-check 100% → ~20%
→
Fields a human must touchTarget

Needs-review sorts to the top, calibrated-high collapses into quiet rows.

Calibrated auto-acceptfixed line → per-field line
→
Analyst throughputTarget

The model takes the fields it is provably good at, nothing else.

Audit trailnone → every value
→
Dispute-resolution timeTarget

A disputed estimate becomes defensible in minutes, not an argument.

Correction loggingdiscarded → captured
→
Model accuracy over timeTarget

Every edit feeds calibration. The tool gets less wrong the more it is used.

Measured from telemetry: correction rate, calibration drift and time-per-record, in a phased rollout.

Reflection

What I validate first, and where Verity goes next

A well-displayed but miscalibrated score is the real danger. First test: does showing confidence calibrate trust, or just add anxiety?

!
Low light · recapturecaught on site, not at the desk
v2

Capture-time QA hints

Catch a poor capture at the door, not at the desk.

PRIORNOWDIFF
v2

Cross-scan change detection

Diff a property against its last verified record for reinspections.

CORRECTIONMODELWEAKEST FIELDS RETRAIN
v2

An active-learning loop

Route the model's most-corrected field types back into training.

verify on site, beside the damage
v2

A field mobile companion

For the second reviewer standing in the property.

Gazi Yeasin ArifinTechnical UX Designer at InsideMaps Inc.

Let's build something people remember

From enterprise teams to growing startups.

Let's talkarifin.yeasin@gmail.com