InsideMaps · Operations Command Center

Running a nationwide 3D-capture network in real time

PulseOps is the internal command center for running InsideMaps' nationwide 3D-capture network: a live capture-to-delivery pipeline, a 3D-model QA workflow, fleet utilization, and multi-tenant SLA analytics. Designed to double delivery volume without doubling the humans watching it.

RoleTechnical UX Designer
PlatformInternal web · operator workstations
At a glance
The problemCoordinators rebuilt the state of a nationwide capture network from three tools every shift, with no view of what was about to breach SLA until it already had.
The outcomeDesigned and handed to engineering for implementation. Target: time-to-triage from 15–20 min to under 2, at 2–3× volume without adding coordinators.
My roleSole designer, with a PM, eng lead, and ops lead.
TimelineInsideMaps · 2021
ApproachA manage-by-exception command center: one risk-ranked triage queue, per-tenant health, and calm real-time.
Capture Pipeline
77 active jobs
Updated 3s ago
⌘K
6
Filtersstatus: active ×SLA: at-risk ×type: residential ×77 jobs sorted most-at-risk first
Capture12
IM-8428018h 40m
Meridian Homes
2,100 sqft · 8mD
IM-842766h 12m
Redwood Realty
3,400 sqft · 17mM
IM-842713h 30m
Sentinel Insurance
1,850 sqft · 31mD
+ 9 more
Upload9
IM-842582h 48m
Cascade Property
5,200 sqft · 12mP
IM-842495h 10m
Meridian Homes
2,700 sqft · 5mD
+ 7 more
Processing24WIP limit 20
IM-842132h 14m
Redwood Realty
4,200 sqft · 41mM
IM-839982h 14m
Cascade Property
6,050 sqft · 38mP
IM-841904h 02m
Sentinel Insurance
3,300 sqft · 22mD
IM-841767h 20m
Meridian Homes
1,900 sqft · 9mD
+ 20 more
QA183 breaching
IM-841090h 52m
Redwood Realty
3,800 sqft · 1h 12mP
IM-840771h 36m
Cascade Property
2,600 sqft · 48mP
IM-840513h 10m
Sentinel Insurance
4,900 sqft · 26mM
+ 15 more
Delivered214
IM-84002delivered
Meridian Homes
2,300 sqft · on timeD
IM-83981delivered
Redwood Realty
3,100 sqft · on timeM
+ 212 more
Flow · arrivals vs completions todayProcessing net +9 backlog
Capture
Upload
Processing
QA
Delivered

The live pipeline board: every job flowing capture to delivered, sorted most-at-risk first.

The 90 minutes that started this

A job went late while everyone was watching

2:10pm. Renata, a coordinator, is toggling between three browser tabs. A Cascade Property Group order is ninety minutes from its SLA and she does not know it. The job finished processing at 11am and has sat in an unwatched QA queue ever since.

She finds out the way she always does: the account manager pastes a red-faced client message into Slack. The truck already went home. A recapture now means the deliverable is a day late at twice the cost. Nobody did anything wrong. The tools simply could not show what was stuck.

The failure was not a person. It was that no screen in the company could answer one question: what is about to be late, right now?

Business stakes

What it takes to run a capture network

A logistics, GPU-compute and human-QA machine wrapped in a 24 to 48 hour promise. The product is the promise, not the model file.

24–48hContractual turnaround, tighter on enterprise rush tiers
~35 miCapture radius per technician: the network is a patchwork of local zones
~12%Of project value lost to rework when a capture is poor
~2×Cost of a recapture: the truck rolls again and the SLA is blown
Before PulseOps

Three tools, none of which could see a deadline

A legacy admin panel, a shared spreadsheet, and a wall of Slack pings. Status was a color word with no clock behind it.

insidemapsOrder AdminA
San Jose — Orders(market: San Jose only)
Order ID
Customer
Address
Status
Scanner
Created
Due date
Actions
IM-84213
Redwood Realty
1428 Camino Real
Processing
mrodriguez
06/28/2020
6/29/2020
Edit
IM-84109
Redwood Realty
88 Winchester Blvd
In QA
pshah
6/28/20
6/29
Edit
IM-83998
Cascade Property
2100 Broadway
processing
pshah
6/27/2020
6/28/2020
Edit
IM-84077
Meridian Homes
512 Alameda
Processing
dana.p
6/28/20
6/30/2020
Edit
IM-84051
Sentinel Insurance
9 Mission St
scheduled
mrodriguez
5/7/21
—
Edit
IM-84044
Cascade Property
41 K Street
PROC.
sam.t
06/28/2020
6/29/2020
Edit
IM-84012
Meridian Homes
512 Alameda
Delivered
dana.p
6/27/20
6/29/2020
Edit
IM-83981
Redwood Realty
77 Park Ave
delivered
mrodriguez
6/26/2020
6/28/2020
Edit
Rows per page: 10 ▾1–10 of 2,481‹ ›
A stock CRUD table: free-text colored status, no SLA, no time dimension anywhere.
▦
SLA_TRACKER_v7_FINAL_use-this-one ☆
File Edit View Insert Format Data Tools Extensions Help
DM
G3ƒx=TEXT(F3-NOW(),"h") → #REF!
MASTER — do not edit without asking Dana. do NOT sort (breaks the HRS LEFT formulas). West tab is ~2-3h behind by afternoon.
·
A
B
C
D
E
F
G
H
I
J
K
1
Job ID
Client
Market
Tech
Ordered
Due
HRS LEFT
Stage (typed)
Last checked
PROBLEM?
Notes
2
IM-84213
Redwood
San Jose
Marcus
06/28/2020
6/29 2pm
#REF!
Processing
11:04a
??
system still says Processing but AM says it's late
3
IM-84109
Redwood
San Jose
Priya
6/28/20
6/29
0.9
in QA
10:40a
yes
blurry, might need recapture
4
IM-83998
Cascade
Oakland
Priya
6/27/2020
6/28 EOD
-1.2
proc.
9:15a
YES
LATE. client emailed. escalated to Sam
5
IM-84077
Meridian
San Jose
Dana
6/28/20
6/30
18.4
PROCESSING
8:50a
no
6
IM-84051
Sentinel
Oakland
Marcus
5/7/21
?
#VALUE!
scheduled
—
?
no due date in contract PDF, asked account mgr
7
IM-84044
Cascade
Sacramento
Sam
06/28/2020
6/29 10a
2.1
capture?
9:55a
?
called tech, no answer x2
8
IM-84012
Meridian
San Jose
Dana
6/27/20
6/29
0
DONE?
yest.
no
delivered I think, double check
9
IM-83981
Redwood
San Jose
Marcus
6/26/2020
6/28
done
delivered
6/28
no
10
IM-83975
Sentinel
Oakland
Leo
6/28/20
6/30 noon
26.5
upload
10:20a
no
11
12
13
+West (MASTER)MiamiNashvilleCharlotteDO NOT DELETEoldSheet14
The coordinator's private spreadsheet: hand-typed status, hand-maintained hours-left, the real source of truth.
Every later decision pays back one of these six failures
  1. 1Status was arbitrary color words with no time dimension.
  2. 2One word, “Processing,” hid capture, upload, compute and QA.
  3. 3There was no SLA anywhere in the interface.
  4. 4Every view was a single-market silo.
  5. 5It was built to edit a record, not resolve a problem.
  6. 6The answer to every hard question was “export to a spreadsheet.”
The real problem

The problem was invisible wait time, not a cluttered dashboard

The private spreadsheet was the tell. The real system of record lived outside the software.

  1. 01

    Work stalled invisibly between stages

    A job finished upload and waited on a free GPU slot; finished processing and waited on a free QA reviewer. Most elapsed time was wait time, not touch time, and no tool measured where jobs sat.

  2. 02

    Exceptions surfaced after the SLA was already blown

    Failed uploads, compute crashes, no-shows, blurry captures: all discovered reactively, often by a client complaint. Easy jobs got cleared while at-risk jobs quietly aged past their deadline.

  3. 03

    Coordination scaled linearly with volume

    Every job needed manual watching, status chasing, and hand assignment. That is the exact constraint that stops a B2B ops org from growing revenue without growing headcount in lockstep.

In the legacy Admin
OrderCustomerStatusDue
IM-84213RedwoodProcessing6/29
No time dimension.
Open the record and do the math yourself.
In PulseOps
Ordered41h ago
Captured38h ago
Uploaded36h ago
Processingstuck 34h · COMPUTE_FAIL 6m ago
QAblocked
Deliverydue in 2h 14m

The data always existed. The legacy tool just could not show it in time to act.

Research & Discovery

The network kept no clock, so I timed it by hand

No tool at InsideMaps recorded when a job entered or left a stage, so "slow turnaround" was a feeling nobody could locate. I spent discovery on the operator workstations themselves: shadowing full shifts, interviewing every role that touches a job, and reconstructing real stage timings from the raw order log. One pattern held across all of it, and it reframed the project from "clean up the dashboard" to "make the network watch itself."

shadowing
6 shifts

Sat beside coordinators through full 8-hour shifts in three markets, tallying every tab-toggle and Slack check that rebuilt the network's state.

interviews
11 sessions

One-on-one across all five roles: coordinators, QA reviewers, a reconstruction engineer, the ops lead, and two account managers who own the QBR.

analytics
18 months

Reconstructed real per-stage dwell times from the raw order log, since no tool had ever recorded when a job actually moved between stages.

workshop
2 sessions

Mapped the true pipeline states and defect codes with ops and eng, splitting one word, "Processing," into capture, upload, compute, and QA.

What discovery surfaced
01

Most of the clock is wait time

Jobs spent hours queued between stages, not being worked, and "Processing" hid four different waits, so the target was dwell time, not touch time.

02

Every shift began from zero

With no time dimension in status, coordinators spent 15-20 minutes each shift rebuilding what was urgent before they could act on anything.

03

Alarms arrived as complaints

Exceptions were discovered reactively through a client's Slack message, always after the SLA had already been lost and the truck had gone home.

04

The average hid the failing account

A healthy global SLA number masked individual enterprise tenants quietly breaching, and per-market silos hid the breach from the account managers who owned it.

Affinity map · from raw notes to five themes
PulseOps · Discovery synthesisEdited 2h ago
GARK+3
Share
The invisible wait
“Processing could mean four totally different things”
Pings the compute eng to ask if stuck
No screen shows where a job sits
Easy jobs clear while at-risk ones age
Re-checks the same job every twenty minutes
Put time-in-stage on every card
Insight

Wait time was untracked and unlabeled, so every stage needs a visible time-in-stage clock and aging heat.

Rebuilding state every shift
“First I figure out what's on fire”
Rebuilds the picture from three tabs at login
Status is a color word, not a clock
Twenty minutes gone before the first action
Keeps a private sheet the software should be
Land on what's about to breach, ranked
Insight

Coordinators reconstructed urgency by hand each shift, so the tool must land them on what is breaching.

Alarms arrive as complaints
“I find out when the client's already angry”
Hears of a no-show from Slack, not the tool
Failures surface after the SLA is gone
Truck's already home when we catch it
Catch the exception before it breaches
One ranked queue, deduped by root cause
Insight

Exceptions surfaced only after breach, so detection has to move upstream and rank itself by root cause.

QA under the clock
“Reject reasons live in my head, not data”
Types free-text notes nobody can aggregate
Fix-or-recapture is a money call made blind
A slow rubric just gets defaulted through
Keyboards through the queue to keep pace
One defect code feeds the whole loop
Insight

Review had to stay fast while producing analyzable reasons, so one defect code rides with the decision.

Numbers that can't explain themselves
“We're at 88, so why is Cascade furious?”
Exports to a spreadsheet to answer anything
A global average hides a breaching tenant
Every market is its own blind silo
Track SLA health per tenant, not overall
Attribute the slowdown to a single stage
Insight

Reporting could show a slowdown but never attribute it, so every metric must drill to a stage and a tenant.

Gazi
you
AFFINITY MAP · 5 THEMES
Legend
Voice of user
Observed behavior
Pain point
Opportunity
Affinity map · shift-floor notes clustered into five themes
−100%+
Design-target personas
RV
Renata Vaughn
Ops coordinator · protects the SLA clock
What they need

A ranked landing surface that says what to touch first, right now.

Goals
  • See what is about to breach before it does
  • Clear the shift's fires without chasing three tools
  • Spend attention only on jobs that need judgment
Frustrations
  • Every shift starts by rebuilding state by hand
  • Finds out a job is late from an angry Slack
TM
Theo Marsh
QA reviewer · accept, fix, or recapture
What they need

A keyboard-first decision where one defect code carries the signal.

Goals
  • Judge a 3D model fast without cutting corners
  • Make the fix-vs-recapture call with cost in view
  • Leave a reason the rest of the system can use
Frustrations
  • A heavy rubric just gets defaulted through
  • Reject notes are free text nobody can aggregate
DO
Dana Okafor
Ops lead · owns SLA and throughput
What they need

Drillable metrics that pin a slowdown to a stage and a tenant.

Goals
  • Double delivery volume without doubling coordinators
  • Point at the real bottleneck, not a gut feeling
  • Defend enterprise SLAs per account, not on average
Frustrations
  • A healthy global average hides a failing tenant
  • Every answer ends in export to a spreadsheet

Design-target archetypes, drawn from the interviews, ride-alongs and the problem space.

The reframe

The problem was never a cluttered dashboard. It was invisible wait time and reactive triage. So the goal was not to show more; it was to make the network watch itself, and interrupt a human only for the ~10% of jobs that need judgment.

Manage-by-exception is not a UI pattern here. It is the business model: the only thing that breaks the linear-headcount curse. And because a capture network is a queueing system, utilization has a ceiling on purpose: the tool has to refuse to reward over-driving.

Constraints & principles

The box I designed inside

Four facts ruled out most of the obvious designs, and generated the five principles below.

01

Density with hierarchy, not minimalism

Operators live in this tool for eight-hour shifts. Information density is a feature. The craft is managing it with a strict altitude system, not hiding it behind whitespace.

02

Exceptions surface themselves

Healthy work collapses into count chips. Only the anomalous earns its own row. The system watches every job so a human only touches the ~10% that need judgment.

03

One source of truth, five lenses

Pipeline, QA, Fleet, Analytics and Client health are views over one shared object model, not five apps. Role sets the default lens and write-permissions, never siloes the data.

04

Every number is drillable

No metric is a dead end. Every aggregate links to its constituents: tile to list to record. A KPI dip is always two clicks from its root cause.

05

Utilization has a ceiling on purpose

A capture network is a queueing system. Past ~85% technician or QA utilization, variance explodes and recapture rises. Slack is designed in, not squeezed out.

Rejected directions

Four tempting versions I did not build

The most important judgment in a project is often what you refuse to ship. The dark command-center goes first, because killing it is why the tool you are looking at is white.

Rejected · dark command center
Command Center
Capture12
IM-842002:14
IM-841994:02
IM-841984:02
Processing24
IM-842002:14
IM-841994:02
IM-841984:02
QA18
IM-842002:14
IM-841994:02
IM-841984:02
Delivered214
IM-842002:14
IM-841994:02
IM-841984:02
Rejected direction 01
01

The dark mission-control wallboard

Glowing blue on dark slate, styled like a NOC. It looks impressive in a thumbnail.

Why I killed it. Operators live here for eight hours, they do not glance at a wall. If everything glows, nothing is an alarm.

02

Separate role-specific apps

One simpler product per persona. Each screen is clean in isolation.

Why I killed it. Three apps means three sources of truth, which is the disease the spreadsheet already was.

03

AI auto-dispatch, first

Let the system assign technicians and auto-resolve exceptions from day one.

Why I killed it. Nobody hands enterprise SLAs to a black box, and you cannot train a routing model before instrumenting the manual decisions. Sequenced to v2 as suggestion-with-override.

04

A Gantt as the primary board

One swimlane per job, bars against the SLA clock. SLA is about time, so show time.

Why I killed it. It optimised for planning, not triage, and buried the at-risk 10% in a wall of on-track bars. Right idea, wrong altitude: it became the job-detail timeline.

Information architecture

The system before the screens

For a data-dense tool the object model is the design. I modelled the domain with engineering before drawing a screen.

01Object modelTenant and Contract are the context an Order hangs off, not steps before it. Exception is first-class: any stage can raise one.
has 1:nslaproducescapturesrunsgatesreleases
Tenant
namestring
tierenum
marketsstring[]
Contract
sla_hoursint
rush_tierenum
penaltymoney
Order
addressstring
stateenum · 8
due_attimestamp
time_in_stageduration
CaptureBundle
scansint
overlap_scorefloat
captured_by→ Technician
Technician
market · radiusgeo
utilizationpercent
rfrt_scorepercent
ProcessingJob
gpu_slotref
queued_forduration
compute_stateenum
QAReview
reviewerref
verdictaccept|fix|recapture
defect_codeenum · 5
Delivery
model_urlstring
delivered_attimestamp
vs_sladuration
ExceptionFirst-class · any stage raises it
severityP1 | P2 | P3
defect_codeBLUR · LOW_OVERLAP · COMPUTE_FAIL · MESH_HOLE · NO_SHOW
owner · opened_atwho · when

Scroll sideways to read the full diagram →

02Pipeline state machineDiscovered, not invented. The loops are the point: a QA verdict can send a job back to reprocess, or roll the truck again.
fixrecapturereprocesstruck rolls again
Ordered
Scheduled
Capturing
Uploading
Processing
QA Pending
Fix
Recapture
Approved
Delivered

Scroll sideways to read the full diagram →

PulseOps · Command center for a 3D-capture network

Two flows where the network hands off to a human

The network runs itself and pulls a person in only for the roughly 10% of jobs that need judgment. These are the two moments it chooses to interrupt someone.

Start / endThe operator actsPulseOps actsDecision
Flow 01

Clearing a P1 before the SLA clock runs out

Ops coordinator (Renata), managing by exception across the live pipeline

  1. Wallboard flags a P1 at-risk job
  2. Open the deduped exception queue
  3. Read wait-time on the job timeline
  4. Recoverable inside the SLA window?No: log breach, alert account manager on client detail
  5. Reassign via the Fleet capacity band
  6. Job re-enters pipeline, clock re-baselined
  7. Exception cleared, board back to green

A cleared exception drops the coordinator back to the queue, and the next at-risk job surfaces itself.

Flow 02

QA verdict: accept, fix, or recapture a model

QA reviewer ruling on a flagged 3D model with keyboard decisions

  1. QA Pending model routed to reviewer
  2. Inspect the model in the 3D viewer
  3. Clean enough to accept?Yes: keyboard-accept, Approved then Delivered
  4. Pin the defect, tag its code
  5. Fixable in compute, or recapture?BLUR / LOW_OVERLAP: recapture, job back to Capturing
  6. Route MESH_HOLE fix to compute pipeline
  7. GPU reprocess, re-enter QA Pending
  8. Approved on re-review, delivered

A reprocessed model re-enters QA Pending for a second verdict. A recapture call rolls the truck again.

Design system for density

Color is a scarce resource I spent only on status

Color is spent on status and nothing else. One taxonomy, never hue alone, and every aggregate drills to its parts.

Healthyon-track · accepted · delivered
UploadingI/O in progress
Processingcompute · reconstruction
At-riskWIP high · SLA nearing
BreachSLA breach · rejected
Idlequeued · neutral
SLA at-risk
7
2 critical< 3h to due
18h 40m3h 30m0h 52mdelivered
Command Overview
Live wallboard
Updated 3s ago
⌘K
6
Network healthy·7 at-risk·+6.2% vs target6 open exceptions
Captures
38
12 in field now
In processing
24
WIP 20backing up
In QA
18
3 breachingwait 7.8h
Delivered
214
6.2%vs 200
SLA at-risk
7
2 critical<3h due
Fleet util
79%
in 75–80 band
Live pipelinethroughput today
Capture
12
Upload
9
1 at-risk
Processing
24
2 at-risk
QA
18
3 at-risk
Delivered
214
The Command Overview wallboard: the whole system composed into one glance layer, exceptions ranked on the right rail.
Sketch to ship

From a paper wall to a white workstation

The whole tool started as five dashboards taped to a wall, each carrying one open question in the margin. Every fidelity pass answered a few of them and threw away the parts that looked impressive but read as noise: a wall of paper first, then a stripped greyscale board, then a white workstation where color is spent only on status.

Lo-fi · paper
Locked the surface inventory and the manage-by-exception bet: five landing boards over one shared object model, with the open questions written in the margins before a single pixel was spent.
Mid-fi · greyscale
1Aging heat on left edge2At-risk sorts to top3WIP limit turns header amber
Killed the fourteen-field card and locked the hierarchy in pure greyscale, so nothing leaned on color: aging heat on the border, an at-risk-first sort, and WIP limits that turn a column header amber.
Hi-fi · shipped
Command Overview
Live wallboard
Updated 3s ago
⌘K
6
Network healthy·7 at-risk·+6.2% vs target6 open exceptions
Captures
38
12 in field now
In processing
24
WIP 20backing up
In QA
18
3 breachingwait 7.8h
Delivered
214
6.2%vs 200
SLA at-risk
7
2 critical<3h due
Fleet util
79%
in 75–80 band
Live pipelinethroughput today
Capture
12
Upload
9
1 at-risk
Processing
24
2 at-risk
QA
18
3 at-risk
Delivered
214
The shipped Command Overview: the whole system composed into one glance layer, exceptions ranked on the right rail.
Lo-fi · paper

Five boards, taped to a wall

Locked the surface inventory and the manage-by-exception bet: five landing boards over one shared object model, with the open questions written in the margins before a single pixel was spent.

Mid-fi · greyscale

One card, stripped to four fields

Killed the fourteen-field card and locked the hierarchy in pure greyscale, so nothing leaned on color: aging heat on the border, an at-risk-first sort, and WIP limits that turn a column header amber.

Hi-fi · shipped

Color spent only on status

Locked the white operator tool and the canonical status taxonomy of color plus icon plus label, with calm real-time on one shared clock tick. This is the version handed to engineering.

Marquee flow

Deep dive: the live pipeline board

It is 2pm. Twenty-four jobs are backing up in Processing and two are two hours from an enterprise breach. Three hard problems had to be solved on one surface.

  • Spotting a stall among hundreds of cards. Aging heat on the left border, an at-risk-first sort, and a WIP limit that turns a column header amber when it backs up. The eye is drawn to the problem, not asked to scan for it.
  • How much on a card before it is noise. Progressive disclosure: the card shows ID, client, SLA countdown and a time-in-stage meter. Everything else waits for the detail workspace. This is the fix from the fourteen-field board that tested badly.
  • Real-time that never makes the board jump. One shared clock tick drives every countdown, and the board never re-sorts under an active cursor. Exactly one live pulse; no ambient jitter.
Job IM-84213
Pipeline · Processing
Updated 3s ago
⌘K
6
PipelineProcessingIM-84213
IM-84213Redwood Realty · 1428 Camino Real, San JoseProcessing24h rush · 2h 14m to due
ReassignExpediteRe-queueOpen in QA
Lifecycle timeline wait time = 66% of elapsed
Ordered08:12·system
◷ 28m to schedule
Scheduled08:40·@dana
assigned Marcus R.
Captured10:20·Marcus R.
92 scan positions · 31% overlap
Uploaded10:42·system
6.4 GB in 22 min
Processing11:04·GPU-07
41 min · 1 COMPUTE_FAIL retry
QA Pending11:45·queue
◷ waiting 1h 12m
Capture summary
Rooms8
Area4,200 sqft
Floors2
OccupancyOccupied
DeviceiPhone 15 Pro
Scan positions92
Coverage96%
Align residual1.2 cm
Assignment
MR
Marcus R.
San Jose zone · load 4/5
Reassign (capacity-aware)
Exceptions & QA history
COMPUTE_FAILretry 1 · resolved
MESH_HOLEflagged in QA · pending fix
Activity log
11:45queued for QA
11:44@marcus expedited
11:04processing started · GPU-07
10:42upload complete
The lifecycle timeline: most elapsed time is wait time between stages.

The hours pool in Processing, so that stage earns its own surface for the reconstruction engineer: GPU queue depth, per-node health, failure and retry rate, and the quality signals that catch a silent regression before QA.

Compute Pipeline
GPU fleet · reconstruction
Updated 3s ago
⌘K
6
Queue depth
18
6 running
Throughput
42
jobs / hr
Failure rate
3.1%
last 24h
Retry rate
6%
auto-retry
Avg runtime
5.1h
bottleneck stage
GPU fleet · 8 nodesrunningidlefailed
GPU-01running
IM-84190
88% util
GPU-02running
IM-84176
76% util
GPU-03running
IM-84120
81% util
GPU-04retry 1
IM-84213
0% util
GPU-05running
IM-84077
84% util
GPU-06idle
no job
—
GPU-07running
IM-83998
96% util
GPU-08running
IM-84051
79% util
Reconstruction queue18 waiting
1IM-842406.4 GBwaiting 2m
2IM-842383.1 GBwaiting 4m
3IM-842365.8 GBwaiting 7m
4IM-842312.4 GBwaiting 9m
Reconstruction quality signalssurfaced upstream of QA
Alignment residual1.2 cm
within 2.0 cm tolerance
Mesh-hole rate4%
trending up 1.2 pts
Measurement confidence96%
stable
Watching the bottleneck stage. A failed GPU node surfaces inline, with the backlog behind it.
The differentiator

Deep dive: the 3D-model QA review workflow

The reviewer inspects a spatial 3D artifact inside a 2D operational tool and decides accept, fix or recapture, at speed and with consistency. This is the surface that makes PulseOps domain-deep, unmistakably distinct from the order-lifecycle study.

QA Review
18 models in queue
Updated 3s ago
⌘K
6
IM-84213Redwood Realty · 1428 Camino Real, San Jose24h rush · 0h 52m to due
RawReconstructedFloorplanMeasure
master bath 8.4 ft · spec 8.0 (+0.4)
LivingKitchenBed 1Master BathBed 2Shower24.0 ft8.4 ft (+0.4)25
Three panes: the SLA-sorted queue, the model viewer with defect pins, and the verdict.

Three moves do the work. A fixed defect taxonomy makes reject reasons analyzable: every code feeds the exception chart. A keyboard-first flow protects review velocity. And the fix-versus-recapture call is made explicit, because it is the money decision: a light fix keeps the 24h SLA, while a recapture is roughly twice the cost and usually blows it. Then the loop closes: reject codes roll up into a per-technician capture-right-first-time score that feeds dispatch. The system learns from itself.

The business layer

Turning “turnaround feels slow” into a bottleneck you can point at

The analytics surface answers one question: where is the network breaking today, and who is affected. Every metric drills through to the underlying orders. The payoff is the bottleneck-attribution bar.

SLA & Client Health
Last 30 days
Updated 3s ago
⌘K
6
Last 30 daysAll regionsAll clientsAll typesnumbers marked target are design goals, not shipped results
SLA met
91%target 96%
Cycle time
31htarget 24h
Throughput
214per day
QA accept
82%target 85%
Exceptions
9%of jobs
Cycle-time distribution23% past SLA
300
0h24h SLA72h+
Arrivals vs completionsarrivalscompletions
140200260
May 1per dayMay 30
Exceptions by reasonshare of 100%
BLUR
24%
LOW_OVERLAP
19%
COMPUTE_FAIL
16%
MESH_HOLE
14%
NO_SHOW
12%
MEASUREMENT_OOT
9%
ACCESS_FAIL
6%
Median time in stage · where the hours gowait time = ~66% of total cycle
5.1h
7.8h
Capture0.6hUpload0.4hProcessing5.1hQA7.8hDelivery0.2h
Enterprise client health91% global average hides the tail below
ClientVol 30dSLA attainmentCycleOpen excTrendHealth
Redwood Realty1,240
94%
26h288 · Healthy
Meridian Homes870
92%
28h381 · Healthy
Sentinel Insurance640
86%
34h763 · At-risk
Cascade Property Group410
78%
41h1147 · Critical
The KPI layer, each tile drilling through to the orders behind it.
The payoff
Median time in stage
Where the hours actually go
91% pools in Processing + QA
5.1h
7.8h
Capture0.6hUpload0.4hProcessing5.1hQA7.8hDelivery0.2h14.1h median end to end
Wait time is about 66% of total cycle. This bar is what reframed the roadmap from make the pipeline faster to attack the queue.

The dangerous 78% is one click from the average. Drill into Cascade and the story is specific and defensible: a 16-point SLA slide over seven weeks, an exec escalation, renewal at risk, and a recommendation pointing back at the Processing bottleneck. This is the surface an account manager walks into a QBR with.

Cascade Property Group
Enterprise · 24h rush tier
Updated 3s ago
⌘K
11
ClientsCascade Property Group
Health 47 · Criticalrenewal at risk
Account teamPMSO
SLA attainment
78%
contractual 95%
Volume 30d
410
of 500 commit
Median cycle
41h
target 24h
Open exceptions
11
3 breaching
SLA attainment · 7 weeks down 16 pts
7 wks ago95% contractual floornow · 78%
At-risk jobs4 of 11 open
IM-839982100 BroadwayProcessing2h 14m
IM-8404441 K StreetCapture3h 10m
IM-840519 Mission StQA1h 36m
IM-8412088 Pine AveUpload5h 40m
Escalations & account timeline
Exec escalation2d ago

12 orders missed the 24h SLA last week.

Renewal flagged at-risk5d ago

CSM logged churn risk ahead of Q3 renewal.

QBR scheduledin 6d

Root-cause review with the account.

Recommended

Attack the Processing bottleneck for Cascade first: it is 5.1h of the 41h cycle. Expedite the 3 breaching jobs today.

Per-tenant health: the account a global average was hiding.
The other half of the equation

Capacity is physical, so the tool refuses to over-drive it

You cannot just process faster: capture is a truck-roll inside a coverage radius. The fleet surface makes that constraint visible and pairs utilization with capture-right-first-time, so the tool never rewards pushing a technician into quality failures.

Fleet & Capacity
6 technicians · 3 zones
Updated 3s ago
⌘K
6
Network utilization
79%
target band 75–80
Captures / tech-day
4.6
target 5.5
On-time arrival
93%
last 7 days
No-show rate
4%
last 7 days
Right-first-time
86%
capture quality
Technician utilization 75–80% healthy band
Marcus R.San Jose
88%5/d
Priya S.Oakland
81%5/d
Dana P.San Jose
72%4/d
Leo M.Oakland
69%4/d
Sam T.Sacramento
61%3/d
Nina K.Sacramento
58%3/d
Marcus R. over-driven at 88%. Rebalance to protect capture quality.
Coverage · ~35 mi service radiiSacramento gap
OaklandSan JoseSacramento
Sacramento demand tomorrow22h needed · 16 avail
TechnicianZoneCaptures/dayOn-timeNo-showRight-first-time
Marcus R.San Jose5.096%2%84%
Priya S.Oakland4.894%3%88%
Dana P.San Jose4.492%5%89%
Sam T.Sacramento3.290%6%82%
Utilization carries a 75 to 80% healthy band. Past it, variance explodes.
Failures & recovery

Three things I got wrong

A case study with zero failures reads as a decorator who was never trusted with a hard problem. Each of these broke when it met reality, and each one forged a principle stated earlier on this page.

Iteration 01 · killed after testing
Ordered8
IM-84200
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84199
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84198
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Assigned6
IM-84193
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84192
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84191
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Accepted5
IM-84186
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84185
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84184
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Capturing12
IM-84179
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84178
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84177
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Uploaded9
IM-84172
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84171
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84170
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Processing24
IM-84165
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84164
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84163
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84162
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
QA18
IM-84158
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84157
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84156
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Rework4
IM-84151
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84150
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84149
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Delivered214
IM-84144
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84143
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84142
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
◷48s to locate the one breaching job
“I can't tell which one is actually on fire.”
Ops lead, first usability session
✗?
Nine columns, 14 fields a card, everything the same weight. It became wallpaper.

The eleven-second board

I built
IM-84213Redwood RealtySan JoseScheduledMara O.2 scans1.4 GBv3GPU-2priority 2created 11/04notetier entrev $840
11 seconds to find the risk
The fix
IM-84213
Cascade Property
0h 52m
IM-84109
Redwood Realty
3h 30m
at-risk first · four fields
It broke

I asked the ops lead to find the most at-risk job. Eleven seconds and a scroll. Everything weighed the same, so nothing was salient.

I learned

Density is not the enemy. Undifferentiated density is.

The wall of red

I built
SLA warning
upload failed
delivered
compute retry
SLA warning
no-show
63 alerts · all muted by day two
The fix
P1SLA breach risk · 2 jobs
P2COMPUTE_FAIL · 3 jobs
P3BLUR · 1 job
6 real fires · deduped by cause
It broke

An SLA breach looked identical to a routine delivery, so within a day people muted every alert. I had built alert fatigue on purpose.

I learned

Unranked alerting is worse than none. It teaches people to ignore the signal.

The defeated rubric

I built
QA rubric · 12 criteria
Coverage
default
Overlap
default
Lighting
default
Mesh integrity
default
2× review time · defaults slammed
The fix
AacceptRrejectJ / Kmove
Reject requires one code
BLURLOW_OVERLAPMESH_HOLE
one keystroke · one code · real signal
It broke

Twelve gated criteria doubled review time on a 24h SLA, so reviewers slammed defaults to reach the decision. Slower AND worse data.

I learned

The signal has to ride along with the decision, not gate it.

Failure 02 · the fix
First alert modelMute all
COMPUTE_FAILIM-84213just now
COMPUTE_FAILIM-84210just now
COMPUTE_FAILIM-84207just now
NO_SHOWIM-84204just now
COMPUTE_FAILIM-84201just now
SLA_BREACHIM-84198just now
COMPUTE_FAILIM-84195just now
BLURIM-84192just now
COMPUTE_FAILIM-84189just now
Every state change fires a toast, so the rational response is to mute all of them, and the real breach hides in the noise.
Alert fatigue I built by accident, then designed out
The fix · quiet by default6 real fires
P1pages on-call
P2sits in the rail
P3rolls into a daily digest
dedup by root cause
20 × COMPUTE_FAIL→1 GPU-node incident
An unranked board surfaces 63 alerts, nearly all red. Ranked and deduped, it lands at six.
Exception Triage
6 open · quiet by default
Updated 3s ago
⌘K
6
P1
2
pages on-call
P2
3
in the rail
P3
1
daily digest
◷
38m
median resolve · target < 45m
severity: allcode: allowner: all 1 unowneddeduped by root cause · 20 COMPUTE_FAIL → 1
P1 · Pages on-call2
COMPUTE_FAIL
IM-84213Redwood Realty
2h 14m
M@marcus6m
Re-queue
SLA_BREACH_RISK
IM-83998Cascade Property
2h 14m
unassigned2m
Expedite
P2 · In the rail3
NO_SHOW
SAC 10:00Sentinel Insurance
—
D@dana14m
Reschedule
BLUR
IM-84109Redwood Realty
3h 30m
P@priya22m
Recapture
LOW_OVERLAP
IM-84077Meridian Homes
4h 02m
D@dana31m
Flag QA
P3 · Daily digest1
MESH_HOLE
IM-84051Cascade Property
—
P@priya38m
Route to fix
Ranked P1/P2/P3, deduped by root cause, quiet by default, a named owner on every row.
Prototyping & validation

What broke when a real coordinator drove it

I cannot dogfood an ops tool, so I sat the ops lead and a working coordinator in front of it on a seeded shift. Every hesitation was a to-do. Three of them moved the design.

Track 01 · Coordinator think-aloud, on a seeded shift
Done
01

A number that could mean two things

Friction
IM-84213Cascade Property4h 20m
elapsed, or left? she ranked it wrong

She read the SLA chip as time elapsed, not time left, and ranked the queue wrong.

Change
time to due
IM-84213Cascade Property0h 52m
IM-84109Redwood Realty3h 30m

Every chip became a countdown to breach, labelled to due, on a colour ramp.

02

A fire with nobody's name on it

Friction
COMPUTE_FAIL · IM-84213unassigned
BLUR · IM-84109unassigned
she reopened one already handed off

Working top-down, she reopened a job already handed off. Nothing said it was taken.

Change
COMPUTE_FAIL · IM-84213claimed @dana
BLUR · IM-84109resolved, out of queue
a triaged fire visibly leaves

Every exception gets a named owner and a claimed state, so a triaged fire leaves the queue.

03

A board with no clock

Friction
SLA at-risk
7
no way to tell if it is current

She distrusted it the way she distrusted the old sheet: is this number current?

Change
SLA at-risk
7
Updated 3s ago · one shared tick

A freshness stamp on one shared tick answers it before she has to ask.

The pivot

Watching the ops lead take eleven seconds to find the one breaching job is what turned this from a denser dashboard into a queue that ranks its own fires.

Track 02 · Moderated study, five coordinators
Scoped

When a job is quietly aging in a queue nobody is watching, does PulseOps put it in front of a human before the SLA is gone?

Five working coordinators who protect a live delivery SLA, none who have seen PulseOps. Sixty-minute moderated sessions on a seeded dataset.

0102030405
0 min5 tasks, think-aloud60 min
Task script
  1. 01Take over the board mid-shift. What is about to be late?
  2. 02Find the Cascade order that has sat unwatched since 11am.
  3. 03Handle a BLUR failure on a job due in three hours.
  4. 04Show me where the network is actually stuck this week.
  5. 05Answer whether Cascade's SLA is healthy.
Success bar
Pre-registered
Time-to-triage, login to first action< 2 min
Breach-spotting task success≥ 4 / 5
Breaches missed in session0
Ease per task (SEQ)≥ 5 / 7
Overall usability (SUS)≥ 80
Collaboration & tradeoffs

Who I decided with, and what each call cost

None of these calls were mine alone, and none of them were free.

Ops team

Co-defined the status taxonomy so UI states matched real database states.

Platform eng

Negotiated the refresh strategy against its performance cost.

Product manager

Scoped the MVP around the highest-leverage surface.

Polling over websockets

streamingno staleness shown
Streaming, v1
Updated 3s agostaleness on screen
Polled, freshness shown
Gave up

Up-to-5s staleness, and a streaming path deferred to v2.

Bought

Shipping inside the eng budget, with the staleness stamped on screen.

One dense board, not three apps

DIS
QA
EXEC
Three role apps
dispatchqa
One board, role defaults
Gave up

A steeper onboarding curve for every role.

Bought

One source of truth. Three apps is the reconciliation tax again.

The SLA path first

board
exceptions
fleet
qa
analytics
All five at once
board
exceptions
fleet
qa
analytics
SLA path first
Gave up

Fleet, QA and analytics, all sequenced after v1.

Bought

The surface that protects the SLA earned the first sprint.

Projected / target outcomes

The targets I designed toward, and how I validate them

Design targets, not shipped results. Each pairs a mechanism I can point to in the interface with the outcome it moves.

On-time delivery within SLA
target 98% · projected
= the sum of four stage wait-times
1
Capture → Upload
surfaced by
Pipeline board
2
Upload → Processing
surfaced by
Job timeline
3
Processing → QA
surfaced by
Analytics bottleneck
4
QA → Delivery
surfaced by
QA queue
The legacy tool measured zero of these four leaves. You cannot manage a number you cannot see.
Design mechanism (leading)Operational outcome (lagging)
Time-to-triage15–20 min → under 2 min
→
SLA adherence88% → 96%+target

One ranked queue replaces reconstructing state from three tools each shift.

Per-stage aging visibility0 → every stage
→
End-to-end cycle time (P90)72h → 40htarget

Wait time is where the hours are. Aging plus WIP limits make dwell a routable signal.

Manage-by-exception queuewatch all → touch ~10%
→
Ops headcount leverage+30–40% jobs / coordinatortarget

The system watches every job and surfaces only the exceptions. This is the pattern tied to the business model.

Fixed QA defect taxonomyfree-text → 5 codes
→
First-pass QA yield70% → 85%target

A fixed taxonomy cuts reviewer variance and pushes capture problems upstream.

Capture-quality feedback loopcaught at QA → at capture
→
Rework / recapture rate12% → under 7%target

Recapture is the most expensive failure, so detection shifts left to the capture moment.

Utilization band, cappeduncapped → 75–80%
→
Cost per delivered modelbounded, not maximizedtarget

Bounded on purpose: past 85% variance explodes. Guarding the ceiling is the defensible move.

Reflection

What I validate first, and where PulseOps goes next

Does per-stage aging actually reduce wait time, or does it only make waiting visible? That is the mechanism the whole thesis rests on, so it is the one to falsify first.

The v2 I scoped: streaming to replace polling, predictive SLA-risk scoring calibrated hard against alert fatigue, suggestion-with-override auto-routing trained on the QA feedback loop, and an explicit, visible multi-tenant priority policy so ops can justify why one enterprise job jumps another. The honest hard part was resisting the urge to show everything. The first three board layouts were too dense to act on, and the version I handed to engineering is the one that hid the most without hiding anything that mattered.

Let's build something people remember

From enterprise teams to growing startups.

Let's talkarifin.yeasin@gmail.com