PulseOps is the internal command center for running InsideMaps' nationwide 3D-capture network: a live capture-to-delivery pipeline, a 3D-model QA workflow, fleet utilization, and multi-tenant SLA analytics. Designed to double delivery volume without doubling the humans watching it.
The live pipeline board: every job flowing capture to delivered, sorted most-at-risk first.
2:10pm. Renata, a coordinator, is toggling between three browser tabs. A Cascade Property Group order is ninety minutes from its SLA and she does not know it. The job finished processing at 11am and has sat in an unwatched QA queue ever since.
She finds out the way she always does: the account manager pastes a red-faced client message into Slack. The truck already went home. A recapture now means the deliverable is a day late at twice the cost. Nobody did anything wrong. The tools simply could not show what was stuck.
The failure was not a person. It was that no screen in the company could answer one question: what is about to be late, right now?
A logistics, GPU-compute and human-QA machine wrapped in a 24 to 48 hour promise. The product is the promise, not the model file.
A legacy admin panel, a shared spreadsheet, and a wall of Slack pings. Status was a color word with no clock behind it.
The private spreadsheet was the tell. The real system of record lived outside the software.
A job finished upload and waited on a free GPU slot; finished processing and waited on a free QA reviewer. Most elapsed time was wait time, not touch time, and no tool measured where jobs sat.
Failed uploads, compute crashes, no-shows, blurry captures: all discovered reactively, often by a client complaint. Easy jobs got cleared while at-risk jobs quietly aged past their deadline.
Every job needed manual watching, status chasing, and hand assignment. That is the exact constraint that stops a B2B ops org from growing revenue without growing headcount in lockstep.
The data always existed. The legacy tool just could not show it in time to act.
No tool at InsideMaps recorded when a job entered or left a stage, so "slow turnaround" was a feeling nobody could locate. I spent discovery on the operator workstations themselves: shadowing full shifts, interviewing every role that touches a job, and reconstructing real stage timings from the raw order log. One pattern held across all of it, and it reframed the project from "clean up the dashboard" to "make the network watch itself."
Sat beside coordinators through full 8-hour shifts in three markets, tallying every tab-toggle and Slack check that rebuilt the network's state.
One-on-one across all five roles: coordinators, QA reviewers, a reconstruction engineer, the ops lead, and two account managers who own the QBR.
Reconstructed real per-stage dwell times from the raw order log, since no tool had ever recorded when a job actually moved between stages.
Mapped the true pipeline states and defect codes with ops and eng, splitting one word, "Processing," into capture, upload, compute, and QA.
Jobs spent hours queued between stages, not being worked, and "Processing" hid four different waits, so the target was dwell time, not touch time.
With no time dimension in status, coordinators spent 15-20 minutes each shift rebuilding what was urgent before they could act on anything.
Exceptions were discovered reactively through a client's Slack message, always after the SLA had already been lost and the truck had gone home.
A healthy global SLA number masked individual enterprise tenants quietly breaching, and per-market silos hid the breach from the account managers who owned it.
Design-target archetypes, drawn from the interviews, ride-alongs and the problem space.
The problem was never a cluttered dashboard. It was invisible wait time and reactive triage. So the goal was not to show more; it was to make the network watch itself, and interrupt a human only for the ~10% of jobs that need judgment.
Manage-by-exception is not a UI pattern here. It is the business model: the only thing that breaks the linear-headcount curse. And because a capture network is a queueing system, utilization has a ceiling on purpose: the tool has to refuse to reward over-driving.
Four facts ruled out most of the obvious designs, and generated the five principles below.
Operators live in this tool for eight-hour shifts. Information density is a feature. The craft is managing it with a strict altitude system, not hiding it behind whitespace.
Healthy work collapses into count chips. Only the anomalous earns its own row. The system watches every job so a human only touches the ~10% that need judgment.
Pipeline, QA, Fleet, Analytics and Client health are views over one shared object model, not five apps. Role sets the default lens and write-permissions, never siloes the data.
No metric is a dead end. Every aggregate links to its constituents: tile to list to record. A KPI dip is always two clicks from its root cause.
A capture network is a queueing system. Past ~85% technician or QA utilization, variance explodes and recapture rises. Slack is designed in, not squeezed out.
The most important judgment in a project is often what you refuse to ship. The dark command-center goes first, because killing it is why the tool you are looking at is white.
Glowing blue on dark slate, styled like a NOC. It looks impressive in a thumbnail.
Why I killed it. Operators live here for eight hours, they do not glance at a wall. If everything glows, nothing is an alarm.
One simpler product per persona. Each screen is clean in isolation.
Why I killed it. Three apps means three sources of truth, which is the disease the spreadsheet already was.
Let the system assign technicians and auto-resolve exceptions from day one.
Why I killed it. Nobody hands enterprise SLAs to a black box, and you cannot train a routing model before instrumenting the manual decisions. Sequenced to v2 as suggestion-with-override.
One swimlane per job, bars against the SLA clock. SLA is about time, so show time.
Why I killed it. It optimised for planning, not triage, and buried the at-risk 10% in a wall of on-track bars. Right idea, wrong altitude: it became the job-detail timeline.
For a data-dense tool the object model is the design. I modelled the domain with engineering before drawing a screen.
Scroll sideways to read the full diagram →
Scroll sideways to read the full diagram →
The network runs itself and pulls a person in only for the roughly 10% of jobs that need judgment. These are the two moments it chooses to interrupt someone.
Ops coordinator (Renata), managing by exception across the live pipeline
A cleared exception drops the coordinator back to the queue, and the next at-risk job surfaces itself.
QA reviewer ruling on a flagged 3D model with keyboard decisions
A reprocessed model re-enters QA Pending for a second verdict. A recapture call rolls the truck again.
Color is spent on status and nothing else. One taxonomy, never hue alone, and every aggregate drills to its parts.
The whole tool started as five dashboards taped to a wall, each carrying one open question in the margin. Every fidelity pass answered a few of them and threw away the parts that looked impressive but read as noise: a wall of paper first, then a stripped greyscale board, then a white workstation where color is spent only on status.
Locked the surface inventory and the manage-by-exception bet: five landing boards over one shared object model, with the open questions written in the margins before a single pixel was spent.
Killed the fourteen-field card and locked the hierarchy in pure greyscale, so nothing leaned on color: aging heat on the border, an at-risk-first sort, and WIP limits that turn a column header amber.
Locked the white operator tool and the canonical status taxonomy of color plus icon plus label, with calm real-time on one shared clock tick. This is the version handed to engineering.
It is 2pm. Twenty-four jobs are backing up in Processing and two are two hours from an enterprise breach. Three hard problems had to be solved on one surface.
The hours pool in Processing, so that stage earns its own surface for the reconstruction engineer: GPU queue depth, per-node health, failure and retry rate, and the quality signals that catch a silent regression before QA.
The reviewer inspects a spatial 3D artifact inside a 2D operational tool and decides accept, fix or recapture, at speed and with consistency. This is the surface that makes PulseOps domain-deep, unmistakably distinct from the order-lifecycle study.
Three moves do the work. A fixed defect taxonomy makes reject reasons analyzable: every code feeds the exception chart. A keyboard-first flow protects review velocity. And the fix-versus-recapture call is made explicit, because it is the money decision: a light fix keeps the 24h SLA, while a recapture is roughly twice the cost and usually blows it. Then the loop closes: reject codes roll up into a per-technician capture-right-first-time score that feeds dispatch. The system learns from itself.
The analytics surface answers one question: where is the network breaking today, and who is affected. Every metric drills through to the underlying orders. The payoff is the bottleneck-attribution bar.
The dangerous 78% is one click from the average. Drill into Cascade and the story is specific and defensible: a 16-point SLA slide over seven weeks, an exec escalation, renewal at risk, and a recommendation pointing back at the Processing bottleneck. This is the surface an account manager walks into a QBR with.
You cannot just process faster: capture is a truck-roll inside a coverage radius. The fleet surface makes that constraint visible and pairs utilization with capture-right-first-time, so the tool never rewards pushing a technician into quality failures.
A case study with zero failures reads as a decorator who was never trusted with a hard problem. Each of these broke when it met reality, and each one forged a principle stated earlier on this page.
I asked the ops lead to find the most at-risk job. Eleven seconds and a scroll. Everything weighed the same, so nothing was salient.
Density is not the enemy. Undifferentiated density is.
An SLA breach looked identical to a routine delivery, so within a day people muted every alert. I had built alert fatigue on purpose.
Unranked alerting is worse than none. It teaches people to ignore the signal.
Twelve gated criteria doubled review time on a 24h SLA, so reviewers slammed defaults to reach the decision. Slower AND worse data.
The signal has to ride along with the decision, not gate it.
I cannot dogfood an ops tool, so I sat the ops lead and a working coordinator in front of it on a seeded shift. Every hesitation was a to-do. Three of them moved the design.
Watching the ops lead take eleven seconds to find the one breaching job is what turned this from a denser dashboard into a queue that ranks its own fires.
None of these calls were mine alone, and none of them were free.
Co-defined the status taxonomy so UI states matched real database states.
Negotiated the refresh strategy against its performance cost.
Scoped the MVP around the highest-leverage surface.
Up-to-5s staleness, and a streaming path deferred to v2.
Shipping inside the eng budget, with the staleness stamped on screen.
A steeper onboarding curve for every role.
One source of truth. Three apps is the reconciliation tax again.
Fleet, QA and analytics, all sequenced after v1.
The surface that protects the SLA earned the first sprint.
Design targets, not shipped results. Each pairs a mechanism I can point to in the interface with the outcome it moves.
One ranked queue replaces reconstructing state from three tools each shift.
Wait time is where the hours are. Aging plus WIP limits make dwell a routable signal.
The system watches every job and surfaces only the exceptions. This is the pattern tied to the business model.
A fixed taxonomy cuts reviewer variance and pushes capture problems upstream.
Recapture is the most expensive failure, so detection shifts left to the capture moment.
Bounded on purpose: past 85% variance explodes. Guarding the ceiling is the defensible move.
Does per-stage aging actually reduce wait time, or does it only make waiting visible? That is the mechanism the whole thesis rests on, so it is the one to falsify first.
The v2 I scoped: streaming to replace polling, predictive SLA-risk scoring calibrated hard against alert fatigue, suggestion-with-override auto-routing trained on the QA feedback loop, and an explicit, visible multi-tenant priority policy so ops can justify why one enterprise job jumps another. The honest hard part was resisting the urge to show everything. The first three board layouts were too dense to act on, and the version I handed to engineering is the one that hid the most without hiding anything that mattered.
From enterprise teams to growing startups.