← Himanshu Kalra

UX Research Case Study

The redesign experts rejected

A security platform shipped a modern redesign, and expert analysts kept preferring the build it replaced. Preferring the old version was not resistance to change. It was a precise signal about what the redesign had removed.

OutcomeDiagnosed the adoption regression across a multi-round program - 45 participants across 11 organizations, plus a 67-analyst operations study - prioritized three Tier-1 blockers, and drove fixes into the roadmap. Two of four shipped within a release.

Enterprise security operations platform 45 participants, 11 orgs (+67-analyst quant) Iterative usability + competitive benchmark + telemetry Q2 2025 to Q1 2026 Lead researcher

The platform had shipped a full redesign of its security operations console. It looked more modern than the build it replaced. Telemetry and preference both pointed the other way: expert analysts kept reaching for the legacy version, and adoption of the new one stalled.

The easy read was change resistance. The research existed to test the harder one. The question was not whether the redesign looked better. It was why analysts who lived in this tool four to eight hours a day could no longer do their core work in it, and what specifically the new build had taken away.

The program ran across three quarters, not one study. Iterative moderated usability rounds on prototypes (21 and 24 participants) tested fixes as they were drafted, benchmarked against the competitor tools analysts came from. Those rounds fed a two-quarter synthesis of 45 participants across 11 organizations, which became a blocker-prioritization readout to product and engineering.

A separate 67-analyst operations study gave the qualitative findings a quantitative floor: how overloaded analysts actually are, how often they pivot tools, and which actions they run every day. Recruitment was deliberate - expert users who spend four to eight hours a day in the tool, because a redesign regression only shows up under real operating load, not in a first-run walkthrough.

Prioritization was not by loudest voice. Workflow regressions were weighted by real usage from product telemetry, so a fix earned its rank by how many analysts hit it daily, not by how vividly it was described in a session.

The redesign removed the query primitive

The most requested fix, named by 11 of 14 analysts in design validation, was query autocomplete and schema discovery. In the redesign an analyst had to know the exact field name to search, and the schema is unforgiving: is it source_ip, src_ip, or IPv4. Without a hint, analysts fell back on senior colleagues or external docs to write a basic query.

That single gap set the ramp time. New analysts reached productivity in six to nine months on this platform, against two to three on the competitors, which all shipped autocomplete. The cost was not only internal: the gap showed up in lost deals, and was framed internally as more than half a million dollars of annual revenue at risk.

It also stripped context and density

Two more regressions came straight from the redesign. The legacy build had shown whether an alerted user was a VIP, with a crown icon and department visible on the alert. The redesign removed both, so analysts could no longer tell a CEO from a contractor at a glance. Six of fourteen named this, and the risk is not cosmetic: it drives wrong prioritization and high-impact mistakes on the wrong account.

The second was density. Host, IP, domain, and file-path details that analysts need on nearly every alert were buried a panel away, costing three to five extra clicks per alert. At peak operations that runs up to 300 alerts an hour, so the clicks were not a nuisance, they were the workload.

Just let me respond from the alert.a SOC analyst

Prefer the old version was the signal

Each blocker on its own was a feature gap. Together they explained the preference the team had read as resistance. The redesign had optimized for looking modern and, in the process, removed the primitives expert analysts had built their speed on: the query hint, the at-a-glance context, the dense single view. For a novice that trade is invisible. For an expert running at 300 alerts an hour it is a tax paid on every alert.

A build that removes what experts move fast on is a regression, whatever it looks like.

The operations data made the stakes concrete. Analysts were overwhelmed daily (79% in the 67-analyst study), pivoted across three to five tools per investigation, and ran bulk actions constantly. That is a population with no spare capacity to relearn primitives the previous build had already given them. Preferring the old version was the most rational thing they could do, and it was pointing directly at the fix list.

Two of four blockers shipped, and the rest were named as risk

Shipped within a release

  • Delivered - detection details tab, addressing two of the blockers
  • Delivered - enhanced alert timeline
  • Delivered - response automation for repeated steps

Named as open risk

  • Open - VIP / role context
  • Open - query autocomplete, phase 1
  • Open - enhanced alert-summary view

Naming the unshipped blockers as explicit, evidenced risk mattered as much as the wins. It moved the two hardest fixes from "nice to have" into roadmap and sales conversations, where the deal-loss evidence could carry them.

Renamed for how analysts read

"Assets" became entity-specific labels (users, hosts, network context) after 8 of 14 read "asset" as a device, not a person or address.

Phased the query fix

Autocomplete sequenced as field names first, then operators and values, then a visual builder, so the highest-pain slice shipped soonest.

Alert-type-aware summary

Surface five to seven critical fields per alert type, lazy-loaded to respect a real render-time constraint.

Fed sales enablement

Autocomplete and VIP context added to competitive battle cards and demo scripts, on the deal-loss evidence.

Two things. First, keep the two open blockers instrumented, because "named as risk" is not "fixed," and the ramp-time and deal-loss costs keep accruing until they ship. Second, the design-validation fractions sit on a small panel (14 users); the weight belongs on the larger synthesis and the 67-analyst study, and any public claim should state its n rather than blend them. One gap the program surfaced but did not resolve: no formal post-incident-review process was observed across these teams, which warrants a study of its own.

Measure a redesign against the workflow it replaces, not against a blank slate. When a new build strips the primitives expert users built their speed on, preferring the old version is data, not resistance, and it points straight at the fix list.