← Himanshu Kalra

UX Research Case Study

Reversibility is not reliability

A security team will trust an autonomous AI to reconstruct an attack. It will not trust the same AI to act on one. That gap is built into the architecture, not the sentiment.

OutcomeReframed the product's autonomy architecture before launch, and reset three shipping decisions about what an AI may act on, how it is corrected, and how its confidence is shown.

Enterprise AI SOC platform ~35 security professionals 4 regions 6 weeks Lead researcher, APAC + Europe cohort Prototype-driven usability + discovery

The product bet on one thing: that a security operations center would hand real investigation and response work to an autonomous AI. The research question was not whether the interface worked. It was what the AI had to do to earn that handoff, and where the handoff stopped.

When the platform build slipped, everything ran on interactive prototypes and a hosted demo rather than live software. That constraint set the pace. The prototype could change between cohorts, so it did, every week, which turned six weeks of fieldwork into six rounds of iteration.

The study ran as a design-partner program: small cohorts, five or so at a time, run and re-run rather than tested once. Each week's prototype was built to answer the previous week's open question, so six weeks of fieldwork became six rounds of iteration instead of one verdict at the end. A pre-launch product whose questions changed weekly needed a method that could change with them.

The audience widened on purpose. The target user was the SOC analyst, but the panel deliberately reached past them to security admins, CISOs, and SOC managers, because the trust question only resolves once the people who sign off on autonomy are in the room. That widening is also why the usability scores moved: a broader audience judged a thickening prototype.

Scale was handled with tooling. A custom rainbow-spreadsheet coding instrument, snapshotted weekly, kept synthesis current across roughly 35 sessions in four regions. Sessions were AI-transcribed into highlight reels so stakeholders acted on evidence within days, not weeks. The round closed on a saturation call, once scores stabilized and new sessions stopped changing the picture.

Trust clears a higher bar than expected

Analysts wanted the AI for the task most people assume they would guard: reconstructing an incident across SIEM, EDR, and NDR. By hand, that means opening tool after tool to assemble a single picture. The AI assembles it in seconds, and analysts trusted it to.

A human can make more mistakes in a complex timeline with multiple sources than an AI.one of the most trust-cautious analysts in the study

Every endorsement carried the same condition. Each entry links to its source, with a path back to the raw data. Remove the link and the trust goes with it.

Legibility, not obedience

When analysts edited the timeline, they reached for direct manipulation. When they opened chat, they asked one kind of question: why did you flag this, why did you rate it low confidence. Chat was for interrogation, and the interrogation grew each week. The trust driver named most often was auditability.

I love that it is showing the reasoning. It matters to understand how it reasoned for this.a SOC analyst

One story set the stakes. An AI-generated root-cause report went badly wrong, reached executives who read it as fact because it came from AI, and had to be walked back as a hallucination. A review gate and legible reasoning are what separate a mistake from a reputational one.

The ceiling is architectural, not emotional

The team's working assumption was that reversibility buys autonomy: make an action undoable and the system can take it alone. For the buyers the product needs most - regulated industries, healthcare, government, finance, infrastructure - that assumption breaks.

Undo is a recovery mechanism. It is not an autonomy license.

Analysts found reject-per-action and undo without prompting, so the reversal primitive was right. Autonomy stayed gated regardless. Host isolation already routes through three people. P1 alerts never auto-execute. Across at least eight participants in four regions, every failure story ended by narrowing autonomy, never widening it.

The reframe split the product into two surfaces. In Zones 1 to 3 - email quarantine, session revoke, IOC block, low-severity playbooks - reversibility earns confidence and autonomy can graduate. In Zone 4 - production, executives, domain controllers, regulated data - reversibility helps recovery and never unlocks autonomy. There the design work is the gating surface, not the automation: what the AI wanted to do, why, who approves, and what the audit trail shows.

The study moved four decisions

  • Chat scoped to clarification. Direct manipulation stayed the editing model.
  • Reversibility and scope-gating split in the architecture, so undo stops standing in for autonomy.
  • A label collision fixed. "High" marked both severity and confidence on the alert card, and analysts scan severity first, so a high-severity low-confidence alert read as high-confidence. Confidence moved off the severity word into a "why this score" drilldown.
  • The round ended on a saturation call, once usability scores stabilized, freeing the team to move to the next surface.
Usability (SUS) across six weeks
System Usability Scale trend. Scores of 90, then 78.8, 71.5, 71.3 and 72 across the study, dropping from the first session then settling into a plateau around 71 to 72. 100 90 80 70 plateau = saturation signal 90 78.8 71.5 71.3 72 wk 1 wk 2 wk 6
SUS charted on a 60 to 100 window to make the trend legible.

The drop is not decline. It tracks a prototype thickening from a labelled skeleton to a full build, and an audience widening past the target analyst to admins, CISOs, and managers. Settling at 71 to 72 was the saturation signal to stop.

Recruit dedicated junior (L1) analysts for the label task. The senior-heavy panel kept missing the silent misread on the alert card, because seniors already hold the mental model that hides it. The people most exposed to a confusing label were the least represented in the room.

Autonomous systems earn trust by removing drudgery and keep it by staying legible. Where a wrong action costs more than undo can recover, trust stops, and better design will not move that line. Build for the gate, not the reversal.