Case 04 · Red Hat · Fixing at scale

The Fix button that didn't fix

There was a button in our product labeled Fix, and pressing it didn't fix anything, it printed you instructions, like calling a plumber and getting a DIY guide in the mail. This case is about making that button honest, across three products that each had their own idea of what fixing means.

Products Red Hat Lightspeed · RHEL · Satellite · Ansible Automation Platform
My role Lead designer · Authored the UX strategy proposal
Timeframe 2023 strategy → ongoing implementation
Status Shipped · In production
The problem

Fixing problems on company servers ran through three different products, one that finds the problems, one that automates the work, one that tracks the machines. Three tools, three teams, three vocabularies. And the button labeled "Remediate" didn't remediate, it generated a script and left you to figure out the rest. The symptom and the disease, both at once.

The solution

We renamed the central object, from "playbook" (a script) to "remediation plan" (a thing you execute), added a shopping-cart pattern for collecting problems from anywhere in the product, and made the "Fix" button actually fix. One workflow that survives all three products.

86%
Less time on remediation. Red Hat's public headline number, validated by a Principled Technologies study (October 2024) on a 90-host RHEL environment.
Try both buttons
Before
outdated packages weak security setting misconfigured service
playbook-4821.yml

After
outdated packages weak security setting misconfigured service

Same three problems on both tables. One button hands you homework; the other does the job.

Research and discovery

The users are the people who fix problems across a company's server fleet. Some live in the monitoring tool, some live in the automation tool, and most work somewhere between the two. The design problem was to make the whole workflow feel like one thing, no matter where it started or ended.

I ran a structured heuristic evaluation of the existing remediation workflow and documented 54 distinct issues across pre-remediation, execution and post-remediation. That single artifact ended up shifting leadership conversations from "polish the existing flow" to "the model is wrong" in one meeting.

The button labeled "Remediate" didn't actually remediate. The button was the symptom; the script-as-only-output was the disease.

Once that pattern had a name, the redesign more or less followed from it. Adoption data backed the direction up too: quarterly readouts in 2023 showed nearly 3,000 remediations executed per quarter, which gave us real usage to work against rather than assumptions.

Key UX moves

1From "script" to "plan", the noun that changed the design

This was a conceptual shift, not a rename. A plan can hold manual steps, automated steps, scheduling and dependencies. A script can't. Changing the noun changed the design problem and gave three teams a shared object that made sense in all three products. It's also the natural next step from Pathways. Pathways bundles the problems into a unit of work; the plan takes that unit and actually executes it.

2A shopping cart instead of a corridor

The old flow was a wizard, a one-way corridor. If you wanted to gather problems from different corners of the product into one fix, every detour reset your progress. The new pattern works like a shopping cart: browse anywhere, drop problems in as you find them, check out once. Nobody loses their place.

Finds the problems Automates the work Tracks the machines One remediation plan manual + automated + scheduled "Fix" finally means fix
Three tools, one plan, one honest button.

3Start from the machine, not the problem

The old entry points started from "a problem" and asked "which machines?" Usage data showed people work the other way round, they pick a machine and want to resolve everything wrong with it at once. The redesign made the machine the primary unit, and the "Fix" button moved to where the machine lives, where it could finally mean what it says.

4What we cut

The old "Create playbook" flow used to end with "Go back to application." Not "Execute." Not "Download." The action that completed the workflow was a navigation, not a resolution. We cut it. The flow now ends where users expect it to end, at a decision about what to actually do next.

Challenges

The three product teams worked in parallel rather than together from the start, so some of the alignment happened after big architectural decisions were already made. The plan model held up, but the seam between "fix it here" and "fix it at scale" still feels like a handoff in places where it should feel like a continuation.

High fidelity walkthrough

Shipped in Red Hat Lightspeed. The remediation workflow is live in the Hybrid Cloud Console, with Red Hat publishing the announcement, the documentation and the customer-portal explainer alongside it.

Read the Red Hat announcement →
Lightspeed Remediations guide →

Final takeaways

  1. When a button label and the action it triggers don't match, the conceptual model underneath is broken. Naming the disease, not just the symptom, is the first real design move.
  2. A drawer-and-cart pattern is the right shape whenever users need to accumulate context across surfaces. Wizards quietly punish any kind of lateral exploration.
  3. Heuristic evaluation can be a leadership tool, not just a research one. Cataloguing 54 issues with severity is what shifted the conversation from "polish" to "rebuild" in a single meeting.
  4. What I'd push harder for next time: joint design with the automation and inventory teams from day one. Seams between products are far easier to design upstream than to retrofit once the model is already set.

Public proof and customer evidence

Principled Technologies study (commissioned by Red Hat, October 2024) tested a 90-host RHEL environment, manual scripting vs Lightspeed automated remediation. (Source)

Measured reduction Time on known issues 86% less time on known issues −86% Time on critical vulnerabilities 79% less time on critical vulnerabilities −79% Unplanned downtime 76% less unplanned downtime −76% Plus $103,500 average annual benefit per 100 servers.
Manual scripting vs the new workflow, 90-host environment. Principled Technologies, October 2024.

"Previously, we could only view monthly reports. With Red Hat Lightspeed we have real-time visibility that helps us make smarter decisions about processes."

Tomago Aluminium · Red Hat customer case study

Read the Tomago Aluminium case →

Further reading

Red Hat Lightspeed 2025: From observability to actionable automation Tackle critical vulnerabilities with the new Lightspeed remediation workflow Customer Portal: What is Red Hat Lightspeed (Insights) Remediation Red Hat Developer learning path: Create a remediation playbook Executing remediation plans documentation Lightspeed Remediations guide Red Hat Insights Remediations improvements Principled Technologies study. Modernize RHEL monitoring and remediation Tomago Aluminium case study Save time, effort, costs with Red Hat Lightspeed RedHatInsights/insights-remediations on GitHub Red Hat Insights is now Red Hat Lightspeed