How might we move an enterprise platform from reacting to what went wrong to preventing it?

Across 23 projects in two years I led design pods of one to three people, taking four unsolicited concepts into the production roadmap, one recognised internally for its use of generative AI.

Policy calendar showing every policy and its jobs across a week
000

Project Overview

Retainer · 2 years · 23 projects

The Client

One of the world's largest technology companies, and a leader in enterprise data protection. Its platform is how enterprises back up, replicate, and restore what they cannot afford to lose across physical, virtual, and multicloud estates. A defect in a routine journey repeats across 98% of the Fortune 500.

The Business Challenge

The product was not broken. It was mature, dense, years deep in accumulated capability, with a widening gap between what it could do and what its users could get out of it. Closing that gap, release after release, was the retainer.

The Engagement

Design Lead, working directly with the client's product, engineering, and innovation teams. Pod composition changed per sub-project.

My Role

Design Leadership, Product Strategy, Concept Pitching, Interaction Design, Design Systems, AI-Assisted Prototyping, Design QA, Stakeholder Management

001

The Delivery Model

Every sub-project ran the same loop, whatever its size. The last stage is the one most engagements skip, and the one that decides whether the design survived.

  1. 01

    Question the Request

    Open with product and engineering. Get past the feature to the need.

  2. 02

    Make It About the User

    Restate it as a person, a journey, and a moment of consequence.

  3. 03

    Ground It in Evidence

    Adjacent studies and field feedback the client already held.

  4. 04

    Design the Journey

    Flows first, then screens. Team sized to the problem.

  5. 05

    Walk It Through

    Review the reasoning with stakeholders, not the styling.

  6. 06

    Revise and Close

    Resolve what is contested. Land a version the room accepts.

  7. 07

    Hand Over to Build

    Specs, states, and component behaviour, ready to build from.

  8. 08

    Design QA the Build

    Check the shipped journey against the spec. Close the drift.

002

Four Interventions

Each began as an observation rather than a brief, was pitched rather than requested, and was accepted. Together they trace the arc of the work: fixing what existed, answering what the platform could not, then preventing failures it once only reported.

  1. Expert Reviews

    A platform-wide audit, then component rules in the design system. The same defects stopped returning each release.

  2. AI-Enabled Global Search

    One search that answers the question instead of listing records. Four screens of cross-referencing became one.

  3. Asset Health and Analysis

    Six separate status readings drawn as one chain, with one rating. A judgement the user can act on.

  4. The Protection Policy Calendar

    Every policy on one calendar, with a warning as a time is picked. Clashes prevented, not reported.

003

Expert Reviews

Nobody in the product organisation could say what good looked like from the user's side. Every team held its own standard, so the same defects came back release after release, in good faith.

What We Did

  1. 01

    Audit Every Major Journey

    Platform-wide, not a sample. Every journey a real user takes.

  2. 02

    Group by Four Lenses

    Navigation, content, presentation, interaction. Not screen by screen.

  3. 03

    Rate the Cost to the User

    Memory, motor, intellectual, or visual, each tied to a broken principle.

  4. 04

    Split Into Two Lists

    Quick wins to ship now, and a northstar to aim the roadmap at.

The Design Interventions

01

One Board for Every Problem in the Product

Every major journey went onto one board, grouped by navigation, content, presentation, and interaction rather than by screen, so a defect read as a pattern across the product instead of a one-off. Each finding named the user it slowed down and the heuristic it broke, then scored the load it added: memory, motor, intellectual, or visual. Severity argued from Nielsen's heuristics and cognitive load theory is severity nobody can dismiss as taste. The board then split into two lists: fixes for the next release, and the product worth aiming at.

Expert review board organised into Navigation, Content, Presentation, and Interaction
02

Rules in the Design System, Not a Slide Deck

The defects that kept returning were written into the design system as rules for how a component behaves: how a toast times in and out, what an empty state says, where an action sits. That is Nielsen's consistency and standards heuristic made enforceable, held in the component rather than in a slide deck nobody reopens. The spec shown here sets toast timing against reading speed, so the same message is not on screen for one second in one place and eight in another. Designers and engineers now build from the same page, so the fixes hold instead of being reintroduced next release.

Toast behaviour specification: anatomy, and a timeline of 200ms reveal, 4000ms dwell, 150ms dismiss
004

AI-Enabled Global Search

Users could find records but not answers. Answering why last night's backup failed meant filtering jobs, cross-referencing policies, and checking storage across four screens. Everything was in the system, almost none of it in one place.

What We Did

  1. 01

    No Research Budget

    Built the case on support tickets and field feedback already held.

  2. 02

    Pitch It Unasked

    Enough evidence to name the problem and take it to product.

  3. 03

    Reframe What Search Is For

    From finding records to answering the question behind them.

  4. 04

    Prototype in the Open

    Working versions in front of stakeholders, not static flows.

The Design Interventions

The screens below are wireframes used to pitch and pressure-test the concept, not final visual design.

01

One Search Box, Results Grouped by Type

Search sits in the top bar on every page. Type a name and results come back grouped as assets, policies, jobs, and alerts, each with a count: Gestalt similarity doing the scanning, so people read the group they want instead of sixty rows. The labels are the product's own words, matching the system to the language users already have rather than teaching them a second vocabulary. If the question is bigger than a lookup, the last row carries the same query into Advanced Search, so nothing is typed twice.

Global search returning results grouped into asset, policy, job, and alert, with a route into Advanced Search
02

Advanced Search Shows What to Ask

Instead of an empty field, Advanced Search opens with real questions you can click, such as which policies had errors in the last 7 days, and below them the searches you saved or ran recently. That is recognition over recall: nobody has to remember the wording that worked last week. An empty state is also the one screen a user reads before they have anything to do, which makes it the cheapest place in the product to teach it.

Advanced Search entry state offering example questions, saved searches, and recent searches
03

The Answer First, the Records Underneath

Ask why last night failed and the top of the screen answers it: 14 failures, up 10% on last week, and the policy behind most of them. Progressive disclosure keeps the detail underneath, where the failures are grouped by cause with the fix inside each group, so eight storage failures restart together without leaving the screen. Putting the action next to the problem it fixes is what removes the trip. Any record is still one click away. It used to take four screens.

Advanced Search results: summary tiles above failure groups with inline recovery actions and optimisation recommendations
005

Asset Health And Analysis

An asset is not one thing. Source, policy, backup, copies, replication, and restore history each fail differently and were each reported separately, so an asset could read as available while its policy was disabled and its copies had expired. Nothing turned those facts into a judgement, so every user made their own.

What We Did

  1. 01

    Agree What Health Means

    Poor, fair, or good, defined with product before anything was drawn.

  2. 02

    Roll It Up to the Asset

    Six component ratings resolved into one judgement per asset.

  3. 03

    Pitch the Concept

    Approved on the definition alone, then designed end to end.

  4. 04

    Keep AI Above the Data

    Generated summaries over deterministic data the user can check.

The Design Interventions

01

The Whole Asset, Drawn as One Chain

Source, policy, backup, copies, configuration, and restore history each used to be reported on a different page. Here they are drawn as one chain, left to right, each part carrying its own rating of poor, fair, or good. Gestalt continuity is what makes it read as one dependent sequence rather than six widgets, and colour does the finding before any label is read. Troubleshooting becomes following the line to the part that is red, and the panel on the right opens that part in place, so the whole picture stays on screen.

Asset chain from source through policy, backup, copies, and restore, with the analysis panel open
02

An AI Summary That Says What Is at Risk

The AI tab reads the whole chain and states it plainly: copies have expired, the policy is disabled, recent jobs failed, this asset cannot be restored right now. Then it recommends what to do, in the order that matters. The summary never replaces the data: everything it says sits one tab away in the Analysis view, so a user can verify it rather than trust it. That is Shneiderman's high automation held under high human control, and it is what stops a generated sentence being the only place a fact exists.

The same asset with the AI summary and recommendations panel
006

The Protection Policy Calendar

Different administrators create policies independently, and nobody could see the whole schedule. Backup windows collided, an hour took on more load than the infrastructure could carry, and jobs failed in a way that looked like an infrastructure fault but was a scheduling one. The most common complaint in the field, and the platform's only answer was to report it the next morning.

What We Did

  1. 01

    Find the Earliest Warning

    The last moment a user could have known about a collision.

  2. 02

    Borrow a Trusted Model

    The shared calendar already solves booking over other people.

  3. 03

    Move the Moment

    Warn while the time is chosen, not in the next morning's report.

  4. 04

    Pitch and Hand Over

    Concept accepted on the pitch. Now in build.

The Design Interventions

01

Every Policy on One Calendar

Every policy sits on one week, at the hour it actually runs. It reads like a shared calendar because that is Jakob's Law: people already use one to avoid booking over each other, so the conventions transfer and nothing has to be taught. The list on the left switches what the week is drawn by, whether policies, streams, assets, or storage target, so each view answers one question rather than one view carrying all four.

Policy calendar showing every policy and its jobs across a week
02

A Heat View of How Busy Each Hour Is

A backup uses streams, and there is a fixed number of them. Too many jobs in one hour and some of them fail. Switching the week to stream counts colours every hour green, orange, or red, which hands the work to preattentive processing: a crowded night is seen before it is read, and nobody adds the numbers up. Optimise Schedule then points at the hours carrying too much and suggests what to move.

The same calendar in stream count view, load rendered as a heat map
03

A Warning While the Time Is Being Set

The person creating a policy has the least idea of what else runs that night, and that is exactly when they pick a time. So the same data comes to them: choose a crowded slot and the form says so, then offers three times that are not. Poka-yoke, and error prevention ahead of error messages, which Nielsen ranks as the stronger of the two. It stops the failure instead of reporting it, where the clash used to surface a week later in a job that looked like an infrastructure fault.

Inline time slot recommendations offered during backup schedule creation
007

The File Behind the Work

Across 23 projects in two years, every pod worked in the same Figma design library, and each project left it better than it found it. We added components, improved the ones already there, made them consistent, and wrote the guidelines for how each is used, so any designer knew which component to reach for and how it behaves.

  1. One Component, Every State

    A separate frame for every state became one component you configure. Fix it once and the fix lands everywhere, instead of in 23 places.

  2. Styles You Pick by Name

    Every colour and text style got a name and a job, so a designer picks error or body copy rather than judging a hex code by eye. That stopped the slow spread of nine greys nobody could tell apart.

  3. Figma Make on Every Project

    Working prototypes in hours, on every project rather than one trial. Stakeholders reviewed three versions that ran instead of one static flow, and an idea nobody asked for became cheap enough to pitch.

  4. Dev Mode Is the Spec

    Sizes, states, and behaviour annotated in the file itself. Engineers build from a link, and design QA has something exact to check the build against.

008

Working With AI, Not Only Designing It

Partway through the retainer I brought AI-assisted prototyping into every pod, on every project, not as a trial on one.

In the Process

Figma Make for prototyping, across every pod. Copilot for transcription and UX copy, so decisions were captured live and wording was argued rather than agonised over.

In the Product

AI went only where a person would otherwise have to work something out: summaries, explanations of anomalies, recommendations. It always sits on top of real data the user can open and check, so nothing the model says is the only place that fact exists.

Not in the Thinking

Every concept here started with something seen in the field, a pattern in support tickets, or a problem nobody had named yet. The noticing was the work, and nothing automated it. Building got faster. Deciding what was worth building did not change hands.

009

Impact

  1. 01

    Four self-initiated interventions accepted into the roadmap. Two live, two in build.

  2. 02

    Internal recognition from the client's innovation team for the use of generative AI in global search.

  3. 03

    Component behaviour written into the design system, so the reasoning outlived the engagement.

  4. 04

    One shared library, extended and given behaviour rules across 23 projects, so a fix landed once and held.

010

Key Learnings

  1. 01

    On a long retainer, the work grows by bringing problems the client has not asked about. The ones that landed were the ones where I could show the cost of leaving it alone. So lead with the cost, not the idea: name what the current behaviour is costing before proposing anything to replace it.

  2. 02

    No research budget constrains the evidence, not the rigour. Field feedback, support patterns, and the client's own prior studies carry a concept a long way. Label what is still assumption when you present it, and the concept survives scrutiny instead of collapsing under it.