Across 23 projects in two years I led design pods of one to three people, taking four unsolicited concepts into the production roadmap, one recognised internally for its use of generative AI.


Retainer · 2 years · 23 projects
One of the world's largest technology companies, and a leader in enterprise data protection. Its platform is how enterprises back up, replicate, and restore what they cannot afford to lose across physical, virtual, and multicloud estates. A defect in a routine journey repeats across 98% of the Fortune 500.
The product was not broken. It was mature, dense, years deep in accumulated capability, with a widening gap between what it could do and what its users could get out of it. Closing that gap, release after release, was the retainer.
Design Lead, working directly with the client's product, engineering, and innovation teams. Pod composition changed per sub-project.
Design Leadership, Product Strategy, Concept Pitching, Interaction Design, Design Systems, AI-Assisted Prototyping, Design QA, Stakeholder Management
Every sub-project ran the same loop, whatever its size. The last stage is the one most engagements skip, and the one that decides whether the design survived.
Open with product and engineering. Get past the feature to the need.
Restate it as a person, a journey, and a moment of consequence.
Adjacent studies and field feedback the client already held.
Flows first, then screens. Team sized to the problem.
Review the reasoning with stakeholders, not the styling.
Resolve what is contested. Land a version the room accepts.
Specs, states, and component behaviour, ready to build from.
Check the shipped journey against the spec. Close the drift.
Each began as an observation rather than a brief, was pitched rather than requested, and was accepted. Together they trace the arc of the work: fixing what existed, answering what the platform could not, then preventing failures it once only reported.
A platform-wide audit, then component rules in the design system. The same defects stopped returning each release.
One search that answers the question instead of listing records. Four screens of cross-referencing became one.
Six separate status readings drawn as one chain, with one rating. A judgement the user can act on.
Every policy on one calendar, with a warning as a time is picked. Clashes prevented, not reported.
Nobody in the product organisation could say what good looked like from the user's side. Every team held its own standard, so the same defects came back release after release, in good faith.
Platform-wide, not a sample. Every journey a real user takes.
Navigation, content, presentation, interaction. Not screen by screen.
Memory, motor, intellectual, or visual, each tied to a broken principle.
Quick wins to ship now, and a northstar to aim the roadmap at.
Every major journey went onto one board, grouped by navigation, content, presentation, and interaction rather than by screen, so a defect read as a pattern across the product instead of a one-off. Each finding named the user it slowed down and the heuristic it broke, then scored the load it added: memory, motor, intellectual, or visual. Severity argued from Nielsen's heuristics and cognitive load theory is severity nobody can dismiss as taste. The board then split into two lists: fixes for the next release, and the product worth aiming at.

The defects that kept returning were written into the design system as rules for how a component behaves: how a toast times in and out, what an empty state says, where an action sits. That is Nielsen's consistency and standards heuristic made enforceable, held in the component rather than in a slide deck nobody reopens. The spec shown here sets toast timing against reading speed, so the same message is not on screen for one second in one place and eight in another. Designers and engineers now build from the same page, so the fixes hold instead of being reintroduced next release.

Users could find records but not answers. Answering why last night's backup failed meant filtering jobs, cross-referencing policies, and checking storage across four screens. Everything was in the system, almost none of it in one place.
Built the case on support tickets and field feedback already held.
Enough evidence to name the problem and take it to product.
From finding records to answering the question behind them.
Working versions in front of stakeholders, not static flows.
The screens below are wireframes used to pitch and pressure-test the concept, not final visual design.
Search sits in the top bar on every page. Type a name and results come back grouped as assets, policies, jobs, and alerts, each with a count: Gestalt similarity doing the scanning, so people read the group they want instead of sixty rows. The labels are the product's own words, matching the system to the language users already have rather than teaching them a second vocabulary. If the question is bigger than a lookup, the last row carries the same query into Advanced Search, so nothing is typed twice.

Instead of an empty field, Advanced Search opens with real questions you can click, such as which policies had errors in the last 7 days, and below them the searches you saved or ran recently. That is recognition over recall: nobody has to remember the wording that worked last week. An empty state is also the one screen a user reads before they have anything to do, which makes it the cheapest place in the product to teach it.

Ask why last night failed and the top of the screen answers it: 14 failures, up 10% on last week, and the policy behind most of them. Progressive disclosure keeps the detail underneath, where the failures are grouped by cause with the fix inside each group, so eight storage failures restart together without leaving the screen. Putting the action next to the problem it fixes is what removes the trip. Any record is still one click away. It used to take four screens.

An asset is not one thing. Source, policy, backup, copies, replication, and restore history each fail differently and were each reported separately, so an asset could read as available while its policy was disabled and its copies had expired. Nothing turned those facts into a judgement, so every user made their own.
Poor, fair, or good, defined with product before anything was drawn.
Six component ratings resolved into one judgement per asset.
Approved on the definition alone, then designed end to end.
Generated summaries over deterministic data the user can check.
Source, policy, backup, copies, configuration, and restore history each used to be reported on a different page. Here they are drawn as one chain, left to right, each part carrying its own rating of poor, fair, or good. Gestalt continuity is what makes it read as one dependent sequence rather than six widgets, and colour does the finding before any label is read. Troubleshooting becomes following the line to the part that is red, and the panel on the right opens that part in place, so the whole picture stays on screen.

The AI tab reads the whole chain and states it plainly: copies have expired, the policy is disabled, recent jobs failed, this asset cannot be restored right now. Then it recommends what to do, in the order that matters. The summary never replaces the data: everything it says sits one tab away in the Analysis view, so a user can verify it rather than trust it. That is Shneiderman's high automation held under high human control, and it is what stops a generated sentence being the only place a fact exists.

Different administrators create policies independently, and nobody could see the whole schedule. Backup windows collided, an hour took on more load than the infrastructure could carry, and jobs failed in a way that looked like an infrastructure fault but was a scheduling one. The most common complaint in the field, and the platform's only answer was to report it the next morning.
The last moment a user could have known about a collision.
The shared calendar already solves booking over other people.
Warn while the time is chosen, not in the next morning's report.
Concept accepted on the pitch. Now in build.
Every policy sits on one week, at the hour it actually runs. It reads like a shared calendar because that is Jakob's Law: people already use one to avoid booking over each other, so the conventions transfer and nothing has to be taught. The list on the left switches what the week is drawn by, whether policies, streams, assets, or storage target, so each view answers one question rather than one view carrying all four.

A backup uses streams, and there is a fixed number of them. Too many jobs in one hour and some of them fail. Switching the week to stream counts colours every hour green, orange, or red, which hands the work to preattentive processing: a crowded night is seen before it is read, and nobody adds the numbers up. Optimise Schedule then points at the hours carrying too much and suggests what to move.

The person creating a policy has the least idea of what else runs that night, and that is exactly when they pick a time. So the same data comes to them: choose a crowded slot and the form says so, then offers three times that are not. Poka-yoke, and error prevention ahead of error messages, which Nielsen ranks as the stronger of the two. It stops the failure instead of reporting it, where the clash used to surface a week later in a job that looked like an infrastructure fault.

Across 23 projects in two years, every pod worked in the same Figma design library, and each project left it better than it found it. We added components, improved the ones already there, made them consistent, and wrote the guidelines for how each is used, so any designer knew which component to reach for and how it behaves.
A separate frame for every state became one component you configure. Fix it once and the fix lands everywhere, instead of in 23 places.
Every colour and text style got a name and a job, so a designer picks error or body copy rather than judging a hex code by eye. That stopped the slow spread of nine greys nobody could tell apart.
Working prototypes in hours, on every project rather than one trial. Stakeholders reviewed three versions that ran instead of one static flow, and an idea nobody asked for became cheap enough to pitch.
Sizes, states, and behaviour annotated in the file itself. Engineers build from a link, and design QA has something exact to check the build against.
Partway through the retainer I brought AI-assisted prototyping into every pod, on every project, not as a trial on one.
Figma Make for prototyping, across every pod. Copilot for transcription and UX copy, so decisions were captured live and wording was argued rather than agonised over.
AI went only where a person would otherwise have to work something out: summaries, explanations of anomalies, recommendations. It always sits on top of real data the user can open and check, so nothing the model says is the only place that fact exists.
Every concept here started with something seen in the field, a pattern in support tickets, or a problem nobody had named yet. The noticing was the work, and nothing automated it. Building got faster. Deciding what was worth building did not change hands.
Four self-initiated interventions accepted into the roadmap. Two live, two in build.
Internal recognition from the client's innovation team for the use of generative AI in global search.
Component behaviour written into the design system, so the reasoning outlived the engagement.
One shared library, extended and given behaviour rules across 23 projects, so a fix landed once and held.
On a long retainer, the work grows by bringing problems the client has not asked about. The ones that landed were the ones where I could show the cost of leaving it alone. So lead with the cost, not the idea: name what the current behaviour is costing before proposing anything to replace it.
No research budget constrains the evidence, not the rigour. Field feedback, support patterns, and the client's own prior studies carry a concept a long way. Label what is still assumption when you present it, and the concept survives scrutiny instead of collapsing under it.
Check Out Other Projects
Rebuilding the UPI App India Was Given
I led the revamp of India's own UPI app, setting the four principles the product now runs on and shipping Split Expenses, Spends Analysis, and Family Mode.
READ CASE STUDYDesigning Inclusive Credit for Neurodiverse Customers
Designed a new design language and credit-scoring product for Barclays' neurodiverse customers.
READ CASE STUDYTransforming Credit-Card CX for 17 Million Customers
Designed and delivered a unified credit-card ecosystem, cutting support queries by 28%.
READ CASE STUDY