From the first partner interview to the pilot that proved it
No brief, no research team. I framed the investigation, ran all ten partner interviews and both rounds of usability testing, then made the call that defined the product — benefits above the fold instead of the performance-banner-first layout the brief had specified — and carried it through SVP-level review and sign-off across markets. Lo-fi to production, token architecture, design QA, pilot.
Two decisions were mine to own and defend: reversing the layout the brief specified, and refusing to let 26 markets each get their own version. The first needed evidence, so I built the test that produced it. The second needed a system, so I built that instead of a backlog.
Late in the project a junior designer joined to support development. I mentored him directly through the build — covered in Deliver.
The team is what made it possible. My Design Manager pushed the thinking further every time I brought them something half-formed. Our Product Manager built the roadmap that got this funded. Two frontend and two backend engineers pushed back on component complexity — and they were right; that’s why the level comparison shipped as a reusable component instead of a one-off. And our Data Analyst built the Eppo experiment from day one, not at QA. Without that, the +120% would have been an anecdote.
Every explanation on the table pointed at the partners. They weren’t ambitious enough. They didn’t care about the badge. The incentives weren’t generous enough to be worth chasing. None of it survived ten conversations with the people the program was built for.
So I reframed the brief before agreeing to it. Not “redesign the rewards page” but something harder and narrower: find out why a well-funded incentive isn’t pulling, prove it, and only then decide what to change. That reframing is what bought the research time — and what made the eventual result attributable to a diagnosis rather than to a redesign.
I ran it as a double diamond: diverge to understand the problem, converge on a diagnosis, diverge on solutions, converge on what shipped.
Four phases, one continuous loop. Research never stopped at the first diamond — testing and partner feedback ran straight through ideation and into the prototype.
Context
The program this case study is aboutA performance incentive nobody was using
Delivery Hero runs food delivery platforms in 50+ countries, with a tiered partner rewards program active in 26 of them, and restaurant and store partners fulfil every single order — their operational performance is the customer experience. The Partner Rewards Program was built to move that performance: partners are scored monthly across seven metrics and placed into one of four levels. Higher levels unlock better placement, lower commissions, ad credits, and a top-restaurant badge — and below the entry bar of 20 orders a month, a partner isn’t in the program at all.
This is also why the brief was never cosmetic. Leadership had set explicit KPI targets against this program, and what landed on my desk read as our performance incentive system has stopped working — not make the rewards page nicer.
It also arrived with a structural risk attached. Markets were each asking for their own version of a fix, which would have produced 26 diverging rewards experiences and no way to attribute a result to any of them. I pushed back on that early and committed the team to a single solution: one system, tokenised per brand, no per-market forks. That constraint shaped every decision that follows — and it is the reason the rollout later became a schedule rather than 26 separate projects.
01 · Discover
Diverge — exploring the problemWhere the motivation actually broke
Before drawing a single screen I needed to see the program the way a partner sees it — not the way the internal deck described it. So I sat with restaurant and store partners across markets and watched them open the rewards page cold: what they looked at first, where they stalled, what they said out loud, and at what point they closed the tab without doing anything. Then I looked outward at how loyalty and performance programs in other industries solve the same motivational problem.
Ten interviews in, the pattern was not subtle. Every partner could find their level badge. Not one could explain what it was worth, which of the seven scored metrics had moved them there, or what specifically to do next. They weren’t unmotivated. They were being asked to chase a target nobody had described.
What ten interviews sounded like
“Which KPI matters most?”
Metric confusion“What do I get for Tier 3?”
Hidden benefits“I feel stuck.”
No momentum“I couldn’t find my tier.”
Poor visibilityFour reactions, ten interviews, one root cause underneath all of them. Every partner here had the data on screen the whole time. The third face — “I feel stuck” — returns at the end of this case study, after the redesign.
The C-level was asking a business question, not a design one
In a monthly business review, Delivery Hero’s SVP and VP put rewards on the agenda as a performance problem, not a design critique. The company had funded a tiered incentive to move orders, rejections and offline time. With 70% of partners stuck at the lowest level, none of it was moving — and those gains were attached to targets.
Then they walked the live flow themselves and came back with three problems: goal visibility too small to act on, benefits below the fold, and locked chips inconsistent with the rest of the platform.
I took each one in the room and sorted it rather than accepting the list wholesale. Two of the three independently matched what partner interviews were already surfacing — which made them evidence, not opinion, and I said so. The third, the locked chips, was a design-system inconsistency rather than a motivation problem; I logged it as one and fixed it there instead of treating it as a redesign requirement.
All three points were acted on. None of them shipped because of who raised it. The two that concerned the rewards experience were tested with partners in both rounds before anything was committed; the third was fixed in the design system. Handling it that way mattered beyond this project: it established that design decisions here would be argued on evidence, including when the person on the other side of the table outranked everyone in it.
The review also reframed the work. The lever wasn’t missing — it was disconnected. The incentives were already generous; partners just couldn’t see what they were worth. I translated that into a brief the business could hold me to: make this program capable of hitting the numbers it was funded to hit — and agreed the metric it would be judged on with the PM and Data Analyst before design started.
What the numbers said
The qualitative signal was consistent, but it needed a floor under it. I pulled a baseline with our Data Analyst before any design work started, so the comparison at the end would be against a recorded number rather than a remembered one. It was worse than the anecdotes suggested.
Pre-redesign baselineBefore any of this changed
| Criteria | Uber Pro | Starbucks | Wolt | Samsung |
|---|---|---|---|---|
| Benefits visible above the fold | ||||
| Progress indicators | ||||
| Side-by-side tier comparison | ||||
| Contextual progress to next tier | ||||
| Evaluation / reset date shown | ||||
| Mobile-first design | ||||
| RTL + multi-brand scalability |
02 · Define
Converge — committing to a problemEvery number was already on the page. None of it told a partner what to do.
Ten interviews, a quantitative baseline and an executive review all pointed at the same screen. Define was about naming what was actually broken underneath it — and resisting the five obvious fixes that would only have treated symptoms.
One page, two different problems
The business problem: Delivery Hero had invested heavily in a tiered incentive and the lever wasn’t pulling. Partners weren’t improving the metrics the business needed to move, so the money behind the program was buying nothing.
The partner problem: a level label with no explanation of what it meant, which metrics controlled it, or what to do next. The program felt arbitrary, so partners treated it that way.
Five failure modes — structural, not surface-level
All five surfaced in every interview — the same walls, hit in the same order.
It pushed the evaluation date and progress data out of sight. Partners didn’t know where to look for what they came for.
No priority, no weighting, no action. Knowing your rejection rate is 8% is useless if nothing says whether that’s the number holding you back.
The whole motivational payload sat below the fold — invisible at the exact moment we asked partners to earn it.
Progress was happening; the interface never showed it — so it may as well not have been.
Level 02 was a label someone else had — no requirement, no date, nothing to act on. Under 20 orders a month the card simply locked. A system can track you without ever guiding you.
None of it was inferred. In their words, at length:
I couldn’t find my current tier. The banner takes up half the screen and tells me nothing.
Which of these actually matters most? There are seven numbers and no order to them.
What do I get for Tier 3? I’d work for it if I knew what it was.
I feel stuck. There’s no forward momentum — nothing tells me I’m getting closer.
Restaurant partners · ten interviews across markets
From five symptoms to one root cause
The temptation was to treat these as five separate fixes — a cleaner banner, a metric hierarchy, benefits moved up, a progress bar, a nudge. That would have been the wrong read. Every one of them traced back to the same thing.
A motivation and clarity problem — not an information problem.
Every number a partner needed was already on the page, rendered and accurate. It was cognitively unavailable: buried below the fold, unranked, undated, stripped of any sense of progress, and attached to nothing a partner could act on.
Reduce, surface, and sequence. Don’t add.
The brief that came out of synthesis was subtractive. Every proposal from here had to earn its place against it — including the ones that would have been easier to build.
Design principles, not just opinions
Three principles governed every decision from here on, each anchored to an established usability heuristic so a design argument could be checked against something other than my own judgment.
Clarity over completeness
Show the highest-value action and remove the rest. Every metric weighted equally is the same as none of them mattering.
“Every extra unit of information in an interface competes with the relevant units of information and diminishes their relative visibility.” NN/g heuristic #8 · Aesthetic and Minimalist DesignFairness & transparency
Only metrics a partner controls, rules in plain language, a visible evaluation date. Otherwise it’s a grade, not an incentive.
“The design should always keep users informed about what is going on, through appropriate feedback within a reasonable amount of time.” NN/g heuristic #1 · Visibility of System StatusScalability by design
One system, 26 markets, no per-brand redesigns. Four brands and full RTL, built in from the first component.
“Users should not have to wonder whether different words, situations, or actions mean the same thing.” NN/g heuristic #4 · Consistency and StandardsQuotations are the published wording of Jakob Nielsen’s ten usability heuristics (1994, NN/g). Naming one heuristic per principle gave the team an external standard to argue against — so disagreements were about a principle, not about taste.
Those principles pointed at two questions worth designing against. Both were written down before any layout work started:
Make partner performance feel like a path to growth rather than a grading system?
Put benefits first, so partners understand the reward before we ask them to earn it?
03 · Develop
Diverge — exploring solutionsTwo rounds. Because the problem was already settled.
By the time design started, the hard part was behind us. We knew what was broken, which principle governed each fix, and what the business needed to move. That is what a properly defined problem buys you: design stops being exploration and becomes execution. Two rounds to sign-off — and partners in the room for both of them.
Two rounds, and what changed between them
Neither round was a restart. Round one tested the structure; round two added the one thing partners said was missing. Same eight partners (a subset recruited from the ten discovery interviews), same script, so the two rounds were directly comparable.
Built straight from the three principles — benefits above the fold, the four levels side by side, the evaluation date on the page. I reviewed it with the PM, engineering and my Design Manager before any partner saw it, so nothing went into testing that couldn’t be built. Partners understood the layout immediately: nobody asked what the levels were or what the benefits were, the two questions that had dominated the interviews. One gap remained, and every partner raised it — “where am I now, and how far is the next level?”
Progress bars with a named target on every metric — “3 more orders”, not “62%” — and a Smart Action beside each one. 7 of the 8 partners were satisfied, up from 3 on the old experience. More telling than the score: 6 of the 8 could identify their next action in the program without any prompting — something none of them could do on the old design. They stopped critiquing and started asking when it would launch.
That last shift mattered more than the satisfaction score. A partner who feels a program was built with them will work toward it. One who feels it was applied to them will ignore it — which is exactly what 70% of them had been doing.
Eight partners in a room is a strong signal, not proof. Confirming it quantitatively was Deliver’s job. The experiment had been scoped with the Data Analyst from the project’s start — not bolted on at QA — so when design reached sign-off the test infrastructure was already in place.
The calls I made, and why
Five decisions defined the product. Each had a real trade-off behind it — an alternative considered, a constraint weighed, someone to convince.
- 1Benefits above the fold, not performance data — this contradicted the original brief
The design spec I inherited led with a 40%-of-screen performance card and buried the benefits underneath it. The cheaper option was to shrink the banner and leave performance on top — far less work, and it would have technically satisfied the complaint. I rejected it because it treated a hierarchy problem as a spacing problem: a smaller banner still puts the ask before the reward. Reversing the page order meant saying the spec was wrong, not just tight, so I didn’t argue it on taste — I built it to the principle and put it in front of partners before asking anyone to approve it. Chosen: benefits-first — matching Starbucks and Uber Pro, the two strongest programs in the audit.
- 2A side-by-side level comparison — the thing no competitor offers
Partners knew their level, not what separated them from the next — and that gap is where motivation died. Engineering flagged the build complexity early and they were right: as a one-off it would have broken the moment a second brand needed it. The trade-off I accepted was a slower first delivery — more upfront component work in exchange for a table that shipped once and held across 26 markets. Chosen: the comparison table — nothing in the audit offered it outright, and only Starbucks came partway.
- 3Progress bars with named targets — and the evaluation date, for the same reason
The “stuck” feeling was perceptual — progress was happening invisibly. A bar plus a named gap made it visible; the date gave it a deadline. Chosen: progress plus date. Partners answered “what do I do next” 40% faster with a named target than with a flat metric — time-to-answer, timed across the moderated sessions.
- 4Smart Actions on every metric — the difference between tracking and guiding
Progress bars showed the gap. They didn’t close it. So every metric got an action — one tap from “50 orders short” to the tool that fixes it. Chosen: a system that only measures you is a report card; one that hands you the next step is a coach.
- 5One design system on multi-brand tokens — 26 markets, 6 brands, one codebase
Semantic tokens handled colour, type and spacing per brand, so a market launch became a token swap rather than a redesign — and RTL was built in from the first component. Chosen: the system, over 26 bespoke versions. This was my answer to the competing market requests, and it cost me the ability to give any single market exactly what it asked for.
The one that pays for the program
Four of those five made the incentive clearer. Smart Actions made it commercial — and it’s the reason this redesign funds itself.
A partner who taps Start a campaign to close an orders gap runs a campaign. That lifts their orders toward the next level and puts more volume through the platform in the same motion. Nobody had to choose between the partner’s interest and the business’s.
“50 more orders” stops being a verdict and becomes a button. The next step is chosen for them, not left as homework.
A campaign, a faster response time, a better menu score — every action a partner takes to level up is an action Delivery Hero wanted from them regardless.
Better performance unlocks better placement and lower commissions, which drives more orders, which feeds the next evaluation. The incentive stops being a cost and starts being a growth loop.
04 · Deliver
Converge — committing to scaleBenefits first. Progress made visible.
The redesign answered all five failure modes named in Define: visibility, metric comprehension, benefit discovery, momentum and next action. Four were fixed outright. One — metric comprehension — was only partly solved: every metric now carries its gap and the action that closes it, but they are still shown with equal weight. Ranking them is the one piece I scoped and did not ship, and it sits at the top of the roadmap for that reason. Mobile-first throughout, since most partners open this on a phone.
What shipped
The final output wasn’t a set of screens — it was a build package. Every state, every edge case and every error was drawn before anything reached a sprint. Design Ready meant engineering could start without a follow-up conversation.
Handoff wasn’t the end of my involvement. I stayed with the two frontend and two backend engineers through the build, reviewing implementation against the specs and resolving the questions that only surface once a component is real. This is also where the junior designer came in. I onboarded him to the token architecture and the three principles, then handed him implementation support — reviewing his QA passes rather than doing them myself, which kept the quality bar where it needed to be while freeing me for the pilot. Delegation only works if the standard is written down; the principles and the spec set were what made his work reviewable.
Motion & delights
Two moments in this experience are emotional rather than structural. I built both animations from scratch in After Effects — everything else is deliberately still.
Shown once, on first open. All three levels land before any performance data — benefits before the ask.
Shown on promotion. The badge lands, then the copy points at the next level instead of closing the loop.
Motion is reserved for these two moments. A system that celebrates everything celebrates nothing — and because level names are tokenised, Advance resolves per brand without a second file.
Across mobile and desktop, every level state, four brands and both text directions. Scored against the same seven criteria as the Discover audit, with the shipped design now in the mix: the gap it was built to close is the one nobody else had closed.
| Criteria | Uber Pro | Starbucks | Wolt | Samsung | DH Partner Rewards New |
|---|---|---|---|---|---|
| Benefits visible above the fold | |||||
| Progress indicators | |||||
| Side-by-side tier comparison | |||||
| Contextual progress to next tier | |||||
| Evaluation / reset date shown | |||||
| Mobile-first design | |||||
| RTL + multi-brand scalability |
Accessibility and global readiness
WCAG, RTL and localisation were built in from day one rather than retrofitted per market — a decision that cost more at the first component and less at every one after it.
dir changes
RTL was built in, not retrofitted. Every component uses CSS logical properties — margin-inline-start, padding-inline-end, text-align: start — so text direction, progress fill, icon placement, badge position and spacing all mirror on their own. The two cards above are the same markup; the only difference is dir="rtl". That meant no separate Arabic design file, and zero per-market layout rework across talabat and HungerStation.
Quantitative validation
Designing the test was the part that decided whether any of this would scale. Comparing two markets wouldn’t have produced clean data — partner density and seasonality differ everywhere, so a cross-market read measures the market, not the design. The only clean test was old versus new in the same market, in the same weeks, split server-side through Eppo so nobody could tell they were in a study.
PartnersRestaurants and coffee shops
Split50 / 50, randomised server-side
PrimaryTier stagnation rate
SecondaryRewards click-through rate
- Benefits below the fold
- No progress, no target
- No evaluation date
- Benefits first
- Progress bars with named targets
- Level comparison and evaluation date
- A Smart Action on every metric
Primary resultPartners stuck at the lowest tier dropped
Secondary — proof of active engagementPartners actively opening their rewards page
Same market, same weeks, randomised split — only the design changed. Tier stagnation was the committed business metric: did partners move up levels? It dropped from 70% to 15%. Click-through rate was the engagement proxy measured in parallel — it doubled, confirming partners were actively using their rewards page rather than ignoring it. Both came from the same experiment run in Eppo. The HungerStation pilot was the basis for approving rollout to 26 markets.
- 1Benefits-first was the change that mattered most (inferred from sessions — the A/B tested the bundle, so it can’t isolate individual elements)
The A/B tested the redesign as a bundle, so on its own it can’t isolate one element — and I won’t claim it does. What points to this one is the moderated sessions: moving benefits above the fold removed more questions than any other change, and by round one nobody was asking what a level was worth. Being right about it mattered less than having the evidence to argue it.
- 2Naming the gap beat showing a percentage
“3 more orders” beat a percentage every time — partners could act on a countable target and stalled on a ratio. The cheapest change in the whole redesign, and one of the most effective.
- 3The comparison table never needed a revision
It went in during round one and shipped unchanged — because it answered a question partners had been asking since the first interview.
HungerStation goes live — then 26 markets
The quantitative validation did something no design argument can do on its own: it changed who had to be convinced. An attributable number turned a 26-market rollout from a request into a funded plan — decided on evidence rather than on whether leadership found the redesign persuasive. Within months the new experience had reached two-thirds of the partner base across the 26 reward-program markets.
The data earned the mandate; the token architecture made it cheap to act on. Had the 26 markets each been given their own version back in Discover, this is the point where the project would have stalled.
One market, chosen because it had the worst stagnation numbers — proving the fix in the hardest case made the mandate to scale impossible to argue against.
Attribution clean enough that funding the rollout stopped being a debate.
Rolled out on the multi-brand token architecture with no per-market redesign required.
Impact
The outcome — measured, not claimedStuck partners started moving. The test made it undeniable.
Every number below is attached to a method. Tier stagnation and click-through rate are from the same controlled A/B test on HungerStation. Satisfaction is from moderated usability sessions. None of it rests on the design being persuasive — it rests on the experiment.
Partners stuck at the lowest tier
Before the redesign, 70% of HungerStation partners were parked at the lowest active level and not progressing. After, only 15% were — a 55-point reduction measured in a randomised 50/50 A/B test against the same market, same weeks.
Tier stagnation and click-through are both from the same controlled A/B test on HungerStation. Satisfaction is from moderated usability sessions. Rollout to all 26 markets is ongoing — the pilot is what earned it.
Five gaps, closed
Define named five failure modes. The shipped product answers each one — four outright, and metric priority partly, with ranking still to come.
What outlasts the project
Three things survive beyond this release. One design system carrying four brands and both text directions, so a new market is a configuration rather than a project. A per-market asset library, so engineering ships without a designer in the loop. And a validated pattern for this kind of work: define the metric with the analyst first, test qualitatively, then prove it quantitatively before asking for scale.
The precedent mattered as much as the artefacts. A change to rewards now arrives with a measurement plan attached — the next proposal will be expected to bring one, including if it’s mine. That is a standard I would rather be held to than a persuasive deck.
Roadmap & Reflection
What’s ahead, and what it taught meValidated and scaling. Not finished.
This is where the story stands now, not where it ends. Two-thirds of the partner base across the 26 reward-program markets has the new experience; the remaining markets, and the ideas the data opened up, are the work still ahead.
What comes next
- 01Full global rollout — in progress. Remaining markets and brands, completing the migration off the old experience by late 2026.
- 02Personalised recommendations — planned. The shipped design gives every metric a gap and an action, but it doesn’t yet rank them. This is the piece that would: live performance data surfacing the one metric most likely to advance a given partner’s level, rather than showing all seven with equal weight.
- 03Deeper quantitative reporting — planned. GA4 and extended Eppo analysis for click heatmaps, drop-off and segment-level engagement.
- 04Gamification pilots — exploratory. Limited tests of leaderboards, milestone recognition and time-bound challenges, to find out whether engagement holds beyond the initial lift.
Underneath the quantitative result is the qualitative one. The four reactions this case study opened with — confused by the metrics, unable to find the benefits, unable to find their own level, and stuck — became seven of eight partners satisfied, up from three — and instead of critiquing the design, they were asking when it would launch. Nothing about the incentive changed. Only whether a partner could see what it was worth, and how close they were to it.
What I learned
- Benefits need visibility, not discoverability. Moving benefits above the fold had an outsized effect relative to how simple the change was. Information that exists but isn’t seen doesn’t effectively exist — and “it’s discoverable” is usually a polite way of saying nobody discovers it.
- Clarity beats feature richness, every time. Partners cared about progress bars and a comparison table. Not animations, not badges, not gamification layers. The best design decision on this project was knowing what to leave out — the brief was reduce, surface, sequence, and holding that line is what made the rest work.
- Partner with your data analyst on day one, not at QA. The quantitative result made the case for global rollout almost self-evident — but only because KPI definition and experiment design started at the beginning of the project, not after the designs were done. Quantitative data brought in late can only validate. Brought in early, it steers.
- Early and continuous feedback is risk management, not ceremony. The gap caught in round one — no visible progress, no named target — would have been expensive at engineering handoff and far worse across 26 markets. Testing early isn’t about being thorough. It’s about being wrong cheaply.
- Seniority is an input, not a decision. The most useful thing I did with executive feedback was sort it — two points were already evidenced by partner research, one belonged in the design system instead. Taking all three at face value would have been faster and would have produced a worse product. The way to disagree with an SVP is not to argue harder; it’s to arrive with the test already designed.
- Saying no to 26 markets was the highest-leverage decision I made. Every market wanted its own fix, and giving them that would have been the path of least resistance. One tokenised system meant no market got exactly what it asked for — and all of them shipped. Constraints defended early are what make scale possible later.
If I ran this again, progress indicators would have been in round one. It was the one thing partners asked for that I hadn’t already built, and it was predictable from the very first interview — “I feel stuck” is a statement about position, not about information. Two rounds was fast. It could have been one.
Credit where it’s due
One name is on this case study; the outcome isn’t the work of one person. The Data Analyst who defined the KPIs and built the Eppo experiment is the reason the +120% is a number and not a claim. The Product Manager held the metrics and the roadmap that funded the rollout, and the Design Manager kept me honest against the principles when scope pushed back. And ten partners gave candid feedback about a program they had every reason to have given up on.