CRO Playbook
The ICE Scoring Framework: How to Prioritize A/B Tests Without the Drama
Most CRO programs don't fail because of bad ideas. They fail because every idea fights every other idea for runway. The ICE framework — Impact, Confidence, Ease — is the simplest, fastest way to pick the next test and move on. Here's how to use it without turning it into a spreadsheet you'll abandon by week three.
What you'll learn
- What ICE stands for and where it came from
- How to score each dimension without lying to yourself
- A 10-row scoring template you can copy today
- Worked example: prioritizing 6 real CRO tests
- How ICE compares to PIE and the LIFT Model
- When ICE breaks — and what to do about it
What is the ICE framework?
ICE is a lightweight prioritization framework popularized by Sean Ellis in the growth-hacking community. You score every experiment idea on three dimensions from 1–10, then either average or sum the scores to get a ranking.
If this test wins, how much will it move the metric you care about (revenue, signups, activation)?
How sure are you the test will win? Backed by data, qualitative research, or a previous result — not gut feel.
How fast and cheap is it to ship — design, dev, QA, and statistical power? Higher score = easier.
ICE Score = (Impact + Confidence + Ease) / 3. Sort descending. Run the top one. Re-score the rest after the result lands.
How to score each dimension honestly
The fastest way to wreck ICE is to score every idea an 8. Use anchored rubrics so everyone on your team scores the same idea the same way.
Impact (1–10)
- 1–3
- Affects a low-traffic page or a metric two layers from revenue.
- 4–6
- Affects a mid-funnel step — visible lift if it wins, but not a step-change.
- 7–9
- Touches a high-traffic primary funnel page (hero, pricing, checkout).
- 10
- Likely to materially change quarterly revenue if it wins.
Confidence (1–10)
- 1–3
- Pure gut feel. No data, no research, no precedent.
- 4–6
- Backed by qualitative data (session replays, surveys) or a heuristic principle.
- 7–9
- Backed by quantitative data (analytics funnel drop-offs, user-test patterns).
- 10
- Repeating a variant that already won on a comparable surface.
Ease (1–10)
- 1–3
- Multi-sprint build, new infra, or needs >6 weeks to reach significance.
- 4–6
- 1–2 weeks of design + dev; reaches significance in a month.
- 7–9
- Few days to build; reaches significance in 2–3 weeks.
- 10
- Copy or styling change shippable today; reaches significance in under 2 weeks.
Worked example: 6 real CRO tests
Here's how we'd score six common test ideas for a mid-stage DTC brand doing ~$2M/mo in online revenue.
| Test idea | I | C | E | ICE |
|---|---|---|---|---|
| Add social proof above PDP add-to-cart | 9 | 8 | 9 | 8.7 |
| Sticky checkout summary on mobile | 9 | 7 | 7 | 7.7 |
| Rewrite hero headline based on top reviews | 7 | 7 | 9 | 7.7 |
| New shipping calculator on PDP | 8 | 5 | 4 | 5.7 |
| Personalized homepage by referrer | 8 | 4 | 3 | 5 |
| Redesign the entire navigation | 7 | 3 | 2 | 4 |
The nav redesign and the personalization play are real opportunities, but they lose on Ease and Confidence. Run the top three first; revisit the long-tail after you've banked wins and learned something about your traffic.
ICE vs PIE vs the LIFT Model
ICE isn't the only prioritization framework in the CRO world. WiderFunnel's PIE framework and the LIFT Model (popularized by conversion.com and others) are the most common alternatives. Each has a place.
Three quick axes. Best for fast-moving teams shipping weekly. Trade-off: easy to game; needs anchored rubrics.
Potential, Importance, Ease. Same shape as ICE but Importance forces you to weigh page-level traffic. Better when test ideas span many pages.
Not a prioritization score — a diagnostic. Evaluates a page across Value Proposition, Clarity, Relevance, Distraction, Urgency, Anxiety. Use it to generate hypotheses; use ICE/PIE to rank them.
In practice we use LIFT to find ideas during a teardown, then drop them into an ICE backlog to rank.
When ICE breaks
- Everything scores 7–8. Re-anchor your rubric or split Impact into "revenue impact" and "learning impact."
- Low-traffic pages always lose. Add a traffic threshold — if a surface can't reach significance in 4 weeks, it's ineligible regardless of score.
- Engineering capacity isn't constant. Weight Ease by current sprint load, or carve out 20% of test slots for higher-effort bets.
- Scoring becomes political. Score independently first, then reconcile. Average the team scores, don't debate them.
Copy-paste ICE backlog template
Drop this into a Google Sheet, Notion table, or your tool of choice. Score independently, paste scores in, sort by ICE descending. That's it.
Hypothesis | Surface | Primary metric | Impact (1-10) | Confidence (1-10) | Ease (1-10) | ICE | Owner | Status
Want us to run the backlog for you?
Rocket Bunny embeds an AI engine + a human strategist to keep an ICE-scored test backlog live, prioritized, and shipping every week. We bring the hypotheses; you ship the revenue.
Book a free CRO auditLast updated June 9, 2026 · Written by the Rocket Bunny CRO team.