RICE scoring ranks work by multiplying reach, impact and confidence together and then dividing the result by effort.
It is popular for a good reason. It forces four separate conversations that teams otherwise have all at once and badly, and it produces a number you can sort by. It is also four estimates wearing a formula, and the difference between using it well and using it badly is entirely about remembering that.
What is rice scoring?
Score each piece of work on four terms, multiply the first three, divide by the fourth.
Reach is how many people this touches in some period you fix in advance. Impact is how much it matters to each of them. Confidence is how sure you are about the first two. Effort is what it costs to build.
Impact and confidence are usually small fixed scales that a team agrees on once and then reuses, so that everybody is scoring against the same ruler. The exact values matter far less than everybody using the same ones, which is the part teams skip.
The demo board shows the same idea in use, open with no account.
What does each letter actually ask for?
Three predictions and one estimate, and it is worth being blunt about which is which.
Reach is a prediction. Nobody knows how many people will use a thing that does not exist. You are forecasting, usually from a number you already have, such as how many people use the nearest existing feature.
Impact is a prediction and a value judgement at once. It asks how much this matters per person, which is not observable even after you ship it, let alone before.
Confidence is a prediction about your own predictions. It is the most honest term in the framework and the most awkward, because it asks you to write down how much you trust the two numbers you just made up.
Effort is the only term that is an estimate about your own work rather than about the world. It is also, in most teams, the number that turns out to be furthest from the truth.
Where does rice go wrong?
Three places, and the first one is structural.
The confidence term multiplies rather than corrects. If you are half sure about the reach, halving the whole score is a reasonable way to express that. But confidence is a single number covering two separate guesses, and a team that is confident about reach and uncertain about impact has no way to say so. The uncertainty gets averaged and disappears.
Effort is the least stable number in software. Everything in that division depends on an estimate that is famously wrong, and being wrong in the denominator moves the ranking more than being wrong anywhere else. A thing you thought was two weeks and is actually six drops to a third of the score it had.
Reach is where a bias hides. It is the term most easily justified after the fact and the one the person filling in the spreadsheet has the most freedom over. If somebody already wants to build a thing, reach is where that shows up, and it will look like arithmetic by the time anybody else sees it. That is the same failure as one loud group looking like a market, moved into a spreadsheet.
None of this makes it a bad framework. It makes it a framework whose output is exactly as good as its inputs, which is true of all of them and is much easier to forget when there is a number at the end.
How does it compare to counting what already happened?
The arithmetic is similar. Where the numbers come from is not.
Our own scoring model is published in full, so this comparison can be concrete rather than hypothetical. Ours counts votes, comments and views, weights each one, adds them, then divides by an effort estimate.
| Term | RICE asks you to | Ours reads from |
|---|---|---|
| How many people | Estimate the reach | The number who actually voted, commented and looked |
| How much it matters | Estimate the impact | Not modelled, except through what each voter pays |
| How sure you are | Estimate your confidence | Not modelled at all |
| What it costs | Estimate the effort | Estimate the effort |
So RICE is four judgements with a formula around them, and ours is three counts and one judgement. That is a real difference and it cuts both ways.
RICE can score work nobody asked for. Ours cannot. If the thing is a strategic bet, an infrastructure project or something your customers could not have imagined, the counts are all zero and a board has nothing useful to say about it. That is a genuine limitation of the way we do it and no amount of weighting fixes it.
Counting removes the term teams get most wrong. Nobody has to forecast how many people want it, because they already told you. And the impact term, the one that resists measurement completely, gets replaced by something that is at least observable, which is how much money the people asking are paying you.
Rice vs moscow, and when does each one win?
Rice wins when you need an order and moscow wins when you need a cut line.
Moscow sorts work into must, should, could and will not, which is four buckets rather than a ranking.
Inside must there is no order at all, so a team with thirty musts still does not know what to build on Monday. The method and a worked example over the same twenty request board are both written out.
Rice produces a number, so it orders everything, and the price of that is the three predictions above. The practical split is that moscow suits a release with a fixed date, because the question there is what drops, and rice suits a backlog with no date, because the question there is what comes next.
Should you use rice or something simpler?
Ask whether the work has been requested. That answers it faster than any comparison of formulas.
For anything already on a feedback board, a count is better evidence than a forecast, because the forecast is trying to predict the thing the count already measured. Adding reach and impact estimates on top of real votes is putting a guess in front of a fact.
For anything not on a board, RICE or something like it is the only tool available, and it is a decent one. The discipline of separating the four questions is worth having even when every answer is soft.
The wider question of which of these inputs you could stop guessing at is feature prioritization, which is the parent of this page and worth reading if you are still choosing a method.
The trap in both directions is treating the output as a decision. It is a sort order, and the moment a team says the score says instead of naming the person who decided, whichever model produced it has stopped helping.
Our board and its scoring are nine dollars a month with nothing gated above it, and a card is required for the fourteen day trial.
14 days, a card at signup, then $9 or $39 a month.