Story points: what they are and how to estimate
The scale, the Fibonacci sequence, story points vs hours and how to calibrate with a reference story.
What story points are
Story points are a relative unit of effort. They summarize three things at once: the work required, the complexity involved and the uncertainty about what isn't known yet.
Because they are relative, points have no absolute value. A 5-point story is not five hours or five days: it means “about this much” when compared with other stories from the same team.
They are also not a measure of individual productivity. Story points size the item, not the performance of whoever built it — using the scale for people reviews breaks its purpose.
The right question is never “how many hours?” but “how does this compare with what we've already done?”. The answer comes as a number, but its content is a comparison.
Story points vs hours
Hours look precise but age quickly: change the person, the tool or the understanding, and the estimate loses meaning. Points compare stories with each other and survive those changes.
Converting points to hours looks tempting but costs dearly: the team goes back to estimating in hours under another name and loses the relative scale. Once the number becomes a deadline commitment, nobody estimates honestly again.
The healthy use is collective. The average points completed per sprint (velocity) helps the team forecast how much it can pull — never to compare teams with each other or to demand productivity.
A practical test: if a number can be converted into a date without an argument, it probably isn't a story point.
The Fibonacci scale and why it works
The sequence 1, 2, 3, 5, 8, 13 grows fast on purpose: the bigger the item, the greater the uncertainty, and the less sense it makes to tell 8 from 9. Planning poker decks usually add 0, ½ and a break card.
| Value | When it makes sense | Example |
|---|---|---|
| 0 | Nothing to do | Copy tweak, configuration |
| ½ | Almost nothing, but it counts | Rename a label, fix a color |
| 1 | Small and known | Adjust form validation |
| 2 | Small with one detail | Add a masked field and its test |
| 3 | Medium, clear path | Simple screen backed by existing data |
| 5 | Medium with integration | Mobile checkout |
| 8 | Large and uncertain | Report with new aggregation |
| 13 | Too large | A sign it needs splitting |
| Break | Context is missing | Take it to refinement or a spike |
The scale isn't sacred: teams use modified Fibonacci, powers of two or even T-shirt sizes. What matters is being relative, known by everyone and stable over time.
Half points and zero exist so no false precision is forced where there is none: a copy tweak isn't half a story, it's zero; a tweak that needs a deploy might be half a point. The whole scale accommodates those differences without fake steps.
How to calibrate: the reference story
The fastest way to give the scale meaning is to pick a reference story: a small, very well understood item the team has already done. It becomes the “1 point” — and everything else is compared with it.
“1 point” is local. What is 1 for one team can be 3 for another, and that's not a mistake: the scale measures the effort perceived by the people who build. Comparing points across teams tells you nothing.
Revisit the reference from time to time: new people joined, the stack changed, the team switched products? Recalibrate with a quick comparison of a few known stories.
The reference doesn't need to be formal: a screenshot, a ticket number or a single sentence usually says enough.
Common mistakes
- Converting points to hours. The conversion destroys the relative scale and creates false precision.
- Comparing velocity across teams. Each team has its own reference; the comparison becomes meaningless competition.
- Using points in performance reviews. Estimates start inflating the moment they become a grade.
- Re-estimating under pressure. Changing the number to fit the sprint flips the logic: the backlog is what adjusts.
- Estimating everything in points. Urgent bugs, spikes and operational tasks don't always need an estimate — only the stories that will be planned.
A calibration example
With the reference defined (“adjust form validation” = 1), the team compares the backlog with it:
| Backlog item | Points | Why |
|---|---|---|
| Social login | 2 | Known path, one extra flow |
| Password reset by email | 3 | New flow with link expiration |
| Mobile checkout | 5 | Integration and payment cases |
| PDF report | 8 | New aggregation and variable layout |
Notice the numbers don't follow an exact mathematical proportion — and shouldn't. They record the team's read at that moment, comparing items with each other.
With the yardstick in hand, future items come in by comparison: “does this look bigger than checkout? then it's 8”. In minutes, the team estimates the whole backlog.
How to estimate in practice
Story points and planning poker go together: the scale provides the language, the round provides the process. In an online room, the team compares, votes in secret, sees the distribution and talks through the differences.
If your team is starting out, it helps to run the ritual on a few finished stories first: comparing with the past calibrates the reference faster than any spreadsheet.
Two tips for the first session: estimate items the team has already delivered, so the yardstick starts calibrated, and record points completed after each sprint — the number is forecasting input, not a stick to beat people with.
Frequently asked questions
Are story points hours?
No. Story points measure relative effort, complexity and uncertainty. Converting points to hours creates false precision and breaks the team's scale.
Why does the scale use Fibonacci?
Because uncertainty grows with item size. The sequence opens up space between large numbers and avoids debates over differences nobody can perceive.
How much is 1 story point worth?
It depends on the team: 1 point is the size of the reference story chosen by the people who build. That's why points aren't comparable across teams.
Does velocity predict delivery?
It helps the team forecast how much it can pull per sprint, based on its own history. It is not for comparing teams or becoming a productivity target.
Can I use another scale, like T-shirt sizes?
Yes, as long as it is relative and known by everyone. Fibonacci is popular because the progression tracks the growth of uncertainty well.
How do I know when an estimate is off?
Compare with what actually happened: if 5-point items always turn into 8-point work, the reference story needs adjusting. Calibration is continuous, not a one-off event.
Estimate the next story
Create a free room and put the scale to work with your team.
Create room