Comparison
Monte Carlo vs historical testing: why retirement calculators disagree
Same plan, two methods. On the same data they usually land within a few points: $40,000 a year from $1 million lasted in 65 of 69 real 30-year retirements (94%) and in 92.9% of 3,000 simulated ones. Enter your numbers to run both.
65 of 69 real retirements, each starting in a year from 1928 to 1996, lasted the full 30 years spending $40,000 a year. So did 92.9% of 3,000 simulated retirements built from real 5-year stretches.
| History | Simulation | |
|---|---|---|
| Lasted 30 years | 65 of 69 (94%) | 92.9% of 3,000 |
| Typical money left (median) | $1.39M | $1.52M |
| Unluckiest 10%: money left | $156k | $113k |
| Ended with at least the starting savings | 43 of 69 | 63.8% |
| Retirements tested | 69 (overlapping) | 3,000 |
What this means
For this plan the two methods agree closely: 94% of real start years lasted, and 92.9% of simulated retirements.
Neither number is a forecast. Both use the same U.S. history; they differ only in how the years are put together. When other calculators disagree by much more than this, the cause is usually their inputs: assumed future returns, fees, inflation, the spending rule and what counts as failure.
The short answer
Historical testing replays the actual sequences of returns since 1928. Monte Carlo builds thousands of new sequences; ours stitches together random 5-year stretches of the same history. On the same plan they usually land close: with $1 million, 60% stocks and $40,000 a year rising with inflation for 30 years, 65 of 69 real start years lasted (94%), and so did 92.9% of 3,000 simulated retirements.
They split in predictable ways. At low spending the simulation is stricter: $37,300 a year lasted in every real start year but in 96.1% of simulations. With all stocks it is stricter too: $35,000 lasted in 68 of 69 real starts (98.6%) but in 93.1% of simulations. But simulation isn’t always the cautious one: at $45,000 it was more generous, 84% of real starts against 86.1% of simulations.
When different calculators disagree by much more than this, the cause is usually their inputs, not the method: assumed future returns, fees, inflation, the spending rule and what counts as failure.
How each method works
Historical testing
Run the plan through every real start year with a complete history, in the order it happened: 69 starts for a 30-year retirement, from 1928 to 1996. Consecutive starts share most of their years, so there are only a few dozen truly different paths, and the late-1960s stretch appears in many of them.
Monte Carlo (ours)
Build 3,000 retirements from random 5-year blocks of the same 1928–2025 data. Each year’s stock, bond and inflation figures stay together, and the blocks keep real crash-and-recovery patterns. A fixed random seed makes the results repeat.
Many other tools draw each year’s return at random from an assumed average and volatility, often a forward-looking estimate rather than history. See how historical testing works and how our simulations work.
Same plan, both methods
| Spending per $1M (30 years) | Real start years that lasted | Simulated retirements that lasted | Typical money left: history / simulation |
|---|---|---|---|
| $35,000, 60% stocks | 69 of 69 (100%) | 97.5% | $1.8M / $1.9M |
| $40,000, 60% stocks | 65 of 69 (94%) | 92.9% | $1.4M / $1.5M |
| $45,000, 60% stocks | 58 of 69 (84%) | 86.1% | $1.2M / $1.2M |
| $50,000, 60% stocks | 50 of 69 (72%) | 76.2% | $817k / $793k |
| $40,000, 100% stocks | 64 of 69 (92.8%) | 89.7% | $4.1M / $3.0M |
Around a 4% withdrawal the methods agree within a point or two, and the typical money left is similar. They also agree on how often a plan ended with at least its starting savings: at $40,000 with 60% stocks, 43 of 69 real starts and 63.8% of simulations did.
Where they agree, and where they don’t
- At low spending, simulation is stricter. $37,300 a year lasted in every real start but in 96.1% of simulations; the simulations only reached 99% at $31,800. They can put bad 5-year stretches back to back in ways history didn’t. For what that means when choosing a target, see what success rate to aim for.
- Above about $41,800 a year (60% stocks), simulation is more generous. History’s success moves in steps, because its failures cluster in neighbouring start years, mostly the late 1960s. $42,000 lasted in 62 of 69 real starts (89.9%) and in 90.1% of simulations.
- With all stocks, the gap is wider. At $40,000, 64 of 69 real starts lasted (92.8%) against 89.7% of simulations, and the typical money left was $4.1 million in history against $3.0 million simulated.
| Highest spending per $1M that met the standard (60% stocks) | History | Simulation |
|---|---|---|
| 95% lasted | $39,000 | $38,100 |
| 90% lasted | $41,700 | $42,100 |
| 80% lasted | $46,400 | $48,100 |
Both methods draw on almost the same years: the average real return the simulations sample is within about 0.1 point of the plain historical average. So the difference comes from how the years are combined, not from different returns. See the worst year to retire for the late-1960s cluster.
Why different calculators disagree much more
On the same data, the method moves results by a few points. Different tools can disagree by far more, because they don’t use the same inputs. When you compare tools, check:
- Assumed future returns and volatility, the biggest lever. A tool that assumes lower returns than history will show lower success for the same plan.
- Fees and inflation assumptions.
- The stock and bond mix, and what “bonds” means.
- Fixed or flexible spending.
- What counts as failure: any short year, running to $0, or falling below a floor.
- Retirement length, Social Security, pensions and taxes.
- How many simulations it runs.
Our engine replays and recombines U.S. history only. It can’t run a simulation with assumed returns, so we don’t show one. If you compare tools, put the same inputs into each, then compare.
How stable are the simulated numbers?
With 3,000 simulations, the results barely move when the random seed changes. At $40,000 a year with 60% stocks, four different seeds all landed between 92.8% and 92.9%; at $50,000 the spread was about one point. The extremes move more: the unluckiest 10% of outcomes shifts by tens of thousands of dollars from seed to seed, so treat that row as approximate. We fix the seed so the same inputs always give the same answer.
How to use both numbers
- Treat them as a range, not a verdict. Neither is a probability of the future.
- If they’re far apart, look at which real start years fell short and why: see the worst year to retire.
- Check what flexibility does. How far spending would have to fall in a bad stretch often matters more than the method: try the guardrails calculator.
- Choose a target with what success rate to aim for.
Assumptions
- Spending rises with inflation every year and never changes otherwise.
- Your savings hold 60% S&P 500 stocks (dividends reinvested) and 40% 10-year Treasuries, rebalanced yearly. Withdrawals happen at the start of each year. No taxes, fees or other income.
- History tests every 30-year stretch since 1928. Simulated retirements stitch together random 5-year blocks of real 1928–2025 history (3,000 runs, fixed seed, so results repeat). Neither is a probability of the future.
New to a term? See the retirement income glossary.
Common questions
Is Monte Carlo or historical testing better for retirement planning?
Neither is better on its own. History shows how a plan held up through real events, but it has only a few dozen truly different 30-year paths. Simulation tests many more combinations, but its results depend on how it builds them. In our test on the same data they landed within a few points around a 4% withdrawal (94% vs 92.9%), so using both gives a range.
Why do retirement calculators give such different results?
Mostly because of their inputs, not their method: assumed future returns and volatility, fees, inflation, the stock and bond mix, whether spending can adjust, what counts as failure, retirement length, Social Security and taxes. Running both methods on the same data, we saw gaps of a few points. A tool that assumes lower future returns than history will show lower success for the same plan.
Is Monte Carlo always more conservative than historical testing?
No. With 60% stocks it was stricter at lower spending ($37,300 lasted in every real start but 96.1% of simulations) and more generous at higher spending ($45,000 lasted in 84% of real starts but 86.1% of simulations). With all stocks it was stricter at almost every level we tested.
How many simulations does a Monte Carlo retirement test need?
Enough that the result stops moving. In our tests, 2,000 and 3,000 simulations gave results within about a point of each other, and at $40,000 a year four different random seeds all landed between 92.8% and 92.9% with 3,000. The extreme results, like the unluckiest 10% of outcomes, move more.
What people ask next
Related tools
How we calculate this
Both methods use the same yearly stock, bond and inflation data and the same spending rule. History replays each real 30-year stretch in order; the simulation builds 3,000 new retirements of the same length from random 5-year blocks of that history, keeping each year’s stock, bond and inflation figures together.
Data: S&P 500 total returns, 10-year Treasury and 3-month Treasury bill returns as compiled by Aswath Damodaran (NYU Stern), and CPI-U inflation from the U.S. Bureau of Labor Statistics, 1928–2025 (January 2026 update). Read the full methodology and limitations.