What We Learned From Simulating Ten Straight Seasons
We ran a decade of league play with no human input to test long-term stability. Here is what broke, what drifted, and what it taught us about balance.
Before certifying the simulation as stable, we ran an endurance test: ten consecutive seasons of a full 30-team league, every affiliate level active, no human decisions at all. Every AI franchise drafted, developed, traded, signed free agents, and managed its own payroll. Then we read the results.
The point of a test like this is not to check whether the game crashes. It is to find slow drift — the imbalances that are invisible in one season and obvious in ten.
What we were looking for
Four failure modes, each of which would quietly ruin a long franchise save:
- Talent inflation or deflation. Does league-average player quality climb or collapse over a decade?
- Competitive collapse. Do a handful of franchises accumulate all the talent and never let go?
- Statistical drift. Do league run scoring, home run rates, and ERA stay in a believable band?
- Roster legality. Does any level of any organization ever end a day with an illegal roster?
What held up
Roster legality never broke. Across ten seasons, 30 organizations, and four levels each, the cascade engine kept every affiliate compliant through injuries, trades, promotions, and releases. Every fill was logged with its source, which meant that when something looked odd we could trace it back to the exact transaction that caused it.
League-wide run scoring stayed inside a believable band across the decade with no manual correction. That was the result we were least confident about going in, since scoring is downstream of player generation, aging, park factors, and the pitch engine all at once.
What drifted
Aging was slightly too gentle in the middle. Players in their early thirties held value a season longer than they should have, which had a knock-on effect: AI teams held veterans longer, traded fewer of them, and the prospect market was thinner than it should have been. Small in one season, structural over ten.
Elite pitching concentrated. By season seven, top-end starting pitching had pooled in high-revenue franchises more than we wanted. The luxury tax was doing its job on total payroll but not on the specific scarcity that matters most in short series.
Draft classes had too little variance. Every class produced a similar number of eventual regulars. Real drafts are lumpy — some years are loaded and some are barren — and that lumpiness is what makes a good pick feel earned.
What we changed
We adjusted the aging curve so decline in the early thirties is a little steeper, which put more veterans back into the trade market and restored prospect liquidity.
We left the pitching concentration mostly alone, with one exception: AI franchises now weigh rotation scarcity more heavily when deciding whether to extend their own starters, which spreads the top end out without adding an artificial cap.
Draft class variance is still on the list. It is a change with wide blast radius, and we would rather ship it with a full ten-season retest behind it than guess.
Why we publish this
A simulation you cannot inspect is a slot machine. If you are going to invest twenty seasons in a franchise, you deserve to know that someone ran a decade of it with the lights on, wrote down what drifted, and fixed the parts that mattered. This is that write-up, and there will be another one after the next engine change.