The Fantasy Football Encyclopedia
Encyclopedia · Concepts · Sample size in a 17-game season

Concepts

Sample size in a 17-game season

The football season is far too short for anything to average out, which is a structural fact about the sport rather than a complaint — and it changes what you are allowed to conclude from what you observe.

as of 2026-08-02 · status seed · format espn-ppr1.0-10tm-std

Up: Fantasy Football · Related: Opportunity share · Touchdown variance · Start-sit — a standing rule, not a research project · Draft capital dominance · Availability and injury risk · Fantasy football glossary

Three numbers that set the whole problem

Seventeen games. That is the NFL regular season. Basketball plays 82. Every intuition you have about how quickly a number becomes trustworthy was formed on a sample roughly five times larger, arriving four or five games a week rather than one.

Fourteen weeks of fantasy regular season. Your qualifying campaign — finishing top six of ten to reach the bracket — is decided across fourteen head-to-head results. One bad start-or-sit decision is one fourteenth of that campaign. In basketball, one bad week of streaming is a rounding error inside a category race that runs for months. Here it is roughly seven percent of the evidence the standings will ever see.

Three playoff games. The bracket is three single games, each against one specific opponent. Between two evenly matched teams, one game is close to a coin flip, and winning three in a row is not a measurement of quality.

Put those together and the season never gets long enough for the law of large numbers to do its usual work. The averaging that basketball performs for you, quietly, in the background, simply does not happen.

Almost nothing reaches significance in-season

Suppose a player has four excellent weeks in a row. Your instinct — trained on a sport where four good games is a meaningful slice of a long, continuous record — says something has changed.

Usually it has not. Four games of football is four samples of a weekly score that is already spiky by construction, because a large share of it is a rare six-point event that either happened or did not (Touchdown variance). If you generated a season of pure noise and went looking for the best four-week window in it, you would find one, and it would look convincing.

The same applies to cold stretches, to a "breakout," and to a back who is "in a slump." At the sample sizes available inside one season, these are almost always stories fitted to noise after the fact. That is not a counsel of despair but a counsel about which observations to weight.

Efficiency is unstable, volume is sticky

The cleanest way to sort signal from noise is to ask what kind of statistic you are looking at.

Efficiency statistics measure production per opportunity — yards per carry, yards per target, catch rate, the share of touches ending in the end zone. These are ratios with small denominators, heavily influenced by things outside the player: blocking, defensive alignment, a couple of long plays that could easily have been tackled. Year over year they wobble badly. A player's yards per carry this season tells you remarkably little about his yards per carry next season.

Volume statistics measure how many opportunities a player got — carries, targets, snaps (a snap is one play he was on the field for). These are decided by a coaching staff's plan, and plans persist. A player who was given a large share of his team's work is very likely to be given a large share again. See Opportunity share.

That asymmetry produces the single most useful heuristic in fantasy football: role beats performance as a predictor. When deciding what a player will do next, "how much work does he get" outperforms "how well did he play" — and it is why draft-day markets misprice players, since the market sees last year's points, which mix the two, and prices the mixture.

Update slowly

If each week delivers thin evidence, then each week should move your beliefs only a little.

Preseason projections are built from a full prior season plus the offseason's information about roles and depth charts — a depth chart being the team's own ordering of who plays ahead of whom. That is a lot of evidence, and three weeks of a new season is not enough to overturn it. The common in-season mistake is not that people update, it is that they update far too fast.

The practical form of this is a standing rule rather than a weekly investigation, which is what Start-sit — a standing rule, not a research project is for. A rule fixed before the season and applied mechanically beats a fresh judgement made every Sunday from four games of data, because the fresh judgement is mostly reacting to noise and the rule is not.

The logic runs backwards into the draft too. If in-season evidence is thin, pre-season inputs carry more of the load than they would in a long sport — and the most durable of those is where the league itself invested, which is Draft capital dominance.

A worked example, from this build's own data

Data/player_context.csv carries observed games played for 2023, 2024 and 2025, joined from nflverse and retrieved 2026-07-25. Across the 134 players on the board with three full seasons behind them:

Pos Player-seasons Mean games played, of 17 Seasons at all 17
QB 72 13.3 32%
RB 123 14.2 35%
WR 138 14.6 33%
TE 69 14.1 30%

Those are counted, not modelled. The column that matters is the last: about a third of seasons are perfect and the rest are not, so one season of games played is a single draw from something noisy. A rookie who played 17 of 17 has not shown he is durable — he has produced one observation that happens to sit at the top of the range.

That is why Availability and injury risk does not use a raw one-season rate. It shrinks each player's observed rate toward his position's average by the weight of exactly one full season, seventeen games, so one clean year moves the estimate part of the way toward a perfect score and three clean years move it further. The shrunk figure is a model and is labelled as one wherever it appears. What it encodes is this note's argument turned into arithmetic: thin evidence should move a belief only a little, and how far it moves should scale with how much of it there is.

The honest counterweight: not all in-season news is noise

This is where a purely statistical reading goes wrong, and it matters enough to state plainly. There is a difference between sampling from a mechanism and changing the mechanism.

A run of good or bad weekly scores is a sample from a mechanism — the player's role, his team's offence, the way work is distributed — that has not changed. Small samples of an unchanged mechanism are noise, and everything above applies with full force.

But some of what happens during a season changes the mechanism itself:

These are not small samples of the old world. They are the first observations of a new one, and they should move your beliefs immediately and by a lot — the exact opposite of the slow updating recommended above.

The test to apply is simply: did the underlying arrangement change, or did the output vary? Points moved is weak evidence. Opportunity moved is strong evidence. That single question sorts nearly every in-season decision you will face.

Open questions