SHOW YOUR WORK · V2
How the odds work
The models, the assumptions, and how much confidence to put in the numbers.
What we’re estimating
The probability that Texas enters the playoffs as a division winner or wild-card team. Every remaining MLB regular-season game is simulated, including interleague games: competitors share opponents, and every game produces one win and one loss.
This remains experimental. V2 adds detail to the player branch; it does not establish calibrated or superior odds. A 35% estimate means about 35 of every 100 simulated seasons qualify under these assumptions.
The dashboard identifies the model version and dates of its actual forecast. If an update fails, a dated earlier forecast—including a V1 fallback when necessary—may remain visible. V2 details below apply only to forecasts labeled V2.
WHAT CHANGED IN V2
Three views of team strength
Each simulated season selects one model: 35% Elo, 35% run differential, or 30% player matchups. That model applies to the whole season, preserving disagreement. V2 retains these provisional weights and the first two branches.
1. Opponent-aware Elo
Teams begin with equal ratings. Completed results update both teams, giving more credit for beating stronger opponents. The update rate is 0.035 in log-odds units. Ratings remain fixed during future simulations so random simulated wins do not feed back into fitted talent.
2. Shrunk run differential
Runs scored and allowed receive a league-average prior worth 45 games. A Pythagorean exponent of 1.83 converts their ratio into team strength. This branch is not fully opponent- or park-adjusted.
3. V2 player and pitching matchups
V2 estimates strikeouts, walks, hit batters, singles, doubles, triples, home runs and contact outs by batter/pitcher handedness. Three seasons of plate appearances receive weights of 1.00, 0.65 and 0.40, newest to oldest. Event counts are park-adjusted before fitting profiles; the scheduled venue is applied once in each matchup.
Estimates shrink toward population averages: hitters use a 600-PA base prior and 1,200-PA split prior; pitchers use 500-BF and 1,500-BF priors. Pitchers retain individual walk, hit-batter, strikeout and home-run estimates, while balls in play use league composition to limit double counting of fielding.
Each player scenario builds legal defensive lineups for both teams, accounts for handedness, and makes limited pinch-hit and pitcher changes at inning boundaries. These are generic decision rules—not a fitted model of Skip Schumaker or another manager. Batting order, base/out state, score leverage and mid-inning tactics are not modeled.
Starters, workload and rosters
Current MLB rosters and probable starters are collected for every update. Announced starters missing from active rosters are verified against MLB identities and included as scenarios, not mislabeled as active players. Unknown starts are sampled from observed starts in the preceding 30 days with four full rest days.
Recent average batters faced divided by 4.25 estimates starter workload, bounded between one and 6.5 innings; relief exposure is the fallback when recent starts are absent. Inferred spot starters supplement undersized rotations and doubleheader coverage. These are assumptions, not announced assignments. Remaining innings use individual relievers; day-to-day bullpen fatigue and injury returns are not modeled.
Park and fielding estimates
Reviewed park and position-defense estimates are frozen through September 14, 2026 and remain visibly dated. They do not refresh nightly. Savant event park indices are regressed with 5,000 PA; hit-batter effects are neutral and ordinary outs are the residual. Unlisted parks in the reviewed dataset use explicit leave-venue-out estimates. Newly uncovered parks stop publication for review.
Fielding uses position-specific runs saved with framing removed, a 1,350-inning prior, and at least nine observed innings for eligibility. Missing fielders use an explicit primary-position neutral prior, counted in the forecast. Rounded data, historical eligibility and residual catcher effects remain limitations.
Shared uncertainty
Home advantage is 0.14 log-odds (about 53.5% for an otherwise even matchup). Each season draws persistent team offsets with standard deviation 0.12. V2 also builds 256 complete player/deployment scenarios spanning the remaining schedule. A player-model season draws one entire scenario, preserving shared talent and deployment effects across games.
This finite scenario bank is rebuilt for each snapshot. It is an approximation: more outcome trials do not eliminate its integration error or uncertain assumptions.
ONCE THE DAY IS IN THE BOOKS
Automatic updates
Lightweight checks look for a completed league-wide baseball date every ten minutes. The private runner checks for queued work on a twice-hourly overnight/morning schedule (03:00–15:59 UTC); GitHub queue delays mean this is not an exact delivery-time promise. Expensive modeling runs once per completed date, not after every final out.
Every run collects results and season statistics explicitly through that date, current rosters and probable starters, and newly completed play-by-play. Recent usage, starter rest/workload, profiles and the 256-scenario bank are rebuilt. Park and fielding estimates remain the dated inputs above.
The result cutoff and collection time differ: rosters/probables describe retrieval time, not an archived historical roster. Today’s live innings and later finals are excluded from a saved prior-day forecast. The dashboard shows these dates and checks for updates every five minutes.
Only validated forecasts replace the public result. Failures retain the last good forecast and notify the owner, with at most three attempts per model/date and a one-hour retry delay. Duplicate jobs cannot publish twice. Postponements require verified replacements; unresolved suspended games or source discrepancies stop publication for review. Automation is limited to 2026.
YOUR WHAT-IFS
Interactive scenarios
V2 runs 40,000 complete seasons for each of three strength profiles: baseline, optimistic and pessimistic. V1 archives used 20,000. Saved counts link every Rangers outcome to its full-league playoff result and chosen model. Multiple picks retain only seasons matching every selection.
These are conditional probabilities, not simulations forcing results. Conditioning can change model shares. We never add single-game percentage changes together. Shared games, persistent strengths and tiebreak dependencies remain intact.
Enough samples to display odds
Estimates and win distributions are suppressed below 200 matching trials. Rare combinations can have no matches even when possible; clear selections to restore supported estimates. The displayed interval covers simulation sampling and unresolved-tie bounds, not total uncertainty.
Strength settings
Optimistic and pessimistic profiles add +0.15 or −0.15 to Texas’s log-odds strength. An otherwise 50/50 matchup becomes approximately 53.7/46.3 or 46.3/53.7 before other adjustments. These are sensitivity settings, not confidence bounds or injury predictions.
Shared URLs preserve model, strength and game selections by stable game ID. Invalid selections are ignored.
Division titles, wild cards & ties
Each league selects division winners first, then wild cards. Winning percentages are compared without display rounding. Implemented two-, three- and four-team tiebreak procedures use the shared completed-and-simulated ledger.
Larger or exhausted cases retain conservative qualification bounds. Every trial stays in the denominator; no random qualifier is invented. Bounds carry through filters; unavailable seed and postseason-opponent probabilities are not displayed.
First-place celebrations use separate logic. Forecast tiebreaks use records relevant to each simulated final standing. No postseason series are simulated.
ASSUMPTIONS & RIGOR
What is checked
- All 30 teams reconcile to 162 regular-season games and MLB standings.
- Raw responses and historical inputs have checksums. Changed and resumed records require explicit reconciliation.
- Completed play-feed scores and PA reconcile with official hits, extra-base hits, walks and strikeouts. Unknown event types stop the run.
- New banks are tied to exact schedules, player profiles, configuration and code.
- Simulations conserve wins and losses; resolved brackets contain 12 distinct qualifiers.
- Public counts match each profile’s engine summaries, including unresolved bounds.
- Tests cover deployment, event parsing, correlated sampling, filters, model isolation, insufficient samples and upload safeguards.
Historical evidence for V2
The populated study reconciled 536,085 plate appearances and 7,114 game scores across 2024–2026. A chronological event test trained through July 31, 2026 and evaluated 45,669 PA across 603 games from August 1–September 14.
Platoon-plus-park estimates modestly improved conditional event log loss over pooled neutral estimates: 1.469706 versus 1.472781 (lower is better). The paired game-level change was −0.003057, with an approximate interval of −0.003771 to −0.002344. Park-only improvement was inconclusive. Prior-year park inputs avoided including test outcomes in park estimates.
The test supplied actual participants and handedness. It did not validate lineup decisions, fielding, injuries, runs, game-win probabilities or playoff calibration. Shared players/series weaken game independence; later-retrieved records may include scoring corrections.
What remains uncertain
No end-to-end calibration establishes V2 ensemble superiority or calibrated playoff probabilities. The matched scalar comparison did not demonstrate a meaningful headline-odds improvement. Manager policies, event-to-run conversion, rosters, fixed park/fielding estimates, priors and ensemble weights remain assumptions.
The 95% Wilson interval excludes model error, data error, injury uncertainty and finite-bank integration error. In the original study, four disjoint 64-scenario banks produced about a 0.59-percentage-point ensemble-weighted spread. That is a diagnostic for that study, not a reusable confidence interval.
Higher detail is not proof of higher accuracy. Full validation requires archived inputs, multiple seasons/horizons and untouched game-level holdouts. Download current counts, dates and fingerprints.
Sources & reproducibility
- MLB Stats API: schedules, standings, rosters, statistics, probables and completed play-by-play.
- Savant park factors and fielding run value: reviewed historical inputs and the conversions above.
- MLB postseason format and tiebreakers: qualification structure.
- Marcel principles: motivation for regression, not an exact implementation or learned aging curve.
Reproduction requires identical inputs, configuration, code, bank, seed and Python runtime. Public counts contain no subscriber data. Private daily audit artifacts have limited retention; this is not a permanent calibration archive.
Return to dashboard →