Building a Baseball Betting Model: From Spreadsheet to Systematic Edge

Spreadsheet on a laptop screen with baseball statistics columns and a colour-coded probability output

Why a Model Beats Your Gut Every Single Time

My first baseball betting model lived in a Google Sheets document with 14 columns and a colour-coded output that told me bet, skip or wait. It was crude. The formulas were basic. The data inputs were limited to what I could copy and paste from free sources. And it outperformed my gut-feel betting by 7% ROI over the course of its first full season. That gap was not because the model was brilliant — it was because my gut was inconsistent, emotional and anchored to narratives instead of evidence.

A model is nothing more than a structured way of converting data into a probability estimate. You feed it inputs (pitcher quality, lineup strength, ballpark factor, schedule spot), it runs those inputs through a set of rules or formulas, and it produces a number: the estimated probability that each team wins. You compare that number to the sportsbook’s implied probability, and if the gap is large enough, you bet. The beauty is that the model makes the same decision framework every single day, immune to the cognitive biases that destroy recreational bettors.

Underdogs win around 43-44% of MLB games outright. That number alone tells you that the sport is more competitive and less predictable than most bettors assume. A model forces you to respect that uncertainty by quantifying it rather than ignoring it. When your model says a team has a 54% chance of winning and the sportsbook implies 52%, the edge is real but thin — and only a model will stop you from treating it as a “sure thing” and over-staking.

Choosing Your Inputs: What Actually Predicts Wins

The biggest mistake beginner modellers make is including too many variables. I know because I made it. My second model had 28 inputs, including humidity, stadium roof status, umpire home-town bias and the catcher’s framing percentage. The model was overfit — it explained past results perfectly and predicted future results poorly, because half the inputs were noise rather than signal.

After years of testing, I have settled on six core inputs that explain the majority of game-level variance in MLB outcomes. Starting pitcher FIP is the single strongest predictor, capturing the quality of each team’s starter independent of defensive luck. Opposing lineup wRC+ against the relevant handedness (lefty or righty starter) measures the offensive threat. Bullpen ERA over the past 14 days captures the current state of the relief corps, which fluctuates more than any other roster component during the season. Home/away split adjusts for the persistent home-field advantage of roughly 53-54% across MLB. Ballpark run factor adjusts the totals projection. And rest/travel status captures the schedule-based fatigue angles that the market systematically underweights.

Six inputs. No more. Each one has a measurable, repeatable relationship with game outcomes, and adding a seventh has never improved my model’s out-of-sample accuracy. The lesson is that simplicity wins in baseball modelling because the sport has so much inherent randomness — about 28% of games are decided by exactly one run — that no amount of additional data can eliminate the noise. A lean model that captures the big effects cleanly will always outperform a bloated model that chases marginal signals.

Weighting and Calibration: Making the Numbers Talk

Gathering the inputs is the easy part. Turning them into a probability is where the modelling skill lives. The simplest approach — and the one I still use as my primary framework — is a log5 method modified with adjustment factors.

Log5 is a formula developed by baseball statistician Bill James that estimates the win probability of a matchup between two teams based on their respective winning percentages. If Team A wins 58% of its games and Team B wins 48%, log5 outputs a probability for Team A winning that specific matchup. The formula accounts for the fact that both teams cannot be average simultaneously — it adjusts for the strength of the opposition implicitly.

My modification layers the six inputs onto the log5 baseline. I convert FIP and wRC+ into implied win percentages using historical regression data (freely available from baseball analytics sites), feed those percentages into the log5 formula, and then apply additive adjustments for bullpen state, home/away, ballpark and travel. Each adjustment is calibrated against three seasons of historical results, with the adjustment size determined by the observed effect in that dataset.

Calibration is the step most amateur modellers skip. A well-calibrated model means that when it says a team has a 60% chance of winning, that team actually wins about 60% of the time across a large sample. If the team wins 65% of the time at that confidence level, your model is underconfident — you are leaving money on the table. If they win 55%, your model is overconfident — you are betting too aggressively. I calibrate my model every 500 games by comparing predicted probabilities to actual outcomes in probability buckets (50-55%, 55-60%, 60-65%, etc.) and adjusting the weights until the calibration curve is as close to the diagonal as possible.

Backtesting Without Fooling Yourself

The temptation to backtest until the model looks profitable is overwhelming — and it is the fastest way to build a model that fails in live betting. Overfitting to historical data is the silent killer of betting models, and baseball’s vast dataset (2,430 regular-season games per year) makes it especially seductive because you can always find a set of weights that produces a stellar backtest.

My backtesting protocol uses strict separation: I build and calibrate the model on data from seasons one through three, then test it blindly on season four. The model never sees season four data during construction. If the blind test shows positive ROI and reasonable calibration, I advance to live paper trading — making hypothetical bets on live games for a full month without risking real money. Only after the paper-trading month confirms the model’s edge do I begin staking.

MLB generated record revenue of $12.1 billion in 2024, and the analytical ecosystem supporting the sport produces cleaner, more comprehensive data than any other league. That data quality is an advantage for modellers, but it also means the market is efficient enough that a poorly constructed model will lose money. The average hold percentage in US sports betting reached 10.15% in 2025 — your model needs to clear that hurdle just to break even. Clearing it by a meaningful margin requires not just good inputs but rigorous testing discipline.

One practical safeguard: track your model’s “closing line value” (CLV). If the line moves in your direction between the time you bet and game time, the market is confirming your model’s assessment. Positive CLV across hundreds of bets is the strongest indicator that your model has a genuine edge, because it means the market’s final price agrees with the price your model identified earlier. Negative CLV means you are consistently on the wrong side of the information flow, and the model needs rework regardless of short-term results.

From Model Output to Bet: The Discipline Layer

A model tells you where the edge is. Staking discipline determines whether you capture it. I use a flat-staking approach — every qualifying bet receives the same unit size, regardless of the model’s confidence level. Some modellers use Kelly Criterion or fractional Kelly to scale stakes by edge size, and I experimented with those methods for two seasons. The theoretical advantage of proportional staking is real, but the practical disadvantage — larger stakes on high-confidence plays that still lose 40% of the time — created emotional strain that compromised my execution.

The minimum edge threshold is equally important. My cutoff is a three-percentage-point gap between my model’s probability and the sportsbook’s implied probability. Below three points, the edge is too thin to survive the bookmaker’s margin and the natural variance of a single game. Above three points, the signal is strong enough to bet with confidence. This threshold eliminates roughly 60% of the daily slate, which means I bet on five to seven games per night rather than twelve to fifteen. Selectivity is the model’s best friend.

Finally, review your model regularly but not reactively. A two-week losing streak does not mean the model is broken — it means variance is doing what variance does in a sport where underdogs win 43-44% of games. I schedule model reviews every 500 bets, not every bad week. At 500-bet intervals, I recalibrate weights, check CLV trends and compare actual ROI to expected ROI. If the three metrics align, I continue. If they diverge, I diagnose and adjust. That cadence keeps the model sharp without chasing noise.

Building a model is the most time-intensive investment a baseball bettor can make, and it is also the most durable. Gut feelings fade. Hot streaks end. But a calibrated, backtested model produces consistent reads on line value that compound over the course of a 162-game season and carry forward from year to year with minimal updates. Start simple, test ruthlessly, and let the data do what data does best: tell the truth when your instincts would rather tell a story.

Do I need programming skills to build a baseball betting model?

Not necessarily. A functional model can be built in a standard spreadsheet application using basic formulas and freely available data. Programming skills (Python or R) become useful for automating data collection and running more complex regression analysis, but the core modelling logic works in any spreadsheet.

How many games of data do I need before trusting my model?

A minimum of three full MLB seasons (approximately 7,300 games) for building and calibrating, plus one additional season for blind testing. Smaller datasets increase the risk of overfitting — finding patterns that appear profitable in historical data but fail to repeat in live betting.

Prepared by the Betting for Baseball editorial staff.