Most "AI baseball picks" you find online are a chatbot dressed up with a logo. Ours are not. ProbWin's MLB predictions come from our own machine-learning models, built and validated on multiple seasons of real data — and every prediction is published before first pitch so you can check it against the result. Wins and losses alike.
This guide explains, in plain English, how the model works, what goes into it, and — just as important — what we refuse to claim. No black box, no "lock of the day," no guaranteed profits.
Baseball is a data sport — so we treat it like one
A baseball game is a chain of matchups: this starting pitcher against that lineup, in this ballpark, in these conditions, with these bullpens waiting. Almost every piece of that chain leaves a statistical trace. That's exactly why a disciplined model can find an edge where gut feeling can't — and why we build for baseball differently than for, say, soccer.
We don't run one giant model for everything. We run a dedicated model per market, because the over/under, the run line, and the first five innings are three different questions with three different sources of noise.
Three markets, three models
| Market | The question it answers | How our model approaches it |
|---|---|---|
| Game total (over/under) | Will the two teams combine for more or fewer runs than the line? | A run-scoring model that projects each team's runs from the opposing starter, the lineup's profile against that pitcher's handedness, the ballpark, the weather and the bullpens — then simulates the game thousands of times to turn that into a probability. |
| Run line (±1.5) | Will the favorite win by 2+, or will the underdog stay within 1? | A gradient-boosting model built on the differences between the two teams: gap in starter quality, overall team strength, bullpen depth, and an opponent-adjusted run rating (more on that below). |
| First 5 innings (F5) | How many runs before the bullpens take over? | A pitcher-centric model focused on the two starters and the top of each lineup. It's the noisiest market in baseball, so we hold it to the highest bar before it ever becomes an official pick. |
Notice what's missing: hype. Each model outputs a probability, not a slogan. When it says 58%, that number is the whole product — and we hold ourselves to it (see "calibration" below).
What actually goes into the model
We won't hand over the recipe, but here's the honest ingredient list — the factors that move a baseball game and that our models learn from:
- The starting pitchers, evaluated as they were on that day — not with end-of-season stats. This matters more than it sounds (it's the single most common way public "backtests" cheat; more on that below).
- The offense versus handedness — a lineup's real threat depends on whether it's facing a lefty or a righty, especially at the top of the order.
- The ballpark — run environments differ enormously from one stadium to the next.
- The weather — temperature and conditions change how the ball carries.
- The bullpens — recent workload and fatigue, because tired relievers give up late runs.
- An opponent-adjusted run rating — our own twist: a team's runs scored and allowed, corrected for the strength of the opponents it actually faced. Think of it as baseball's version of "opponent-adjusted expected goals." A 5-run night against the best pitching staff in the league is worth more than five runs against the worst.
- Recency — recent games are weighted more heavily than games from months ago, so the model tracks form instead of a stale season average.
The part almost nobody does: proving the model is real
Here's where most "AI picks" fall apart, and where we spend most of our effort.
1. We validate on two full seasons — with no data leakage. Every feature is computed strictly "as of" the moment before the game. That sounds obvious, but the most flattering backtests in this industry quietly use information that wasn't available yet (season-long pitcher stats applied to April games, final box scores fed back in). That's how a broken model shows a fake +30% return. We banned those shortcuts and rebuilt the data so the past can't peek at the future.
2. We only publish an official pick where the model has proven it's calibrated. Calibration is the honest test: when the model says 58%, does it actually win ~58% of the time? We check that, market by market, across both seasons. A model only earns "official" status in the ranges where it repeatedly hit its announced probability — and only when both starting pitchers are confirmed. Everything else stays a trend, clearly labeled, never sold as a pick.
3. We kill our own mirages. A backtest that looks brilliant on one season and can't repeat on the other is thrown out — no matter how good the number looked. That single rule has retired several "edges" that would have cost real money.
You don't have to take our word for any of this. We publish a live "announced vs. actual" table — the model's stated probability next to what really happened — so you can audit us the same way we audit ourselves.
See the live announced-vs-actual journal for MLBWhat we do not claim
Being useful means being honest about the limits:
| We do | We don't |
|---|---|
| Publish a probability for every pick, before the game | Promise you'll win — variance is real, and short-term swings are normal |
| Keep a permanent, timestamped record — wins and losses | Hide the losing days or reset the scoreboard |
| Mark a pick "official" only where the model proved calibrated | Sell "locks," "sure things," or a mythical 90% strike rate |
| Tell you plainly what the model does and doesn't cover | Pretend we beat the sharpest closing lines every time — sometimes our number lands ahead of the market, sometimes it doesn't, and we say which |
That last row matters. We're not promising to out-trade the sharpest books on every line — that's a different sport, and anyone guaranteeing it is selling something. Some of our reads land ahead of where the market closes, others don't, and honestly not as often as we'd like — we'd rather tell you that than dress it up. What we do deliver, game after game, is a calibrated, independent read published in the open, so you decide with better information than a hot take.
How to use it
- Start with today's board. Open the MLB picks page: every game gets an analysis, and official picks are separated from trends.
- Read the reasoning, not just the pick. The analysis is the product — the pick is the summary.
- Track it over time. The record is public and timestamped, so you can judge us on a real sample instead of a lucky week.
Rejoignez ProbWin gratuitement
Accédez à nos prédictions IA pour MLB, NHL et NFL. Inscription en 30 secondes.
Baseball rewards patience and process — exactly what a well-built model is for. If you want a transparent, calibrated second opinion on every MLB game, that's what we're here to provide.






