Methodology
How LineGrade Grades Every Bet
A transparent, honest look at the math — and the guardrails that keep it honest.
The Problem We Solve
Most bettors don't know if a line has value. They place bets based on gut, hype, or last night's box score — without ever quantifying whether the price they're getting actually beats the true probability of the outcome.
LineGrade quantifies that gap on every available line in real time, then filters the output through a strict publication gate so the grades you see are the ones worth acting on — not model noise.
Market Anchor — Sharp-Consensus Devig Across 13 Books
The starting point of every grade is a fair, no-vig probability built from 13 sportsbooks. Instead of the naïve proportional devig — which systematically overprices favorites and underprices longshots — we use the power method: each side's implied probability is raised to a power k, and k is solved so the sides sum to 1. This corrects the favorite-longshot bias baked into raw prices.
Individual book probabilities are then combined in logit space (ln(p/(1−p))) — averaging in raw probability space biases the consensus, especially on longshots. Each book carries a sharpness weight before that average is taken.
- Pinnacle — sharpest and highest-limit book in the world, 4× weight. When available, the fair price is anchored directly to Pinnacle's power-devigged line.
- Bet365 — 2.5× weight.
- DraftKings / FanDuel / BetMGM / Caesars — liquid U.S. market, standard weight.
- Soft U.S. retail books — lower weight; useful for finding mispriced lines when their number strays from the sharp consensus.
Reference-only feeds. Prediction markets (Kalshi, Polymarket) and pick'em apps (Underdog Fantasy, PrizePicks) are shown alongside picks for context but are never graded as picks. Only real sportsbook lines are eligible for a LineGrade grade — pick'em performance is tracked in its own section on the Track Record page.
Deterministic Data Adjustments
The sharp market is the best public predictor of outcomes. Model v4 doesn't try to overrule it — it applies deterministic, sport-specific adjustments to nudge the anchor when the market hasn't priced in publicly available information. Each adjustment is capped, weighted by sample size, and disclosed on the pick page.
- MLB — probable-starter form and xERA, per-hand platoon splits (batter vs LHP/RHP), bullpen fatigue (recent innings load, back-to-back usage), per-park run factors, and live wind/temperature/precipitation for outdoor games. Dome parks are exempt from weather.
- WNBA / NBA — minutes trends and role changes over the last 5–10 games, pace and defensive-rating differentials, rest and travel spots (back-to-back, 3-in-4), and home/away context when the sample supports it.
- Team games (soccer, NHL, football) — recent form vs opponent strength, rest differential, and market-implied ratings for teams without a reliable independent rating.
- Player props — recency-weighted negative binomial (or Poisson where dispersion is well-behaved) fit to the player's actual game log, with a hard 15-game minimum. Below that, the sim abstains and the pick is graded from market signal alone rather than a guess.
Sample-size-aware weighting. The influence of every adjustment is a function of how much evidence supports it. A thin sample gets a small nudge; a rich, consistent sample gets a bigger one. Nothing gets to overwhelm the sharp anchor on its own.
Scratch / validity voiding. If a listed starting pitcher is scratched after odds were pulled, every pitcher-prop row tied to that starter is voided and pulled from the active feed — instead of silently grading against a name that isn't playing.
Per-Segment Calibration From Realized Results
Settled results feed back into the model on a weekly cadence. For each sport × market × edge-bucket × odds-bucket segment, we measure realized hit rate versus the model's predicted probability and derive a calibration factor that pulls future predictions in that segment toward the truth.
Calibration is applied where the sample is large enough to be meaningful and shrunk toward "no adjustment" everywhere else. This is why the model looks like it does what it says it does: overconfident segments get pulled back, underconfident segments get room to run, and segments with no evidence yet aren't allowed to steer.
The Publication Gate
The gate is the single most important layer in Model v4. Every candidate pick has to clear it before it publishes at B+ or above; anything that doesn't is graded C-tier or lower and excluded from the record. C, D, and F are "no edge" calls, not recommendations.
- Segment gating. A pick only publishes B+ or higher if its (sport, market family, edge tier, odds tier) segment has shown positive expectancy on settled data. New segments start with no publication rights and earn them by performing.
- Edge shrinkage by evidence quality. Fewer books quoting the line — or higher disagreement between books that do — shrinks the displayed edge toward the anchor. Rich markets with tight consensus shrink little; thin markets shrink a lot. The size of the cut is disclosed on the pick page.
- Hard 8% edge cap. Any post-shrinkage edge above 8 percentage points is capped at 8% and flagged with a data-quality caveat. Outsized edges are usually stale odds or a mispulled line, not real opportunity.
- Confidence gating. On thin markets with low consensus confidence, the letter grade is capped at B+ regardless of the raw Kelly.
- Coherence guards. The model never publishes both sides of the same line at B+ or above, and only the strongest direction on any totals ladder or spread family survives.
- Grade hysteresis. Grades don't flap between refreshes — an EWMA smoother holds the previous grade unless the new number is a half-step or more away.
The tradeoff is intentional: LineGrade's edges look smaller than what naïve models publish, and fewer picks earn A-tier grades. The ones that survive are the ones worth acting on.
Grading Scale — What A+ Through F Mean
Once the calibrated probability has passed the gate, every bet is scored with the Kelly Criterion — the mathematically optimal bet-sizing formula, which accounts for both edge and win probability simultaneously.
f = (p·b − q) / b where p = calibrated fair probability, b = decimal odds minus 1, q = 1 − p.
This is why a −200 favorite with a real edge can grade A+ while a +500 longshot with a paper-thin edge grades C-. The letter is a signal of quality, not just price discrepancy.
- A+ — Kelly ≥ 10.0%.
- A — Kelly ≥ 7.0%.
- A− — Kelly ≥ 4.8%.
- B+ — Kelly ≥ 2.8%. Also the ceiling for confidence-gated thin markets.
- B — Kelly ≥ 1.8%.
- B− — Kelly ≥ 1.0%.
- C+ and below — "No edge" calls. Displayed for transparency but excluded from the track record.
Accountability — Locked Before Tip, Settled After
- Logged before start. Every graded pick is written to the record before its game begins and stays there permanently — win, lose, or push.
- Settled automatically. Grades are reconciled against official game data on a rolling basis; nothing is selectively curated.
- CLV vs closing lines. Every pick's price is compared to the sharp closing line to measure Closing Line Value — the best available signal of long-term skill vs luck.
- Public track record. Settled results and CLV are visible to everyone on the Track Record page.
- Ongoing recalibration. Settled data feeds back into the calibration factors on a weekly cadence.
What We Don't Do
- We don't guarantee outcomes. Even A+ bets lose. The grade reflects expected value, not certainty. Past grades don't guarantee future results.
- We don't place bets. LineGrade is analytics only — no sportsbook integration, no deposits, no wager handling.
- We don't grade pick'em or prediction-market lines as picks. Those feeds are reference signals only.
- We don't hide losses. Every graded bet stays in the record — winners and losers.
- We don't publish outsized "20% edge" grades. The gate exists specifically to prevent that. LineGrade is informational only — not financial advice.