Methodology · version 2.1 · the anti-black-box
How greatness
gets measured
Promotions publish rankings without publishing rules. Judges publish scorecards without publishing reasons. This page is the opposite: every parameter, every weight, every exclusion, and every known limitation of the system that produces our numbers — including the parts that make our own rankings less flattering to famous names.
§ 01 One replay of history
We take the entire recorded history of professional MMA — 351,930 bouts across 155,352 athletes, 1980 to Sep 13, 2026 — and replay it chronologically, bout by bout, through the GOAT Rating — a probabilistic rating engine that tracks not just how good a fighter is, but how certain we can be about it. Beat better opponents, your rating rises more. Lose to weaker ones, it falls further. Every athlete starts equal, at 1600. Nobody is seeded by reputation.
A fighter's peak rating is the highest value they ever reached — the basis for the all-time leaderboards. Their current rating is where the replay left them. Both are shown everywhere, unrounded by narrative.
Current rankings order athletes by current rating among those active — a professional bout within 18 months of the data date. Inactive athletes keep their rating (uncertainty grows, skill is never silently subtracted) but leave the current list until they return. The window is an editorial parameter, published here precisely so it can be argued with.
§ 02 Parameters nothing up the sleeve
| Parameter | Value | What it does |
|---|---|---|
| Starting rating | 1600 | every athlete's first-bout rating |
| New-fighter uncertainty | high | early bouts move a rating fast, in either direction |
| Inactivity | uncertainty grows | layoffs make us less sure — they never silently subtract skill |
| Updates | per bout | ratings move after every single fight |
§ 03 The points table how you win matters
A finish is stronger evidence of superiority than a split decision — so it moves ratings more. As of version 2.1 these weights are not set by hand: they are fit by a machine-learning approach — optimized for predictive accuracy over 230,000 historical fights, trained on 2010–2019 and validated on held-out 2020–2024 results:
| Method of victory | Weight | Reading |
|---|---|---|
| Finish (KO, TKO, submission) | 0.5 | the strongest signal there is |
| Unanimous decision | 0.39 | clear, but judges were required |
| Other / unclassified | 0.22 | technical decisions, unusual endings |
| Split decision | 0.12 | one judge disagreed — weak evidence |
| Disqualification / no-contest | excluded | not evidence of fighting ability at all |
In the published runs that means 897 DQs and 2,124 no-contests were left out entirely. An eye poke should not mint a legend. Method weighting is not a stylistic choice: in backtesting it independently improved prediction over plain win/loss scoring — and the learned weights delivered the sharpest verdict in the system: the data prices a split decision at barely a quarter of a finish.
§ 04 Pools from the fight graph not from a spreadsheet
Men's and women's ratings are computed separately — but nobody sat down and tagged 155,352 fighters by hand. Pool membership is derived from the data: seeded label propagation over the opponent graph, which naturally splits by gender at 99.99% purity. The handful of cross-edges the algorithm finds are flagged for data-quality review, never silently resolved.
§ 05 Layoffs and comebacks uncertainty, not punishment
We apply no ad-hoc inactivity decay. When a fighter is away, the engine's uncertainty about them grows on its own — the system becomes less sure of them, not convinced they got worse. A returning legend keeps their number but has to re-prove it; the first bouts back move their rating faster in either direction. That is exactly how uncertainty should behave.
§ 06 What we deliberately don't correct
Ratings at the top of the sport inflate an estimated +8–14 points per year as the talent pool deepens. Version 2.1 applies no era normalization: we publish raw numbers and this caveat instead of a hidden adjustment we can't yet defend. All-time comparisons across decades carry that drift. When a normalization is adopted, it will arrive as a new methodology version with its own superseding decision record — not as a quiet edit.
A second known trade-off, also accepted with eyes open: short, meteoric careers are discounted. A fighter with a dozen bouts simply hasn't generated enough evidence for a top-tier peak, no matter how spectacular those bouts looked — the most famous example lands around #151 all-time rather than in the top ten. The system prices uncertainty. That is the product, not a bug.
§ 07 Why these rules chosen by evidence, not taste
Version 2 was selected by backtest, not committee: 241,127 scored bouts (2010–2024), judged on pre-bout win-probability accuracy. The current engine beat the v1 system (log-loss 0.6091 vs 0.6133; 65.6% accuracy) and every variant tested. The backtest also settled an argument: making ratings move less (lower K) strictly worsened prediction — v1's defects were in the shape of its updates, not their size. Version 2.1's learned method weights were adopted the same way: fit on 2010–2019, then required to improve on held-out 2020–2024 results before adoption (validation log-loss 0.6081 → 0.6075). When the evidence says our own system is wrong, the system changes.
§ 08 The receipt reproducibility
Rankings here are versioned publications, not a live feed that quietly shifts. Every computation is a run with an immutable id and a full parameter snapshot stored beside its results. The footer of every page on this site names the runs behind it:
| Pool | Run id | Bouts replayed | Athletes rated | Computed |
|---|---|---|---|---|
| Men | 20260915T133944Z-b60d591a | 337,355 | 148,662 | 2026-09-15 |
| Women | 20260915T134044Z-a68e3338 | 14,575 | 6,690 | 2026-09-15 |
Same data, same run id, same parameters → same rankings, forever. A methodology change requires a superseding decision record and produces new runs under a new version — old runs stay in the database for comparison. If we ever move a GOAT, you'll be able to see exactly why.
§ 09 The GOAT designation 365 cumulative days at №1
Athletes marked with the ringed oxide goat hold the site's GOAT designation: 365 or more cumulative calendar days at world №1 in their pool, summed across their whole career — separate reigns add up; the rule is cumulative, not consecutive. Days are computed from the official run's daily-rankings materialization, calendar-day weighted between snapshots, and the designation is permanent once earned: falling out of the rankings later never subtracts earned days.
Like every number here, it is run-versioned — re-derived from the daily rankings on each published recompute, never hand-assigned. Each designated athlete's profile shows the date the threshold was crossed and their total days at №1.
Divisional reign badges
Division boards carry a second ladder: the divisional reign badge, the goat glyph in a tier color, for cumulative days spent at №1 of a division while an active fighter. Days at the top while inactive count nothing — the ladder rewards showing up. Tiers: 1 year (steel), 2 years (bronze), 3 years (silver), 5 years (gold),8 years (oxide red, the GOAT color — the ladder's crown). The ring belongs to the global designation alone: a bare red goat marks an eight-year divisional reign, and only the ringed red goat is the GOAT of the world. Badges are permanent once earned and re-derived from each run's divisional daily rankings; a fighter's division on any day is where their most recent real-weight-class bout put them, so the whole ladder is checkable from the published boards.
Divisional careers
All-time division boards rank athletes by the peak rating reached while a member of that division, with in-division records. Athletes who competed in several divisions appear on each division's all-time board with the peak they earned there — greatness is credited where it happened.
§ 10 Honest limits
- The corpus currently ends Sep 13, 2026; more recent bouts land in the next ingest and a new run.
- Ratings measure results against opposition — they don't yet see in-fight dominance, and judges' decisions are taken as recorded outcomes (weighted by how contested they were).
- Source data is imperfect; contradictions and gaps are quarantined in a quality log for review rather than silently patched.
- Popularity is a real, separate axis. We plan to measure it rigorously — and never blend it into these numbers.