How Do Elo Ratings Work? The Maths Behind Chess Rankings
By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.
When Magnus Carlsen's rating reached 2882 in 2014, the number meant something precise: that against a player rated 2682 he was expected to score about three points out of four, and against a club player rated 1882 to win essentially every game. The system that produces such numbers was designed by a physics professor and chess master, Arpad Elo, for the United States Chess Federation in 1960, and it has since been adopted, often without the name, by football, tennis, video games, dating apps and the rankings of universities and restaurants. It is simple enough to write on a napkin and subtle enough to have kept statisticians busy for sixty years.
The idea
Elo's insight was to treat a player's strength as a number that could be estimated from results and to treat each game as a measurement of it. A rating is a guess; a game is evidence; and after each game the rating is moved a little towards what the evidence suggests. Beat a much stronger player and your rating rises a lot, since the result was surprising; beat a much weaker one and it barely moves, since it was expected. Lose to a stronger player and you lose almost nothing. Over many games the guesses converge on values that predict results well, and the scale is anchored so that a difference of 400 points means the stronger player is expected to win about ten games in eleven.
The formula
Two steps, both short. First, the expected score of player A against player B is computed from the difference in their ratings: E equals one divided by one plus ten to the power of the rating difference over 400, where the difference is B's rating minus A's. That gives 0.5 for equal ratings, about 0.76 for a 200-point advantage and about 0.91 for 400. Second, after the game, A's new rating is the old rating plus K times the actual score minus the expected score, where the actual score is 1 for a win, 0.5 for a draw and 0 for a loss, and K is a constant that sets how fast ratings move. Some consequences:
- •A 1500 player who beats a 1700 player with K equal to 20 gains about 15 points; the 1700 player loses the same 15
- •The same 1500 player beating a 1300 gains about 5
- •Drawing against a player 200 points higher gains about 5, since a draw is better than the expected 0.24
- •Points are conserved between the two players, so the whole pool's average never changes through play, only through players joining and leaving
- •K is set high for new players, whose rating is a poor guess, and low for established ones; FIDE uses 40, 20 and 10
What the numbers mean
The scale has no fixed zero and no top; it is entirely relative, and a rating only means something against other ratings in the same pool. In chess the conventional bands are that 1200 is a beginner who knows the rules, 1600 a strong club player, 2000 an expert, 2200 a national master, 2500 a grandmaster and 2700 the world's top few dozen. The gap between adjacent bands, 200 to 300 points, corresponds to the stronger player winning roughly three games in four. Because the scale is relative, ratings in one era cannot be compared directly with another, and the argument over whether Carlsen is stronger than Kasparov or Fischer cannot be settled by comparing their peak numbers; inflation, deflation and the size of the pool all shift the scale over decades.
Where it falls short
Elo assumed that each player's performance varies around their true strength in a bell curve, which is close to right, and that the true strength is fixed, which is not; a rapidly improving junior is always underrated, and an ageing master overrated. The system also says nothing about how confident it is. A player who has played three games and one who has played three thousand can hold the same rating with very different reliability, and later systems address that: Glicko, used by many online chess servers, attaches a rating deviation to each player that shrinks with play and grows with inactivity, and TrueSkill, developed by Microsoft for Xbox matchmaking, extends the idea to team games with any number of players. All of them keep Elo's core, a number moved by the surprise of each result.
Beyond chess
The method has spread wherever pairs compete. FIFA's world football rankings switched to an Elo-based formula in 2018; the World Football Elo Ratings had been published unofficially since 1997 and predict results better than the old points table did. Backgammon, Go, Scrabble, table tennis and esports use variants, and statistical models of tennis and American football are built on the same expected-score idea. Elo's original paper noted that the method would work for any contest with a winner and a loser, and it has turned out to work for things that are not contests at all, from ranking the attractiveness of faces on a website in 2003, the origin of Facebook, to sorting which of two chatbot answers people prefer. The professor who designed it to settle arguments at chess clubs built a general instrument for turning comparisons into a scale.
The takeaway
An Elo rating is an estimate of a player's strength that moves after every game by an amount proportional to how surprising the result was: the expected score follows from the rating difference on a scale where 400 points means ten wins in eleven, and the rating shifts by a constant times the gap between the actual and expected score. Points are conserved between opponents, the scale is purely relative, later systems add a measure of uncertainty, and the same formula now ranks football teams, video game players and chatbots.