GRADING ANOMALIES

General discussions about ratings.
Roger de Coverly
Posts: 22609
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 01, 2009 1:51 am

Robert Jurjevic wrote:Why would you stop after 15 games
It's just a simple illustration that a Clarke type system doesn't have to oscillate when confronted when lighthouse keepers.
Robert Jurjevic wrote:Let us assume that Roger 120 and Robert 120 play 30 games Roger scoring 70% and Robert scoring 30%.
Robert Jurjevic wrote:According to GS (current Grading System) Roger's new grade would be 140 and Robert's new grade would be 100, which IMHO makes no sense as Roger's new grade is not 20 grading points greater than Robert's new grade,
70% against a 120 field is 21/30 which gives a new grade of (30*120 + 12*50) / 30 = 140 . 30% against a 120 field is 9/30 which is (30 * 120 - 12*50) /30 = 100. The ECF system uses the previous period's performance as the strength estimate for the next period (and has done so for fifty years). Why is this very simple rule a problem? Please note that the ECF system makes no distinction between playing a range of opponents and playing a single opponent. You may be trying to demonstrate that the ECF system works badly for matches. Don't worry about it, there are hardly any matches least of all 30 game ones. In the following season, the 140 player will need to score 90% against 100 players to maintain his grade.

The biggest perceived problem in the ECF system is the treatment of improving players in that their improvements may not be recognised quickly enough. When you have a player playing 20 or 30 points above their last published grade, you may not want to exacerbate the lag issue by including their old obsolete grade in your calculation of their new grade. Elo and Glicko type systems do this of course. In mitigation they can be used to update the strength estimate more frequently (after every game if data and publication issues permit) In practice in these systems an improving player's rating is converging towards his new strength even if it take a K-factor linked game count to get there.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 01, 2009 11:36 am

Hello Roger,

Thank you for willing to discuss ECF grading issues in this thread, please let me know if I become boring, as this may happen to me my not noticing.
:)

In a nutshell, MHO is:
1) ECF is a sound grading system, but unfortunately it stretches the grades,
2) a small correction is need to advise a system which would be very similar to the current ECF grading system, but which would not stretch the grades, this is AGS3,
3) with a little bit of more effort current ECF grading system can be converted into a system which should be better than the current FIDE Élo, that is ÉGS6 (if the calculation is done using a computer program on a computer practical efforts of implementing either AGS3 or ÉGS6 would be virtually the same)
4) all GS, AGS3 and ÉGS6 have ECF grading scale which in my opinion makes more sense than FIDE scale.
Roger de Coverly wrote:
Robert Jurjevic wrote:Why would you stop after 15 games
It's just a simple illustration that a Clarke type system doesn't have to oscillate when confronted when lighthouse keepers.
Assuming that one applies the same rules to lighthouse keepers as to other players, if the lighthouse keepers play enough games every season for G30 rule not to apply, then I am afraid their grades would oscillate, IMHO I think that you should accept that the lighthouse keepers problem exposes GS weakness, though IMHO it is just one of many examples which show that GS stretches the grades, and that is why I do not wish to treat it as an exceptional case which could be neglected as it is unlikely to happen in practice.

I do not know what the Clarke type system is, but if by Clarke type one meant to say a system with a linear approximation for 'p = f(d)' then even the Clarke type system need not to stretch the grades, as then say AGS3 would be a Clarke type system, the only difference between AGS3 and GS is in 'k' factors in grading formulae in "The formulae..." section at http://www.jurjevic.org.uk/chess/grade/ ... malies.htm (both systems have the same linear approximation for 'p = f(d)').
Roger de Coverly wrote:
Robert Jurjevic wrote:Let us assume that Roger 120 and Robert 120 play 30 games Roger scoring 70% and Robert scoring 30%. ...
According to GS (current Grading System) Roger's new grade would be 140 and Robert's new grade would be 100, which IMHO makes no sense as Roger's new grade is not 20 grading points greater than Robert's new grade,
70% against a 120 field is 21/30 which gives a new grade of (30*120 + 12*50) / 30 = 140 . 30% against a 120 field is 9/30 which is (30 * 120 - 12*50) /30 = 100. The ECF system uses the previous period's performance as the strength estimate for the next period (and has done so for fifty years). Why is this very simple rule a problem? Please note that the ECF system makes no distinction between playing a range of opponents and playing a single opponent. You may be trying to demonstrate that the ECF system works badly for matches. Don't worry about it, there are hardly any matches least of all 30 game ones. In the following season, the 140 player will need to score 90% against 100 players to maintain his grade.
Aren't you worried that in the above example the difference between new GS grades of Roger and Robert is 40 rather than 20 grading points, this is IMHO a huge error (40 grading points difference would correspond to Roger scoring 90% rather than 70%, IMHO a huge error).

IMHO the problem with the GS rule is simply that it applies a grade correction which is twice as big than it should have been, simple as that. When calculating Robert's grade the rule assumes that Roger played at the expected level and when calculating Roger's grade the rule assumes that Robert played at the expected level, which is a contradiction (logical fallacy), as either both Roger and Robert played at the expected level (in which case grade correction would be zero), or only one of the players played at the expected level or neither of the players played at the expected level, you can't assume that Robert and Roger both did and did not play at the expected level, this is simply a logical mistake.

Please note that GS, AGS3 and ÉGS6 all use the previous period's performance as the strength estimate for the next period, what is different are the formulae (or rules) used for grading individual games.

Please note that the same would apply if say Robert rather than Roger played a pool of players with average grade of 120, Robert's new grade would be calculated assuming that both Robert played below his level and that on average pool players played above their level, not that only Robert played below his level and that pool players played exactly at their level. The fact that the players are in the pool should not affect grading of individual games (the pool players played against Robert) as being in the pool should not give them any special status, as for each player in the pool there is an associated pool of the opponents that particular player did play, and it happens that Robert is a member of each of those pools, so then one could argue that Robert is also a member of one or more player pools and therefore has to be treated in a special way. Every player eventually plays a pool of prayers with some average grade, this pool should have no bearing on a decision how to distribute grade corrections when grading individual games.
Roger de Coverly wrote: The biggest perceived problem in the ECF system is the treatment of improving players in that their improvements may not be recognised quickly enough. When you have a player playing 20 or 30 points above their last published grade, you may not want to exacerbate the lag issue by including their old obsolete grade in your calculation of their new grade. Elo and Glicko type systems do this of course. In mitigation they can be used to update the strength estimate more frequently (after every game if data and publication issues permit) In practice in these systems an improving player's rating is converging towards his new strength even if it take a K-factor linked game count to get there.
The error induced due to so called "junior problem" IMHO can be neglected when compared to the error caused by GS grade stretching. Nevertheless, Glicko 2 would address this issue, and if required I can advise a system. But, IMHO AGS3 would be a significant (and good enough) improvement over GS, ÉGS6 would bring fine refinements over AGS3, and if you wish at the top of that to add corrections for "junior problem" even better.

Kind regards,
Last edited by Robert Jurjevic on Tue Dec 01, 2009 1:16 pm, edited 1 time in total.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22609
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 01, 2009 1:13 pm

Robert Jurjevic wrote: In a nutshell, MHO is:
1) ECF is a sound grading system, but unfortunately it stretches the grades,
,
I don't accept the stretching point - please show the practical demonstration using actual data over several years. I am aware of the Grist graphs but to my mind these could just demonstrate an unknown mixture of non-linearity and lag.

Robert Jurjevic wrote: IMHO I think that you should accept that the lighthouse keepers problem exposes GS weakness,
I call it the Clarke system after Sir Richard Clarke, its designer in the 1950s. I don't think it's designed to handle players playing large numbers of games against the same person, rather league conditions which approximate to all play all tournaments. You would expect the grades resulting from all play all tournaments to have the same order as the finishing places in the tournaments.
Robert Jurjevic wrote:
Aren't you worried that in the above example the difference between new GS grades of Roger and Robert is 40 rather than 20 grading points, this is IMHO a huge error (40 grading points difference would correspond to Roger scoring 90% rather than 70%, IMHO a huge error).
Not in the slightest as the Clarke ststem takes no note of whether the 120 opposition is one player or a "stable distant universe". The main point is that the winner slots in between players who score 69% and 71% against equivalent fields or indeed is equal to a player scoring 50% against the 140 field.
Robert Jurjevic wrote: . When calculating Robert's grade the rule assumes that Roger played at the expected level and when calculating Roger's grade the rule assumes that Robert played at the expected level,
The rules of the grading system assume that both players play against a "stable distant universe". This is in practice a completely valid assumption. If there were extended matches taking place in the ECF universe, then it might be worth looking at whether special rules were needed for the players involved.
Robert Jurjevic wrote: Please note that GS, AGS3 and ÉGS6 all use the previous period's performance as the strength estimate for the next period, what is different are the formulae (or rules) used for grading individual games.
If as part of the strength estimate for the next period, you bring in the strength estimate for the previous period, then for a rapidly improving player you are giving them a lower strength estimate than you would if you ignored it. This presumably increases the players in the system with under stated or lagged ratings. The ECF system doesn't really express an opinion about the grading of an individual game. It's % score against the average field that counts. You could perhaps argue that if a player has played 30 games against a field of 140 and scored 50% (no assumption needed about his start grade) then on the 31st game if he wins, his new grade would be (31*140+50) /31 = 141.6. On the 32nd it will be (32*140+100) /32 = 143.1. So his gain per game starts to drop. This is also probably asymmetric against the effect on opponents but it doesn't seem to matter in practice. In your system you would hand the player (after 30 games) a new grade of 130 if he started at 120 and of 120 if he started at 100. Equally if he started at 160, he would remain in the relative elite at 150.

It come back to what a grading system is for. An intention has to be to put players in a correct ranking order. To do this you have to blend historic and recent performance. The ECF system dumps historic performance where the number of games exceeds a minimum. Your proposed system would bring back historic performance as a 50% weighting. It might be more stable than the current system but it would increase lag. Elo systems have a some advantages, not least that they enable a new rating to be meaningfully computed after every game. Their implicit memory of past performances can be a disadvantage since a new player may earn a different rating to an established player for the same performance.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 01, 2009 2:31 pm

Hello Roger,

I am sorry that we cannot reach agreement virtually on any of the points risen. I guess this is because verbal communication might be better suited for exchanging ideas of such nature.
Roger de Coverly wrote:I don't accept the stretching point - please show the practical demonstration using actual data over several years. I am aware of the Grist graphs but to my mind these could just demonstrate an unknown mixture of non-linearity and lag.
Sure, I would not advocate adopting AGS3 or ÉGS6 before testing it in practice (though if I would be in the position of decision making and taking responsibility for eventual consequences I think that I would take the risk and adopt AGS3 right away, as I think the risk would be relatively small in comparison to eventual benefit). My intention is to calculate AGS3 grades for every season, I have already calculated them for 2009 (I can do that without knowing the actual game results), please see http://www.ecforum.org.uk/viewtopic.php ... &start=385.
Roger de Coverly wrote:
Robert Jurjevic wrote:Aren't you worried that in the above example the difference between new GS grades of Roger and Robert is 40 rather than 20 grading points, this is IMHO a huge error (40 grading points difference would correspond to Roger scoring 90% rather than 70%, IMHO a huge error).
Not in the slightest as the Clarke system takes no note of whether the 120 opposition is one player or a "stable distant universe". The main point is that the winner slots in between players who score 69% and 71% against equivalent fields or indeed is equal to a player scoring 50% against the 140 field.
If one calculates Roger's and Robert's new grades based on their performances in the 30 games (no G30 rule applies) the difference between their new grades (calculated based on the 30 games) should be 20 grading points (as Roger scored 70% and Robert 30%). GS clearly makes a huge mistake here assigning 140 to Roger and 100 to Robert. GS would assign 130 to Roger and 110 to Robert, and not only that this would match the performance in the 30 games (Roger scoring 70% and Robert scoring 30%) but would also assure that if in the next season Roger and Robert play 30 games with the same score, Roger scoring 70% and Robert 30%, their grades would remain unchanged, 130 and 110, while according to GS if Roger and Robert played two seasons in a raw, on both occasion Roger scoring 70% and Robert 30%, Roger would be assigned a grade of 120 and Robert of 120, which makes no sense at all, Roger for scoring against Robert 70% in two seasons twice in a row is assigned equal grade as Robert.

Please note that the same would apply if say Robert rather than Roger played a pool of players with average grade of 120, Robert's new grade would be calculated assuming that both Robert played below his level and that on average pool players played above their level, not that only Robert played below his level and that pool players played exactly at their level. The fact that the players are in the pool should not affect grading of individual games (the pool players played against Robert) as being in the pool should not give them any special status, as for each player in the pool there is an associated pool of the opponents that particular player did play, and it happens that Robert is a member of each of those pools, so then one could argue that Robert is also a member of one or more player pools and therefore has to be treated in a special way. Every player eventually plays a pool of prayers with some average grade, this pool should have no bearing on a decision how to distribute grade corrections when grading individual games.

Maybe the term 'stretching' is not appropriate as it may be likely that errors induced in one season would be canceled in the subsequent season and that some sort of oscillations similar to those in the lighthouse problem may occurs (say like assigning back original grades of 120 to both Roger and Robert after Roger scoring 70% two seasons in a row), but form the examples it should be evident that AGS3 is 'healthier' than GS. Because of that it is possible that GS's and EGS3's median, mean and standard deviation stay approximately equal during course of the seasons, if on the other hand the errors do not cancel then I guess this would result in standard deviations of GS and AGS3 distributions drifting apart.
Roger de Coverly wrote:I call it the Clarke system after Sir Richard Clarke, its designer in the 1950s. I don't think it's designed to handle players playing large numbers of games against the same person, rather league conditions which approximate to all play all tournaments. You would expect the grades resulting from all play all tournaments to have the same order as the finishing places in the tournaments.
If Clarke system is a current ECF grading system, with all due respect, I am afraid that I should say that IMHO it does 'stretch' the grades (I know that maybe the 40 point rule has been introduced later, but unfortunately it does not help in regard to grade 'stretching').

Kind regards,
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22609
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 01, 2009 3:56 pm

Robert Jurjevic wrote:Every player eventually plays a pool of prayers with some average grade, this pool should have no bearing on a decision how to distribute grade corrections when grading individual games.
That as I have repeated many times is exactly the point. The ECF system does NOT I repeat NOT make any decision about distributing grade corrections when grading individual games. If you want a system which gives you a new grade after each individual game, then use the the Elo system or some derivative of it. The ECF/Clarke system calculates a performance measurement (which is virtually independent of the initial estimate) and then uses it to set the end period estimate for the following season. Elo type systems also use performance measurement but only in the limited applications of making initial estimates for ungraded players and for establishing eligibility for title norms. Otherwise Elo systems chain one grade to the next so the equivalent performances may give different strength estimates.
Robert Jurjevic wrote:My intention is to calculate AGS3 grades for every season, I have already calculated them for 2009 (I can do that without knowing the actual game results), please see viewtopic.php?f=4&t=22&start=385.
I took a look at your file. The effect of your calculations is to slow down the rate at which improving players have increases in their strength reported in their grade and declining players drop back. For example you've cut back one high flying junior from 199 to 183.5. The general belief is that his strength is equal to players at the 200 level rather than the 180s level. Equally you are protecting players with poor results from the consequences of their play. A player who dropped from 110 to 72 only goes to 91 in your system.

This is the point I made several months ago, that you are just increasing lag by bringing the prior year grade into the current year calculation.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 01, 2009 5:09 pm

Roger de Coverly wrote:
Robert Jurjevic wrote:Every player eventually plays a pool of prayers with some average grade, this pool should have no bearing on a decision how to distribute grade corrections when grading individual games.
That as I have repeated many times is exactly the point. The ECF system does NOT I repeat NOT make any decision about distributing grade corrections when grading individual games.
May I ask then to what words "win", "draw", "loss" and "points-per-game" in the rule 1a below refer to?

Rule 1a: For a win you score your opponent's grade plus 50; for a draw, your opponent's grade; and for a loss, your opponent's grade minus 50. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule G30 applies only when one plays less than 30 games in a season, and has no baring on the rule (or formulae) for grading individual games.

Rule G30: The Grade is calculated by dividing the total number of points scored by the number of games played. If there are at least 30 games in the current period, then the Grade is based on these games alone. If there are not, results are brought forward from the previous period to make the total up to exactly thirty. If there are not 30 games in the two seasons together, results are taken from the season before that. Games are never taken from further back than this; the maximum is two prior grading periods.

When in rule 1a one says "for a draw you score your opponent's grade" this implies that one assumes that both your opponent and you played on your opponent's level, but one may have said "for a draw you score average grade" which would imply that one assumes that both your opponent and you played on average level between you and your opponent, etc.

If say we grade a game Roger 140 drew against Robert 120 then when grading the game for Roger "for a draw you score your opponent's grade" would read "for a draw Roger scores Robert's grade" which implies that both Roger and Robert performed at 120 level, and when grading the game for Robert "for a draw you score your opponent's grade" would read "for a draw Robert scores Roger's grade" which implies that both Robert and Roger performed at 140 level, which contradicts the assumption made when grading Roger's game.

AGS3's rule for grading individual games reads:

Rule 2b: For a win you score average grade plus 25; for a draw, average grade; and for a loss, average grade minus 25. Average grade is half of the sum of your and your opponent's grade. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.
Last edited by Robert Jurjevic on Tue Dec 01, 2009 5:56 pm, edited 2 times in total.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22609
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 01, 2009 5:44 pm

Robert Jurjevic wrote:e]
May I ask then to what words "win", "draw", "loss" and "points-per-game" in the rule 1a below refer to?

Rule 1a: For a win you score your opponent's grade plus 50; for a draw, your opponent's grade; and for a loss, your opponent's grade minus 50. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.
(a) You write down the grade of each opponent, you adjust it by the 40 point rule if necessary, you add up the total adjusted grades. You add 50* the excess of wins over losses. You divide by the number of games. You repeat the process for all players in the grading set.

Or

(b) You write down the grade of each opponent, you adjust it by the 40 point rule if necessary. You add 50 points if a win and subtract 50 points if a loss. You add up the resulting number of points. You divide by the number of games. You repeat the process for all players in the grading set.

I usually use method (a) because it also gives the average strength of the opposition. The two approaches obviously produce the same answer but IMHO it's (a) that is the performance measurement and (b) that is a calculation method producing the same result.

Your method can be described as
(c) You write down the grade of each opponent, you adjust it by the 40 point rule if necessary, you add up the total adjusted grades. You add 50* the excess of wins over losses. You then add your start grade multiplied by the number of games played. You then divide by twice the number of games. You repeat the process for all players in the grading set.

For players with start and end grades of about the same, the differences are minor. For players where the start grade differs from the end grade, method (c) halves the change.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 01, 2009 6:01 pm

Roger,

Sorry, maybe you missed the following (which I added in my previous post while you were answering)...

When in rule 1a one says "for a draw you score your opponent's grade" this implies that one assumes that both your opponent and you played on your opponent's level, but one may have said "for a draw you score average grade" which would imply that one assumes that both your opponent and you played on average level between you and your opponent, etc.

If say we grade a game Roger 140 drew against Robert 120 then when grading the game for Roger "for a draw you score your opponent's grade" would read "for a draw Roger scores Robert's grade" which implies that both Roger and Robert played at 120 level, and when grading the game for Robert "for a draw you score your opponent's grade" would read "for a draw Robert scores Roger's grade" which implies that both Robert and Roger played at 140 level, which contradicts the assumption made when grading Roger's game.

It should be obvious that the following two rules are about how to grade individual games...

Rule 1a: For a win you score your opponent's grade plus 50; for a draw, your opponent's grade; and for a loss, your opponent's grade minus 50. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule 2b: For a win you score average grade plus 25; for a draw, average grade; and for a loss, average grade minus 25. Average grade is half of the sum of your and your opponent's grade. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22609
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 01, 2009 6:26 pm

Robert Jurjevic wrote:Roger,

Sorry, maybe you missed the following (which I added in my previous post while you were answering)...

When in rule 1a one says "for a draw you score your opponent's grade" this implies that one assumes that both your opponent and you played on your opponent's level, but one may have said "for a draw you score average grade" which would imply that one assumes that both your opponent and you played on average level between you and your opponent, etc.
I don't accept this implication. All that's happened is that you drew a game of chess. It does happen even between players of disparate strength. According to Elo, playing strength can be considered as a random variable of unknown distribution so any result is possible, just some are more likely than others, so you have to limit the inference that can be drawn from a single game. So it's quite possible for a 140 player to draw with or even beat a 180 player without implying that the 140 player is better than 140 or the 180 player worse than 180. A draw is a less likely result than between two 140 players of course. That's why in the Elo system, there's a K factor which acts to slow down the acceptance of the hypothesis that the strengths have changed. The 30 game rule is the ECF system's equivalent of the K factor and has the same dampening effect on fluctuations.
Last edited by Roger de Coverly on Tue Dec 01, 2009 7:15 pm, edited 1 time in total.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 01, 2009 7:10 pm

Roger de Coverly wrote:
Robert Jurjevic wrote:When in rule 1a one says "for a draw you score your opponent's grade" this implies that one assumes that both your opponent and you played on your opponent's level, but one may have said "for a draw you score average grade" which would imply that one assumes that both your opponent and you played on average level between you and your opponent, etc.
I don't accept this implication. All that's happened is that you drew a game of chess. It does happen even between players of disparate strength. According to Elo, playing strength can be considered as a random variable of unknown distribution so any result is possible, just some are more likely than others, so you have to limit the inference that can be drawn from a single game. So it's quite possible for a 140 player to draw with or even beat a 180 player without implying that the 140 player is better than 140 or the 180 player worse than 180. A draw is a less likely result than between two 140 players of course. That's why in the Elo system, there's a K factor which acts to slow down the acceptance of the hypothesis that the strengths have changed. The 30 game rule is the ECF system's equivalent of the K factor and has the same dampening effect on fluctuations.
Statement 1: All that's happened is that Roger and Robert drew a game and any result would have been possible though not equally probable.

My I ask if the above statement 1 (with winch I agree) can be used to 'prove' that grading rule 1a is better than say 2b and 3? Thanks.

Rule 1a: For a win you score your opponent's grade plus 50; for a draw, your opponent's grade; and for a loss, your opponent's grade minus 50. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule 2b: For a win you score average grade plus 25; for a draw, average grade; and for a loss, average grade minus 25. Average grade is half of the sum of your and your opponent's grade. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule 3: For a win you score your opponent's grade plus 53; for a draw, your opponent's grade minus 12; and for a loss, your opponent's grade minus 27. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22609
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 01, 2009 7:39 pm

Robert Jurjevic wrote:
My I ask if the above statement 1 (with winch I agree) can be used to 'prove' that grading rule 1a is better than say 2b and 3? Thanks.
Fifty points for a win gives you the net result that players 25 points apart have to score 75% to maintain their grade. 25 points for a win introduces lag so that a player who improves from 150 to 175 is ranked below one who has always been 175. Any non symmetric scoring scheme will cause the whole structure to move. We used to do this once. Juniors had published grades of their previous season performance, but when you played them in effect you got 60 points for a win, 10 for draw and lost only 40.

The real empirical proof is that the grading system has been running on a plus 50/ minus 50 basis for over 50 years without exploding into chaos or engendering a complete lack of confidence. One future GM rocketed from about 175 to 220 in his teens. He would have been well undergraded at 195 to 200. In fact as 175 wasn't his first grade, he wouldn't even have started at 175, more like 160 with the lag from his younger days. Look at the system as a whole, you are trying to rank players in strength order. If you presume that at least 30 games are needed to establish a reasonable estimate of a player's strength and you are using performance as the strength estimate, it cannot be correct to rank a player playing 30 games at 175 below one who also plays 30 games at 175 just because the previous year's strength estimate was 150.

I think I'm now repeating myself so I'm not going to add any more
Last edited by Roger de Coverly on Tue Dec 01, 2009 8:03 pm, edited 1 time in total.

Peter Rhodes
Posts: 246
Joined: Thu Aug 20, 2009 10:53 pm

Re: GRADING ANOMALIES

Post by Peter Rhodes » Tue Dec 01, 2009 7:59 pm

The real empirical proof is that the grading system has been running on a plus 50/ minus 50 basis for over 50 years without exploding into chaos or engendering a complete lack of confidence.

A child could develop such a system that would meet these requirements. It seems to me that as Robert drills deeper and deeper into the detail, his detractors pull further and further from it.

The real empirical evidence tells us that we run a unique system which no-one else would touch with a barge-pole.
Chess Amateur.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Wed Dec 02, 2009 1:36 pm

Roger,
Roger de Coverly wrote:
Robert Jurjevic wrote:When in rule 1a one says "for a draw you score your opponent's grade" this implies that one assumes that both your opponent and you played on your opponent's level, but one may have said "for a draw you score average grade" which would imply that one assumes that both your opponent and you played on average level between you and your opponent, etc.
I don't accept this implication. All that's happened is that you drew a game of chess. It does happen even between players of disparate strength. According to Elo, playing strength can be considered as a random variable of unknown distribution so any result is possible, just some are more likely than others, so you have to limit the inference that can be drawn from a single game. So it's quite possible for a 140 player to draw with or even beat a 180 player without implying that the 140 player is better than 140 or the 180 player worse than 180. A draw is a less likely result than between two 140 players of course. That's why in the Elo system, there's a K factor which acts to slow down the acceptance of the hypothesis that the strengths have changed. The 30 game rule is the ECF system's equivalent of the K factor and has the same dampening effect on fluctuations.
I think that statement 1 below is not in contradiction with statement 2 below as one does correct players' grades based on their game results regardless of the fact that any game result would have been possible but not equally probable.

In the ECF official grading statement below it is said that grading "points are allocated in respect of each game." and that one's final "grade is calculated by dividing the total number of points scored by the number of games played" which is an average of grading points scored in each game.

If there are different rules for scoring grading points in each game there should exist an explanation about what makes the rules different. If a rule says "for a draw Roger scores Robert's grade" isn't it obvious that this should imply that the rule assumes that Roger played at Robert's level and is therefore assigning to Roger Robert's grade? If a rule says "for a draw Roger scores average grade" isn't it obvious that this should imply that the rule assumes that Roger played at the level which is an average between his and his opponent's and is therefore assigning to Roger the average grade?

If one can accept that different rules make different assumptions about the players' levels of play then GS's rule 1a can be refuted on the following example:
If say we grade a game Roger 140 drew against Robert 120 then when grading the game for Roger "for a draw you score your opponent's grade" would read "for a draw Roger scores Robert's grade" which implies that Roger played at 120 level, and when grading the game for Robert "for a draw you score your opponent's grade" would read "for a draw Robert scores Roger's grade" which implies that Robert played at 140 level, so it is assumed that Roger played at 120 level and Robert at 140 level which contradicts the fact that it should have been assumed that they have played at the same level as they drew.

AGS3's rule 2b cannot be refuted on the above example:
If say we grade a game Roger 140 drew against Robert 120 then when grading the game for Roger "for a draw you score average grade" would read "for a draw Roger scores average grade" which implies that Roger played at (140+120)/2=130 level, and when grading the game for Robert "for a draw you score average grade" would read "for a draw Robert scores average grade" which implies that Robert played at (140+120)/2=130 level, so it is assumed that both Roger and Robert played at 130 level which is in accord with the fact that it should have been assumed that they have played at the same level as they drew.

ECF official grading statement: Points are allocated in respect of each game. For a win you score the opponent's Grade plus 50, for a draw the opponent's Grade, and for a loss the opponent's Grade minus 50. There is a proviso that if your opponent's Grade differs from yours by more than 40 points it is assumed to be exactly 40 above (or below) yours. This is to prevent a player increasing his Grade by losing to a much stronger player, or decreasing his Grade by beating a much weaker player. If an opponent (or the player himself) is ungraded, a Grade is estimated, using all available information. * The Grade is calculated by dividing the total number of points scored by the number of games played. If there are at least 30 games in the current period, then the Grade is based on these games alone. If there are not, results are brought forward from the previous period to make the total up to exactly thirty. If there are not 30 games in the two seasons together, results are taken from the season before that. Games are never taken from further back than this; the maximum is two prior grading periods. * Results are brought forward in two different ways, depending whether the Grade is Rapid or Standard. With Rapidplay, any games brought forward from a previous period will be the most recent games in that period. This is possible because the dates of Rapid games are (almost) always known. With Standardplay, unfortunately, this is not the case. So, instead, the required number of (notional) games is brought forward at the average score for the period.

Statement 1: When in rule 1a one says "for a draw you score your opponent's grade" this implies that one assumes that you played on your opponent's level, and that when in rule 2b one says "for a draw you score average grade" this implies that one assumes that you played on average level between you and your opponent, etc.

Statement 2: In a chess game between two players any result is possible though not equally probable.

Rule 1a: For a win you score your opponent's grade plus 50; for a draw, your opponent's grade; and for a loss, your opponent's grade minus 50. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule 2b: For a win you score average grade plus 25; for a draw, average grade; and for a loss, average grade minus 25. Average grade is half of the sum of your and your opponent's grade. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Both rules 1a and 2b can be expressed with the following formulae:

Code: Select all

a2 = a + ka*(q - p);
b2 = b + kb*((100 - q) - (100 - p));
where 'a' is your grade, 'b' grade of your opponent, 'p' your expected performance (expected performance of your opponent is then '100 - p'), 'q' your actual performance (actual performance of your opponent is then '100 - q'), 'a2' your new grade (your grading points allocated for the game) and 'b2' your opponent's new grade (your opponent's grading points allocated for the game).
(if the players played only one game in the season 'q' is either 100, 0 or 50, if they played more than one game it can be a number between 0 and 100 inclusively)

Grading points allocated for you for the game are 'a2 = a + ka*(q - p)', i.e. one applies a correction 'ka*(q - p)' to your grade 'a'.
Grading points allocated for your opponent for the game are 'b2 = b + kb*((100 - q) - (100 - p))', i.e. one applies a correction 'kb*((100 - q) - (100 - p))' to your opponent's grade 'b'.

The only difference between the two rules is in 'ka' and 'kb' factors:
rule 1a: 'ka = kb = 1'
rule 2b: 'ka = kb = 1/2'

I found that a necessary and sufficient condition for a grading system not to stretch (nor shrink) the grades is that:
'ka + kb = 1'

Sum of the 'k' factors:
rule 1a: 'ka + kb = 2' (stretches the grades)
rule 2b: 'ka + kb = 1' (does not stretch nor shrink the grades)

Correction applied to your grade:
rule 1a: (q - p)
rule 2b: (q - p)/2

Correction applied to your opponent's grade:
rule 1a: ((100 - q) - (100 - p))
rule 2b: ((100 - q) - (100 - p)) /2

Mathematical requirement for a grading system (using the above formulae) not to stretch nor shrink the grades is that '(a2 - a) + (b - b2)' is equal to 'q - p' for any 'a', 'b', 'p' and 'q', i.e., one requires that the sum of grade corrections for both players '(a2 - a) + (b - b2)' matches the difference between actual and expected performance 'q - p'. It can be proven that iff '(a2 - a) + (b - b2)' is equal to 'q - p' for any 'a', 'b', 'p' and 'q' then 'ka + kb = 1'.

Note that the above formulae also hold for ÉGS5 and ÉGS6, ÉGS5 differs from AGS3 only in definition of 'p = f(d)' (i.e. for a given grade difference 'd' ÉGS5's expected performance is different form that of AGS3, the difference is noticeable in practice approximately for '|d| > 30'), ÉGS6 differs from ÉGS5 only in 'ka' and 'kb' factors (ÉGS6 'ka' and 'kb' factors may range from 0 to 1 but their sum is always 1, the more trusted one's grade is the closer is the 'k' factor to 1, the grades trust estimate in ÉGS6 is based on frequency of play).

Kind Regards,
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22609
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Wed Dec 02, 2009 2:55 pm

Robert Jurjevic wrote:If there are different rules for scoring grading points in each game there should exist an explanation about what makes the rules different. If a rule says "for a draw Roger scores Robert's grade" isn't it obvious that this should imply that the rule assumes that Roger played at Robert's level and is therefore assigning to Roger Robert's grade? If a rule says "for a draw Roger scores average grade" isn't it obvious that this should imply that the rule assumes that Roger played at the level which is an average between his and his opponent's and is therefore assigning to Roger the average grade?
You are reading far too much into a simple calculation method. In the ECF system, you use 50 points rather than 25 because if you use 25 and thus average in the previous grade, then players who are inconvenient enough to play above their grade will not be completely revalued to their new higher level. If you really want ex-150 players who have improved to 175 to be given grades of 162.5 feel free to advocate your system. I don't think the rest of us want it because annual lists cause enough trouble with lag as it is. One of the points of a grading system is to rank players by strength. You are not doing this if you "remember" the previous grade when comparing otherwise identical performances. Think about it , you've got 175 standard players with 162.5 grades - what's that going to do to expected scores?

Thinking about it further, I think even the Elo system introduces paradoxes when confronted with a match and one player playing 200 points above his previous standard. Let's suppose a 2 player, 2 game mini match between players of equal rating (say 2000) and that one player always win 1.5 to 0.5. So with a K factor of 15, each match gains the winner 0.5 * 15 = 7.5 Elo points and the loser loses the same. So each match causes their rating to move 15 points apart, so by 30 games they've overshot their "true" difference of 200 by 25 points to 225. Even worse, one of the players is still under rated at 2112 and the loser probably under rated at 1887.

I think you have to consider a grading system as measuring performance against a pool of players, because if you don't you get tied up in paradoxes. Could I suggest that you treat the ECF's description of the grading system as an explanation of a calculation method rather than a statement of the underlying theory? Again if you want a grading system which can compute game by game ratings then you use Elo or one of its derivatives.

Your proposed system can be summed up very simply as stating that even for the most active players, their new grade is an average of their performance for a season and their previous season's grade. As a consequence it might plausibly be more stable than the current ECF system but is almost certainly more out of date and demoralising to rapidly improving players.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Wed Dec 02, 2009 5:38 pm

Hello Roger,

Rule 1a: For a win you score your opponent's grade plus 50; for a draw, your opponent's grade; and for a loss, your opponent's grade minus 50. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule 1c: For a win you score your grade plus 50 minus grade difference; for a draw, your grade minus grade difference; and for a loss, your grade minus 50 minus grade difference. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule 2b: For a win you score average grade plus 25; for a draw, average grade; and for a loss, average grade minus 25. Average grade is half of the sum of your and your opponent's grade. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule 2c: For a win you score your grade plus 25 minus half grade difference; for a draw, your grade minus half grade difference; and for a loss, your grade minus 25 minus half grade difference. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.
Roger de Coverly wrote:In the ECF system, you use 50 points rather than 25 because if you use 25 and thus average in the previous grade, then players who are inconvenient enough to play above their grade will not be completely revalued to their new higher level.

Please note that if expected performance 'p' is 'p = 50 + d' (which is the case for GS and AGS3), where 'd' is grade difference, then rule 1a is equivalent to rule 1c, and rule 2b to rule 2c.

GS: rule 1a or equivalent rule 1c
AGS3: rule 2b or equivalent rule 2c

Rules 1c and 2c are convenient for comparison as they both express your new grade as a function of your grade and some correction. It can be easily seen that the only difference between rules 1c and 2c is in the amount of correction they apply, and that correction applied by rule 1c is twice of that applied by rule 2c.
Roger de Coverly wrote:If you really want ex-150 players who have improved to 175 to be given grades of 162.5 feel free to advocate your system. I don't think the rest of us want it because annual lists cause enough trouble with lag as it is. One of the points of a grading system is to rank players by strength. You are not doing this if you "remember" the previous grade when comparing otherwise identical performances. Think about it , you've got 175 standard players with 162.5 grades - what's that going to do to expected scores?

The point is that I think that it is more likely that GS assigns to ex-150 players who have improved to 162.5 level grades of 175, rather than AGS3 would assign to ex-150 players who have improved to 175 level grades of 162.5. Of course, one cannot prove which of the two statements is true without a priori assuming correctness of one of the systems (you probably assuming GS and me AGS3 being correct), but a number of examples (also known to you) showing oddities in the GS system (not present in the AGS3 system) as well as mathematical analysis of grading rules formulae may indicate that 'more correct' system could be AGS3.
Roger de Coverly wrote:I think you have to consider a grading system as measuring performance against a pool of players, because if you don't you get tied up in paradoxes. Could I suggest that you treat the ECF's description of the grading system as an explanation of a calculation method rather than a statement of the underlying theory? Again if you want a grading system which can compute game by game ratings then you use Elo or one of its derivatives.
In my opinion the fact that every player eventually plays a pool of prayers with some average grade cannot be used in assessing which of the grading rules, 1c or 2c, is better.
Roger de Coverly wrote:Your proposed system can be summed up very simply as stating that even for the most active players, their new grade is an average of their performance for a season and their previous season's grade.
Yes, exactly (that is because 'a + c/2 = 1/2*a + 1/2*(a + c)', where 'a' is a grade from previous season and 'c' rule 1c correction applied to the grade this season).
Roger de Coverly wrote:As a consequence it might plausibly be more stable than the current ECF system but is almost certainly more out of date and demoralising to rapidly improving players.
I was thinking if I would keep calculating AGS3 grades for upcoming seasons (so far I have them for 2009) after a couple of seasons (when the difference accumulates, 2008 GS and AGS3 grades were assumed to be equal) somebody could perform similar statistical analysis which has led to grade correction (resulting in recent introduction of new grades) and draw some conclusions about GS vs AGS3 (note that due to a small difference between GS and AGS3 grading rules I can calculate GS grades without having actual game results and all I need is a published GS grades).

Kind regards,
Robert Jurjevic
Vafra

Post Reply