GRADING ANOMALIES

General discussions about ratings.
Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Mon Dec 07, 2009 6:24 pm

Robert Jurjevic wrote:
Roger, please could you let me know if you would agree with the following:
(for the example in my previous post where Carlsen draws all of his 28 games)

1) FIDE penalizes Carlsen for his draws against Howell for (3.9 + 3.6 + 3.3 + 3)/4 = 3.45 rating points in each game on average (one corrects players ratings after each draw)
At 2801 v 2597, the difference is 204 Elo points. This is a score expectation of 0.76 which means a loss to the higher rated player of 2.6 Elo points (K=10 for ratings over 2400). If they play another 3 games in the same rating period, then it's 2.6 for each game. If they play again in the next rating period, then Howell's rating may be higher and Carlsen's lower so the loss of points is less. The international system is not a continuous system with re-rating after every game.

If the difference is rounded to 3 Elo points and there are no other changes, then it becomes 2798 v 2600, a difference of 198 points which is still a score expectation of 0.76 because of the grouped approach used in the FIDE tables. If we move on to the next period, the ratings are 2795 v 2603 and 192 diff gives an expectation of 0.75, so the loss is 2.5 points. It's really quite stable - so over 4 games the loss is about 10 Elo points whether they are in the same rating period or not.

I don't think the Elo system is trying to express opinions as to whether Carlsen is playing at Howell's level or vv. What it's saying is that the assertion that Carlsen is better than Topalov is challenged if he can only draw with Howell. Similarly Howell might now be a little better than someone rated 2598 because he's got a draw with Carlsen.

Robert Jurjevic wrote: 2) GS (current ECF grading system) would penalize Carlsen for his draws against Howell for (2801 - 2597)/30 = 6.8 rating points in each game
For a game count of 30, then yes. If the game count is lower, then it's more. If the game count is higher the effect is reduced. This is the problem that EMW thinks we should worry about - that the player playing 60 games gains/loses less than the player playing 30.
Robert Jurjevic wrote: 3) AGS3 (amended ECF grading system tree) would penalize Carlsen for his draws against Howell for (2801 - 2699)/30 = 3.4 rating points in each game
It's the same effect as if they played 60 games in the season.


Robert Jurjevic wrote: 4) the AGS3's correction of 3.4 rating points is much closer to the FIDE's average correction of 3.45 than the GS's correction of 6.8
Yes because both are systems with built in lag and a memory of the previous grade/rating. The FIDE correction is 2.6 though.
Robert Jurjevic wrote: 5) if ECF would wish to make rating corrections of the order of magnitude as FIDE they should abandon GS and adopt (at lest) AGS3
No - they would fall in line with the rest of the world's national grading systems and use an Elo approach. Issues which would need to be resolved amongst others would be the frequency of rebasing the list ( FIDE now use 2 months), the estimation of new players and the treatment of rapidly improving players including juniors. In addition, everyone would also have to accept that equal performance no longer meant equal grades.
Last edited by Roger de Coverly on Mon Dec 07, 2009 7:15 pm, edited 1 time in total.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Mon Dec 07, 2009 7:13 pm

Roger de Coverly wrote:
Robert Jurjevic wrote:1) FIDE penalizes Carlsen for his draws against Howell for (3.9 + 3.6 + 3.3 + 3)/4 = 3.45 rating points in each game on average (one corrects players ratings after each draw)
At 2801 v 2597, the difference is 204 Elo points. This is a score expectation of 0.76 which means a loss to the higher rated player of 2.6 Elo points (K=10 for ratings over 2400). If they play another 3 games in the same rating period, then it's 2.6 for each game. If they play again in the next rating period, then Howell's rating may be higher and Carlsen's lower so the loss of points is less. The international system is not a continuous system with re-rating after every game.
I understand that as the number of games Carlsen draws against Howell increases the average Carlsen's rating correction (in each game) tends towards zero (until it reaches and stays at zero). Nevertheless, if Carlsen and Howell only play each other and in every game the result is always a draw, Carlsen should never be rated bellow Howell (not after a thousand, a million of games played), FIDE will never rate Carlsen below Howell (once they have played enough games their grade will become and remain equal). On the other hand GS would simply after only 30 draws (in a season) assign to Carlsen a rating of 2597 and Howell a rating of 2801, which is highly illogical, and a direct consequence of the logical flaw in the GS grading rule.
Roger de Coverly wrote:
Robert Jurjevic wrote: 2) GS (current ECF grading system) would penalize Carlsen for his draws against Howell for (2801 - 2597)/30 = 6.8 rating points in each game
For a game count of 30, then yes. If the game count is lower, then it's more. If the game count is higher the effect is reduced. This is the problem that EMW thinks we should worry about - that the player playing 60 games gains/loses less than the player playing 30.
Okay, agree with you, I should have said that the GS penalty would be '(2801 - 2597)/n' where 'n >= 30' is a number of games taken into account for grading.
Roger de Coverly wrote:
Robert Jurjevic wrote:3) AGS3 (amended ECF grading system tree) would penalize Carlsen for his draws against Howell for (2801 - 2699)/30 = 3.4 rating points in each game
It's the same effect as if they played 60 games in the season.
Yes, AGS3 penalty would be '(2801 - 2699)/n', where 'n >= 30' is a number of games taken into account for grading, so '(2801 - 2699)/30' is equal to '(2801 - 2597)/60', but I can't see why would that matter as for a given number of games played 'n', '(2801 - 2699)/n' will always be half of '(2801 - 2597)/n' (they either played 30, 60, etc. games in the season, they couldn't have played at the same time 30 and 60 games).
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Mon Dec 07, 2009 7:33 pm

Robert Jurjevic wrote: I understand that as the number of games Carlsen draws against Howell increases the average Carlsen's rating correction (in each game) will tend towards zero (until it reaches and stays at zero). Nevertheless, if Carlsen and Howell only play each other and in every game the result is always a draw, Carlsen should never be rated bellow Howell (not after a thousand, a million of games played), FIDE will never rate Carlsen below Howell (once they have played enough games their grade would become and remain equal).
This isn't actually true with a "discrete" Elo system like the international one. If you play all the games inside the same rating period ( difficult now but possible when there were only annual lists), then because it's a constant addition for each result it's possible to overshoot and for Howell's rating to go above Carlsen's. In fact it takes 78 games ( 204 / 2.6). Jack Rudd pointed this out regarding the rating of lighthouse keepers.
Robert Jurjevic wrote: On the other hand GS would simply after only 30 draws (in a season) assign to Carlsen a rating of 2597 and Howell a rating of 2801, which is highly illogical, and a direct consequence of the logical flaw in the GS grading rule.
We know the ECF system doesn't work for lighthouse keepers who play the exact number of games needed to qualify for a replacement grade. Elo systems can also overshoot for a suitably high number of games. The ECF system has as a principle "equal grade for equal performance for a suitably high number of games against different people". In terms of a system intended to rank players in order of strength ,that's a valid premise. I can get the equal grade result even for the lighthouse keepers by the simple expedient of publishing a grading list halfway through the match.



Roger de Coverly wrote:It's the same effect as if they played 60 games in the season.
Just to clarify - that's 60 games using the ECF system where the sensitivity of the final grade to one result is halved compared with playing 30 games.

Edited because it's 2.6 not 2.4
Last edited by Roger de Coverly on Tue Dec 08, 2009 3:36 pm, edited 1 time in total.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 08, 2009 10:58 am

Hello Roger,
Roger de Coverly wrote:
Robert Jurjevic wrote: I understand that as the number of games Carlsen draws against Howell increases the average Carlsen's rating correction (in each game) will tend towards zero (until it reaches and stays at zero). Nevertheless, if Carlsen and Howell only play each other and in every game the result is always a draw, Carlsen should never be rated bellow Howell (not after a thousand, a million of games played), FIDE will never rate Carlsen below Howell (once they have played enough games their grade would become and remain equal).
This isn't actually true with a "discrete" Elo system like the international one. If you play all the games inside the same rating period ( difficult now but possible when there were only annual lists), then because it's a constant addition for each result it's possible to overshoot and for Howell's rating to go above Carlsen's. In fact it takes 85 games ( 204 / 2.4). Jack Rudd pointed this out regarding the rating of lighthouse keepers.
As the numbers draws tend to infinity the rating correction applied to either of the players tends to zero, and in the limit (mathematical limit when number of draws is infinite) Carlsen and Howell ratings would be equal to (2801+2597)/2 = 2699 (accurate to infinite number of decimal places) and after that the rating correction applied to either of the players would be (mathematical) zero.

If you think that at some point Howell can be assigned a higher rating than Carlsen, could you please show me how, either by a numerical example using the official FIDE online calculator at http://ratings.fide.com/calculator_rtd.phtml, or by a mathematical argument. Thanks.

Roger, I do not know your mathematical background, but I hope that you know that say:

Code: Select all

1/2 + 1/4 + 1/8 + 1/16 + 1/32 + 1/64 + 1/128 + 1/256 + 1/512 + 1/1024 + ... = 1
i.e., that you can add to '1/2' an infinite number of positive numbers an yet never exeed '1'.
Roger de Coverly wrote:
Robert Jurjevic wrote: On the other hand GS would simply after only 30 draws (in a season) assign to Carlsen a rating of 2597 and Howell a rating of 2801, which is highly illogical, and a direct consequence of the logical flaw in the GS grading rule.
We know the ECF system doesn't work for lighthouse keepers who play the exact number of games needed to qualify for a replacement grade. Elo systems can also overshoot for a suitably high number of games. The ECF system has as a principle "equal grade for equal performance for a suitably high number of games against different people". In terms of a system intended to rank players in order of strength, that's a valid premise. I can get the equal grade result even for the lighthouse keepers by the simple expedient of publishing a grading list halfway through the match.
To me assigning a rating of 2597 to Carlsen and 2801 to Howell after either 30 or a million draws is a serious flaw with a large numerical error and what is more important a clear evidence of a logical mistake in reasoning.

You can call it however you like, a lighthouse effect, a jungle men, but once an example which falsifies a theory is found, IMHO one should correct the theory and not ignore the falsification example (or experiment). There are number of examples in science of a kind, say Einstein vs Newton mechanical theory (when Einstein proposed his theory of special relativity first what people have done was conducted an experiment in attempt to refute Einstein's theory by falsifying it on an example where it predicts different results from Newton's, the experiment in question was now famous "deflection of light by the Sun" experiment, which BTW failed to falsify Einstein's theory, it falsified Newton's theory, Newton theory is not 'wrong' it just gives 'wrong' results in some for Newton theory 'extreme' cases, so Newton's theory is still applicable in a vast number of cases, giving virtually identical results as Einstein's theory; unfortunately GS makes non-negligible error in every single case where the grade correction is not zero, the larger the correction the greater the error, and GS theory is only applicable in cases were the grade correction is zero).

IMHO logical mistakes should be corrected.
Say, would you accept the following logical argument:

Code: Select all

P1: Roger is a man.
P2: Robert has a dark hair.
-----------------------------------------------
Q: Therefore, Roger owe to Robert a 100 pounds.
If one does not correct logical mistakes one may end up with pretty any sort of false conclusions. Validity of a logical argument is a minimum one should require in rational thinking (if a logical argument is valid you can still get a false conclusion, if one or more of the premises are false, but with invalid logical argument you can get a false conclusion even if all of the premises are true).

Kind regards,
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 08, 2009 12:51 pm

Robert Jurjevic wrote:If you think that at some point Howell can be assigned a higher rating than Carlsen, could you please show me how,
It's very simple and a function of the operational rules of the international rating system.

The rating at the start of each rating period applies to all games in that rating period.

Each game that Howell draws earns him 2.4 Elo points. He starts off 204 points behind Carlsen and that difference applies until the end of the rating period and the publication of the next list.

After 85 games in the same rating period, he will have gained 204 points in the end period rating list and he will overtake on the 86th. Players trying to reach a specific rating target such as 2400 for an IM or 2500 for a GM sometimes try to use this effect by playing lots of games over a short period if they are "in form".

You have to realise that rating systems are full of practical compromises. As a consequence extreme examples such as match playing English lighthouse keepers or Carlsen and Howell playing 86 games in two months are not reasons to redesign the whole system.

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 08, 2009 1:14 pm

Robert Jurjevic wrote:To me assigning a rating of 2597 to Carlsen and 2801 to Howell after either 30 or a million draws is a serious flaw with a large numerical error and what is more important a clear evidence of a logical mistake in reasoning.
The simple fix to abolish lighthouse keepers is to add a rule which says that you don't have a grade unless you play at least x different people. x would need to be at least 2. Alternatively you notionally rebase the grades at some point during the match. The ECF system presumes that if you play a number of different people then it's reasonable to assess relative performance using the grades at the start of the period. In part this rule exists because you don't necessarily know the sequence in which games are played. The simple point is that if A and B play all of C,D,E,F,G etc. then you assess whether you think A is a better player than B based on their relative results.

Do you accept the view that the result of a game of chess is random if the players are of equal strength? If they are of unequal strength then you try to infer the difference in strength from the fact that the results seem to be biased in one direction or other. So a tournament winner is a stronger player than the guy who came last. Or is he? If you took 16 players of equal rating and played an all-play-all, one of the least likely results is that they all tie with the same number of points. So someone always wins. You then have to decide how much weight to give to their past reputation and how much to the current result to get a working rating system.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 08, 2009 2:32 pm

Hello Roger,
Roger de Coverly wrote:
Robert Jurjevic wrote:If you think that at some point Howell can be assigned a higher rating than Carlsen, could you please show me how,
It's very simple and a function of the operational rules of the international rating system. The rating at the start of each rating period applies to all games in that rating period. Each game that Howell draws earns him 2.4 Elo points. He starts off 204 points behind Carlsen and that difference applies until the end of the rating period and the publication of the next list. After 85 games in the same rating period, he will have gained 204 points in the end period rating list and he will overtake on the 86th. Players trying to reach a specific rating target such as 2400 for an IM or 2500 for a GM sometimes try to use this effect by playing lots of games over a short period if they are "in form".
Okay, I see, it is possible to assign higher rating to Howell than to Carlsen if one does not grade after every game. May I ask if FIDE changes ratings after every game or after every event such as say a tournament?
Roger de Coverly wrote:You have to realise that rating systems are full of practical compromises. As a consequence extreme examples such as match playing English lighthouse keepers or Carlsen and Howell playing 86 games in two months are not reasons to redesign the whole system.
Wouldn't you agree that a logically sound grading system should never assign higher rating to Howell than to Carlsen in the above example?

FIDE penalizes Carlsen for a draw against Howell 3.9 rating points.

GS would penalize Carlsen for (2801 - 2597)/30 = 6.8 rating points if Carlsen played 30, for (2801 - 2597)/60 = 3.4 grading points if Carlsen played 60 games in the season, etc.

AGS3 would penalize Carlsen for (2801 - 2699)/30 = 3.4 rating points if Carlsen played 30, for (2801 - 2699)/60 = 1.7 grading points if Carlsen played 60 games in the season, etc.

So, both GS's and AGS3's penalties depend on the number of games taken into account for grading and comparison with FIDE penalty may be difficult.

Am I right in understanding that your main objection to AGS3 is that you think that the GS penalty is spot on, while the AGS3 penalty would be too small?

If you could accept that the AGS3 penalty is okay, then IMHO at least you would have a logically sound system.

Can you prove or show why do you think that AGS3 would change grades too little (or equivalently why is GS a system which changes the grades for the exactly right amount)?

Thanks.

Kind regards,
Robert Jurjevic
Vafra

Alex Holowczak
Posts: 9308
Joined: Sat May 30, 2009 5:18 pm
Location: Oldbury, Worcestershire
Contact:

Re: GRADING ANOMALIES

Post by Alex Holowczak » Tue Dec 08, 2009 2:39 pm

Robert Jurjevic wrote:May I ask if FIDE changes ratings after every game or after every event such as say a tournament?
Every 2 months, regardless of how many games have been played in the preceding 2 month period. If an event traverses two rating periods, then all games go in the last one - apart from things like the 4NCL, which get drip-fed in as appropriate (I think).

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 08, 2009 3:03 pm

Roger de Coverly wrote:The simple fix to abolish lighthouse keepers is to add a rule which says that you don't have a grade unless you play at least x different people. x would need to be at least 2.
The fact is that you allocate grading points for each individual game. My formulae at http://www.jurjevic.org.uk/chess/grade/ ... malies.htm are equivalent to the grading rules for allocating points for matches between two players (which in most cases in practice consist of only one game). I would rather impose a rule in which you need to play the same person more than once, as performance against that person is better assessed if you played more than one game, but as in practice such a rule would be practically impossible to enforce without causing great inconvenience to the players, one puts up with assessing one's performance bases on one game only (note how crude estimate that is, there is either score of 0%, 50% or 100%, while if players played more than one game the score could have been say 75%, etc.)

The whole point about lighthouse keepers is not to swipe them under the carpet but to open one's eyes and ask a question why is this happening?

IMHO for a good estimate of one performance one would need to play a number of games against the same opponent but also to play enough different opponents (say I can play well against one player as his style may suite me, but no so well against other player even their grades may be close).
Roger de Coverly wrote:Do you accept the view that the result of a game of chess is random if the players are of equal strength? If they are of unequal strength then you try to infer the difference in strength from the fact that the results seem to be biased in one direction or other. So a tournament winner is a stronger player than the guy who came last. Or is he? If you took 16 players of equal rating and played an all-play-all, one of the least likely results is that they all tie with the same number of points. So someone always wins. You then have to decide how much weight to give to their past reputation and how much to the current result to get a working rating system.
Yes, I agree that any game result is possible though not equally probable regardless of the grades of players involved. The point is that once two players performed in a mini-match (usually only one game) if their performance is not as expected one has to correct their grades based on the difference between actual and expected performance in that game by allocating grading points to the players. What I logically require in AGS3 (and my other rules) is that the grading points are allocated so that the total grade correction due to grade allocation for that game equals the difference between actual and expected performance, while GS rule requires that the grading points are allocated so that the total grade correction due to grade allocation for that game equals double the difference between actual and expected performance, which to me is not logical. Note that the actual grade correction per game depends on the number of games you take into account for grading, and you would take the same number of games if you used GS or AGS3 rule, we are talking about differences in rules applied for allocating points in individual games, which should not depend on the number of games you take into account for grading (if a player played 60 games you have to take 60 games, if a player played 40 games you have to take 40 games, etc.).
Robert Jurjevic
Vafra

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Dec 08, 2009 3:53 pm

Alex Holowczak wrote:
Robert Jurjevic wrote:May I ask if FIDE changes ratings after every game or after every event such as say a tournament?
Every 2 months, regardless of how many games have been played in the preceding 2 month period. If an event traverses two rating periods, then all games go in the last one - apart from things like the 4NCL, which get drip-fed in as appropriate (I think).
Thanks Alex, may I ask if the rating published every 2 months is calculated by rating after every game using live ratings or a constant rating at the beginning of the period is used in rating of each game in the period? Thanks.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 08, 2009 4:09 pm

Robert Jurjevic wrote:Am I right in understanding that your main objection to AGS3 is that you think that the GS penalty is spot on, while the AGS3 penalty would be too small?

If you could accept that the AGS3 penalty is okay, then IMHO at least you would have a logically sound system.

Can you prove or show why do you think that AGS3 would change grades too little (or equivalently why is GS a system which changes the grades for the exactly right amount)?

As I see it the issue is that we have (or perhaps had) an operational grading system which worked for many years on the basis of 50 points for a win and has the logical underpin that equal performances generate equal grades. It's used for tournaments and leagues, so any difficulties with the lighthouse keeper swop effect are irrelevant.

The decision point is to either retain this system or move over to a version of an Elo system which more explicitly tracks game by game performance.

There's no point in modifying the ECF system to slow down grading changes by averaging everyone against their previous grade. Fundamentally if 3 players all play equivalent people over at least 30 games with a performance of 175 (based on the most recent published grading list), then they should all get grades of 175 regardless of whether (a) they had a previous grade of 175 or (b) they had a previous grade of 150 or (c) whether they had a previous grade of 200 or (d) they are a player new to English chess.

An Elo style system would only give the same grade to (a) and (d). It's long known that Elo style systems remember earlier ratings. In some circumstances, this is a disadvantage but this issue is perhaps outweighed by coping better with more frequent list publication and less active players.

Grading systems have to be practical to apply rather than logical in extreme circumstances.

Robert Jurjevic wrote:Can you prove or show why do you think that AGS3 would change grades too little (or equivalently why is GS a system which changes the grades for the exactly right amount)?
Suppose you consider the very top of the food chain. Suppose for many years the top players are first amongst equals at around 250. Suppose a player comes along who can score 75% against all these 250's.
How does his grade ever reach 275 in AGS3? He scores 15 draws and 15 wins. First season his grade is ( 250 *30 + 25 * 15) /30 = 262.5. Second season (30*( 250 + 262.5)/2 +25 *15) /30 = 268.75. Third season ( 30*(250+268.75)/2 +25*15) = 271.88. As he gains fewer and fewer points each season, he never reaches 275. I'm assuming that the 250s manage to retain their 250 status by beating someone else.

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Dec 08, 2009 4:39 pm

Robert Jurjevic wrote: Thanks Alex, may I ask if the rating published every 2 months is calculated by rating after every game using live ratings or a constant rating at the beginning of the period is used in rating of each game in the period? Thanks.
You could try looking at the actual site. Here's Magnus
http://ratings.fide.com/individual_calc ... 2010-01-01

Each and every calculation is based on the ratings at the start of the period namely the 1st November published list. I don't think any other approach is practical because FIDE cannot control the order in which events are submitted. I don't even think the order shown in an event is always the order in which games were played.

"Live" ratings are a prediction of what the next published list will say assuming no more games are played. We're not expecting Topalov to play any more before Christmas, so Magnus becomes the new number one if he doesn't lose points at the London event.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Wed Dec 09, 2009 4:12 pm

Hello Roger,

In the ECF official grading statement it is said that grading "points are allocated in respect of each game." and that one's final "grade is calculated by dividing the total number of points scored by the number of games played" which is an average of grading points scored in each game.

Various rules for allocating grading points with respect of each game can be expressed with the following formulae:

Code: Select all

a2 = a + ka*(q - p);
b2 = b + kb*((100 - q) - (100 - p));
where 'a' is your grade, 'b' grade of your opponent, 'p' your expected performance (expected performance of your opponent is then '100 - p'), 'q' your actual performance (actual performance of your opponent is then '100 - q'), 'a2' your new grade (your grading points allocated for the game) and 'b2' your opponent's new grade (your opponent's grading points allocated for the game).
(if the players played only one game in the season 'q' is either 100, 0 or 50, if they played more than one game it can be a number between 0 and 100 inclusively)

What makes the rules different is a choice of factors 'ka' and 'kb' and function 'p = f(d)'.

The difference between GS rule and AGS3 rule is in factors 'ka' and 'kb' only, for GS 'ka = kb = 1' and for AGS3 'ka = kb = 1/2'. Both GS and AGS3 use the same linear 'p = f(d)' shown in the figure 1 below (green line).

Image
Figure 1: Relationship between expected performance 'p' and grade difference 'd' as defined in GS (green line), CGS, AGS and AGS2 (blue line), ÉGS, ÉGS2, ÉGS3 and ÉGS4 (red line), ÉGS5 and ÉGS6 (yellow line), and (normal relationship 'p = 100*(1 + Erf[d/50])/2', where the error function Erf[z] is the integral of the Gaussian distribution) as originally defined by Élo (brown line above yellow). Expected performance 'p' is a function of grade difference 'd', i.e., 'p = f(d)'. Note that both FIDE and USCF switched from normal (brown line) to logistic (yellow line) relationship 'p = f(d)' which they found provides a better fit for the actual results achieved.

Although grading points are allocated with respect of each game, this is done once a season after the players had played all of their games. This fact can be utilized to asses how to distribute penalty and reward when allocating grading points in respect of each game, and instead of using 'ka = kb = 1/2' we could calculate 'ka' and 'kb' using the following formulae:

Code: Select all

ca = Abs[(qa - 50 - da)/2];
cb = Abs[(qb - 50 - db)/2];
c = ca + cb;
ka = If[c > 0, ca/c, 1/2];
kb = If[c > 0, cb/c, 1/2];
where 'da' is grade difference between your grade and average grade of your opposition (positive if your grade is above average grade of your opposition), 'db' is grade difference between your opponent's grade and average grade of his opposition (positive if your opponent's grade is above average grade of his opposition), 'qa' your actual performance against your opposition, 'qb' your opponent's actual performance against his opposition (if your grade differs from average grade of your opposition by more than 40 points, it is taken to be exactly 40 points above (or below) yours, if your opponent's grade differs from average grade of his opposition by more than 40 points, it is taken to be exactly 40 points above (or below) his.

AGS4: We define AGS4 (Amended Grading System four) to be AGS3 with factors 'ka' and 'kb' calculated using above formulae (rather than taking 'ka = kb = 1/2').

Image
Figure 2a: Factor 'ka' (used in AGS4) as a function of 'qa' and 'qb' where 'da = -40' and 'db = 0'.

Image
Figure 2b: Factor 'ka' (used in AGS4) as a function of 'qa' and 'qb' where 'da = -20' and 'db = 0'.

Image
Figure 2c: Factor 'ka' (used in AGS4) as a function of 'qa' and 'qb' where 'da = 0' and 'db = 0'.

Image
Figure 2d: Factor 'ka' (used in AGS4) as a function of 'qa' and 'qb' where 'da = 20' and 'db = 0'.

Image
Figure 2e: Factor 'ka' (used in AGS4) as a function of 'qa' and 'qb' where 'da = 40' and 'db = 0'.
Roger de Coverly wrote:
Robert Jurjevic wrote:Can you prove or show why do you think that AGS3 would change grades too little (or equivalently why is GS a system which changes the grades for the exactly right amount)?
Suppose you consider the very top of the food chain. Suppose for many years the top players are first amongst equals at around 250. Suppose a player comes along who can score 75% against all these 250's. How does his grade ever reach 275 in AGS3? He scores 15 draws and 15 wins. First season his grade is ( 250 *30 + 25 * 15) /30 = 262.5. Second season (30*( 250 + 262.5)/2 +25 *15) /30 = 268.75. Third season ( 30*(250+268.75)/2 +25*15) = 271.88. As he gains fewer and fewer points each season, he never reaches 275. I'm assuming that the 250s manage to retain their 250 status by beating someone else.
If the 250 players played approximately as expected (i.e. approximately at their level of 250, say every 250 player performed on average 50% against approximately 250 field) AGS4 would assign to the player who scored 75% a grade of 275, i.e. AGS4 would reward the 75% player for the maximum amount in every game he played against the 250's, but it would not penalize (nor reward) either of the 250 players for his game against the 75% player (games the 250 players played against the 75% player wouldn't count in grading of the 250 players).

GS would assign to the player who scored 75% a grade of 275, i.e. AGS4 would reward the 75% player for the maximum amount on average in every game he played against the 250 players, but it would penalize each of the 250 players for the maximum amount on average for his game against the 75% player (though this penalty may be relatively small as it may be done for relatively few games in respect to number of games each 250 player did play, it is still a penalty which should not have been done, also there are cases where this error may not be small at all, say the example of the lighthouse keepers).

AGS3 would assign to the player who scored 75% a grade of 262.5, i.e. AGS3 would reward the 75% player for a half of the maximum amount on average in every game he played against the 250 players, and it would penalize each of the 250 players for a half of the maximum amount on average for his game against the 75% player (AGS3 simply does not have a means to asses how to split the penalty and the reward, so it does it 50/50).

Note that it is a coincidence that in GS rule you do not need to know your grade in order to calculate your new grade (it is enough to know your actual performance and grades of your opponents), for all other rules you need to know your grade as well, and in general you need to know your grade. GS rule can also be expressed in a way so that you need to know your grade, say rules 1a and 1c below are equivalent:

Rule 1a: For a win you score your opponent's grade plus 50; for a draw, your opponent's grade; and for a loss, your opponent's grade minus 50. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

Rule 1c: For a win you score your grade plus 50 minus grade difference; for a draw, your grade minus grade difference; and for a loss, your grade minus 50 minus grade difference. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

So, there is a maximum allowed amount (in grading points) for which you can reward o penalize the 75% player (as if he is not un-graded he must have a grade), if you choose to reward the 75% player for the maximum allowed amount then you should not penalize the 250 players for their games they have played against the 75% player (as already all it could have been given or taken was given to the 75% player), but you can choose to splilt the reward and the penalty 50/50, which AGS3 does. AGS4 is more sophisticated and would spit the reward and the penalty depending on the circumstances (using actual performances of the 250 players against their oppositions AGS4 tries to asses how to distribute the penalty and the reward in each game the 75% player has played).
Roger de Coverly wrote:There's no point in modifying the ECF system to slow down grading changes by averaging everyone against their previous grade. Fundamentally if 3 players all play equivalent people over at least 30 games with a performance of 175 (based on the most recent published grading list), then they should all get grades of 175 regardless of whether (a) they had a previous grade of 175 or (b) they had a previous grade of 150 or (c) whether they had a previous grade of 200 or (d) they are a player new to English chess.
Assuming that (as stated above) there is a maximum allowed amount (in grading points) for which you could reward or penalize players, the answer depends on the rule one did use and how the rule splits the reward and the penalty. 'Correct' rules should IMHO obey the axiom which states that grading points should be allocated in each game so that the total grade correction equals the difference between expected and actual performance (which is to me a perfectly logical and natural requirement), i.e, '(a2 - a) + (b - b2) = q - p' for any 'a', 'b', 'p' and 'q'.

Why in GS you do not need to know your grade in order to allocate grading points in each game? One allocates 'a2' grading points to you in each game according to the formula:

Code: Select all

a2 = a + ka*(q - p);
where 'a' is your grade, 'b' grade of your opponent, 'p' your expected performance (expected performance of your opponent is then '100 - p'), 'q' your actual performance (actual performance of your opponent is then '100 - q'), 'a2' your new grade (your grading points allocated for the game) and 'b2' your opponent's new grade (your opponent's grading points allocated for the game).

GS's 'p = f(d)' (for 'd <= 40') is:

Code: Select all

p = 50 + d;
so 'a2 = a + ka*(q - p) = a + ka*(q - 50 + d)' and as in GS 'ka = 1' one has 'a2 = a + (q - 50 + d) = a + d + q - 50' and as 'a + d = b' is your opponent's grade you do not need to know your grade 'a' in order to calculate 'a2', as 'a2 = b + q - 50', where 'b' is your opponent's grade. Nevertheless, the correction one applies to your grade is 'ka*(q - p)', and the correction one apples to your opponent's grade is 'kb*((100 - q) - (100 - p)) = kb*(p-q)'. The question is which 'ka' and 'kb' to choose, i.e. how much to correct your grade and the grade of your opponent?

It can be proven that if '(a2 - a) + (b - b2) = q - p' for any 'a', 'b', 'p' and 'q'. then 'ka + kb = 1'. Unfortunately, GS's 'ka = ka = 1', so 'ka + kb = 2'.

Kind regards,
Last edited by Robert Jurjevic on Wed Dec 09, 2009 4:48 pm, edited 2 times in total.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Wed Dec 09, 2009 4:45 pm

Robert Jurjevic wrote:'Correct' rules should IMHO obey the axiom which states that grading points should be allocated in each game so that the total grade difference equals the difference between expected and actual performance (which is to me a perfectly logical and natural requirement).
For all the fancy graphs, you are a minority of one in believing this (and in holding the belief that there is something metaphysical and important about allocating grading points game by game). In the context of practical grading systems (which will always have an element of estimation about them), it is a much better result that a player scoring 75% against 250 opposition over a reasonable number of games should have a grade of 275 rather than 262.5 even if the consequence is that the 250 opposition is diluted to 249. You either accept that the system is deflating a little, or add a fiddle point back by means of a junior increment or otherwise. The deflation dilution is only conditional anyway since the results of players playing a fewer or greater number of games than 30 have an impact. Again we're looking for a rating system which ranks players in strength order whilst also expressing opinions on the likely results when they meet.

In the Elo approach you add something which is a function of the actual v expected performance. The function depend on how "surprising" the result is, So if Howell beats Carlsen, you are still fairly sure (K=10) that the respective ratings are "about" correct and just make a 2.6 change. The interpretation of this is that you don't now believe that Carlsen is better than Topalov. If they were 1000 points lower and inside their first 30 games, you would be less confident that the ratings were correct and therefore adjust them by 6.5 points. (K=25).

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Wed Dec 09, 2009 5:33 pm

Hello Roger,
Roger de Coverly wrote:For all the fancy graphs, you are a minority of one in believing this (and in holding the belief that there is something metaphysical and important about allocating grading points game by game).
Sure, you can say that the axiom which requires that '(a2 - a) + (b - b2) = q - p' for any 'a', 'b', 'p' simply needs not to hold! Fair enough.

Note that for GS it holds that '(a2 - a) + (b - b2) = 2*(q - p)' for any 'a', 'b', 'p'. May I ask if you are suggesting that this should be the axiom, or that one should not impose any restraint on the realtionship between '(a2 - a) + (b - b2)' and '(q - p)'? Thanks.
Roger de Coverly wrote:In the context of practical grading systems (which will always have an element of estimation about them), it is a much better result that a player scoring 75% against 250 opposition over a reasonable number of games should have a grade of 275 rather than 262.5 even if the consequence is that the 250 opposition is diluted to 249.
Note that AGS4 (a newly introduced rating rule with 'fancy' graphs) would allocate to the 75% player a grade of approximately 275 (rather than 262.5) if it assessed that every 250 payer performed approximately as expected (i.e., at approximately 250 level) against his opposition (you will find some details about AGS4 in my previous post).

If my axiom should hold than AGS4 should be better than GS, as though in this example both of the systems would assign to the 75% player a grade very close to 275 (GS would assign exactly 275 but it would do that by chance), when grading the 250 players for their games against the 75% player AGS4 would allocate for each 250 player a grade very close to his grade from previous season (effectively when grading 250 players AGS4 would neither penalize nor reward the 250 players in the games they have played against the 75% player), while when grading the 250 players for their games against the 75% player GS would allocate for each 250 player a grade which is lower than his grade from previous season (effectively when grading 250 players GS would penalize the 250 players in the games they have played against the 75% player, even this must not have been done if my axiom should hold).

Kind regards,
Robert Jurjevic
Vafra

Post Reply