GRADING ANOMALIES

General discussions about ratings.
E Michael White
Posts: 1420
Joined: Fri Jun 01, 2007 6:31 pm

Re: GRADING ANOMALIES

Post by E Michael White » Fri Dec 04, 2009 10:34 am

Roger de Coverly wrote:I think Clarke and the BCF first noticed that treatment of improving players was a design issue back in the 1960s
They have to be more active as well as improving which I dont think was mentioned. It is possible for a junior to join the system at grade 80 improve every year and give up at 220 and cause some inflation ! It all depends on the playing patterns of who plays whom.

In your 31 player scenario if the x150 player had another harmless looking drawn game against another 150 player from outside it becomses deflationary. Similarly if you remove one game between 2 of the 175s in the original scenario it becomes deflationary. This latter point has been understood for a few years in the effect of withdrawls from all play alls which can also be inflationary depending on the playing and result patterns.

Roger de Coverly
Posts: 22608
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Fri Dec 04, 2009 10:58 am

E Michael White wrote:In your 31 player scenario if the x150 player had another harmless looking drawn game against another 150 player from outside it becomses deflationary.
If it's a 150 with a 150 standard of play, then his results have not quite made 175. If it's a 150 playing to a 175 standard then yes. If we use the junior analogy, then they play each other of course whilst improving.
E Michael White wrote:Similarly if you remove one game between 2 of the 175s in the original scenario it becomes deflationary.
I chose 31 players playing 30 games to be in the most stable zone of the ECF system and where the game count has its minimum effect. Players playing 1*150 and 28 * 175 would get 1*175 averaged back in from the previous year, so there would be no effect. With a larger or smaller game count, there would be an issue of course. It just goes to show how finely balanced the interactions between inflationary and deflationary forces actually are. The other point is that we shouldn't worry, since over 30 games the grading system is only accurate ( in the sense that player A is better than player B) to about 8 ECF points. I suppose this is why they banded the earliest lists to cover about 8 ECF points ( 1a, 1b etc.) .

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Fri Dec 04, 2009 2:12 pm

Hello all,
Paul McKeown wrote:Elo system updates game per game. Damping is needed.
BCF system takes an annual average. Damping would damage its accuracy.
ECF grade of a player is an average (over the games) of allocated grading points (to the player) in respect of each game (the player had played). Consequently, the ECF grade of a player is directly dependent on the choice of the rule which advises how to allocate grading points (to a player) in respect of each game (the player had played).

GS (rule 1a is equivalent to rule 1c)
Rule 1a: For a win you score your opponent's grade plus 50; for a draw, your opponent's grade; and for a loss, your opponent's grade minus 50. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.
Rule 1c: For a win you score your grade plus 50 minus grade difference; for a draw, your grade minus grade difference; and for a loss, your grade minus 50 minus grade difference. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

AGSS3 (rule 2b is equivalent to rule 2c)
Rule 2b: For a win you score average grade plus 25; for a draw, average grade; and for a loss, average grade minus 25. Average grade is half of the sum of your and your opponent's grade. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.
Rule 2c: For a win you score your grade plus 25 minus half grade difference; for a draw, your grade minus half grade difference; and for a loss, your grade minus 25 minus half grade difference. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly not 40 points above (or below) yours. At the end of the season an average of points-per-game is taken, and that is your new grade.

The above rules are used to provide a recipe on how to allocate grading points in respect of each game, and they, as well as ÉGS5 and ÉGS6 rules (which cannot be expressed in simple words like rules 1a and 2b), can be mathematically represented with the following formulae:

Code: Select all

a2 = a + ka*(q - p)
b2 = b + kb*((100 - q) - (100 - p))
where 'a' is your grade, 'b' grade of your opponent, 'p' your expected performance (expected performance of your opponent is then '100 - p'), 'q' your actual performance (actual performance of your opponent is then '100 - q'), 'a2' your new grade (your grading points allocated for the game) and 'b2' your opponent's new grade (your opponent's grading points allocated for the game).
(if the players played only one game in the season 'q' is either 100%, 0% or 50%, if they played more than one game it can be a number between 0% and 100% inclusively)

What makes the rules different is a choice of factors 'ka' and 'kb' and function 'p = f(d)'. The difference between rule 1a and 2b is only in factors 'ka' and 'kb', for rule 1a 'ka = kb = 1' and for rule 2b 'ka = kb = 1/2' (they both use the same linear 'p = f(d)' which graph is shown as green line in figure 2 at http://www.jurjevic.org.uk/chess/grade/ ... malies.htm).

The above formulae are a bit more general than rules 1a and 2b, as they allow for the cases where two players play more than one game, accepting actual performance 'q' to be virtually any number between 0% and 100%, not only 100%, 0% or 50% which are the only possible actual performances in one game.

When trying to compare different rules (different 'ka' and 'ka' factors and different 'p = f(d)' functions) it may be more convenient to assume that above formulae are applied in cases where the two players played a match of say 30 games rather than they played only 1 game (the only difference being is that actual performance of the players is more accurately estimated if they played 30 rather than only 1 game, nevertheless, if they played only 1 game then one is left with the only choice to estimate their actual performances from that single game which is rather crude as it can be either 100%, 0% or 50%, there are no cases such as 75%, 25%, etc.). (the above formula are in ECF system in general applied to allocate grading points in one game only as the same two players rarely play two or more games in a season)

!
I argue that the 'k' factors 'ka' and 'kb' in the above formulae cannot be arbitrarily chosen (this is my 'axiom' and can be attacked, but, as you will see, I think it makes a lot of sense, as every 'axiom' should). (FIDE mentions only one 'k' factor, that is possibly because in their system it could be that always 'ka = kb = k')

My mathematical requirement ('axiom') for a grading system (using the above formulae) not to stretch nor shrink the grades (that is how I refer to this requirement) is that

Code: Select all

(a2 - a) + (b - b2) = q - p
for any 'a', 'b', 'p' and 'q'.

In other words one requires that corrected grades match the actual performance of the players (or that a total correction applied to both grades is equal to the difference between expected and actual performance). I think you would agree that this makes a lot of sense, as you simply require that the grades should be corrected in such a way that if the players continued playing like this between themselves (their actual performances did not change) their grades would not change.

!
It can be proven that if '(a2 - a) + (b - b2)' is equal to 'q - p' for any 'a', 'b', 'p' and 'q' then it must be

Code: Select all

ka + kb = 1
So, you are still left with lot of choice for 'ka' and 'kb' factors providing that you satisfy that 'ka + kb = 1', i.e., you are free to chose how much to penalize one and reward other player but must take care that the total grade correction equals the difference between actual and expected performance.

Unfortunately for GS (rule 1a) 'ka = kb = 1', so 'ka + kb = 2' and GS does not satisfy condition 'ka + kb = 1'.

Here is an example:
If Robert 150 drew against Roger 175 (this can be one game or a match, if it is a match then Robert scored 50%), allocated grading points for that game (or match) for Robert and Roger would be as follows:

GS: 175.00 to Robert and 150.00 to Roger
AGS: 162.50 to Robert and 162.50 to Roger
ÉGS5: 162.99 to Robert and 162.01 to Roger
ÉGS6: 162.99 to Robert and 162.01 to Roger (assuming that Robert and Roger played equal number of games in the previous season)

I am arguing that allocating 175 to Robert and 150 to Roger for the game (or a match) in which Robert drew (or scored 50%) against Roger is a mistake.

Allocating 162.50 to Robert and 162.50 to Roger, or 162.99 to Robert and 162.01 to Roger, would be acceptable (even allocating 175 to Robert and 175 to Roger or 150 to Robert and 150 to Roger would be acceptable).

The problem with allocating 175 to Robert and 150 to Roger is that (at least to me) if you assumed that Robert played at 175 level then you should have assumed that Roger too played at 175 level (simple case of Roger playing his standard game and Robert improving), and if you assumed that Roger played at 150 level then you should have assumed that Robert too played at 150 level (simple case of Robert playing his standard game and Roger worsening), assuming both that Robert played at 175 level and Roger at 150 level is (at least for me simply) a logical contradiction, as clearly their actual performance suggests that they played at the same level.

In general there is not enough info to see if only one of the players improved or one player improved more that the other worsened, and the best bet should be then to assume that as much as one player improved so much the other worsened, and to allocate 162.50 to both players.

Note that a slightly asymmetric distribution of reward and penalty, allocating 162.99 to Robert and 162.01 to Roger, in Élo cases above is caused by non-linearity of 'p = f(d)'. The logistic (non-linear) relationship between 'p' and 'd' used in ÉGS5 and ÉGS6 was shown to fit the experimental data the best of all so far examined relationships. That is to say, on average it is more likely that a weaker player improved a bit more than a stronger player worsened, than as much as a weaker player improved so much a stronger player worsened (the greater the grade difference the greater the imbalance, but for grade difference of approximately 30 grading points or less it can be assumed that the imbalance is approximately zero, i.e. as much a weaker player improved so much a stronger player worsened).

Please note that grades are relative, so when one says that Roger worsened from 175 to 150 his absolute playing strength may have remained constant, as say absolute playing strength of Robert and of all his chess fellows with grade of 150 may have increased for the same amount, so that Robert relatively to his peers remained a 150 player, while at the same time the absolute playing strength of Rogers chess fellows with grade of 175 may have increased even more than Robert's, so that Roger is now relatively to his peers a 150 player. That is how Roger can worsen even not playing worse in absolute terms.
Roger de Coverly wrote:
Robert Jurjevic wrote:That is because you cannot assume that all opposition players played on the level suggested by their grades
Why not?
Actually, you can but I think you shouldn't!

If there is a difference between actual and expected performance, not having on yours disposal any specific info (like this player has learnt a lot more about chess than his opponent, etc.), the best bet should be to split the reason for discrepancy between actual and expected performance equally on both players, and penalize the under-performer for the same amount (of 12.5 grading points on average) as award the over-performer.

You cloud though assume that the player (over-performer) should be rewarded for a full amount (25 grading points on average) in each game he played against the opposition (though I really see no reason why you would assume such a thing), but then you must not penalize a single player in the opposition, which GS unfortunately does (GS penalizes every single player in the opposition for 25 grading points on average, even though it rewards the player for a full amount of 25 grading points on average in each game).

Maybe you think that the players can only improve and if Robert 150 drew against Roger 175 one has to allocate 175 to Robert and 175 to Roger, but please be reminded that the grades are relative and the case where one allocates 150 to Robert and 150 to Roger is perfectly possible, please see my paragraph above where I show how Roger can worsen grade-wise even not playing worse in absolute terms. Nevertheless, still the best bet should be to allocate 162.5 to Robert and 162.5 to Roger, unless there is some special info on how one should distribute the award and the penalty (say if Robert was un-graded and Roger an established player then the best bet should have been to allocate 175 to Robert and 175 to Roger, as Robert's grade should not be trusted at all in respect to Roger's and one should assume that Roger played at his level of 175 and that Robert's estimated grade should have been 175 rather than 150).

Kind regards,
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22608
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Fri Dec 04, 2009 3:37 pm

Robert Jurjevic wrote: In other words one requires that corrected grades match the actual performance of the players
This is an axiom which neither an Elo system nor an ECF system fully accepts. I think of both of them as a bit Bayesian in that they treat the result of a game as evidence that the playing strength may have shifted but not conclusive. ECF requires 30 games to accept evidence that a player has shifted and in Elo it's derived from the K factor. So the fact that a 150 player plays a 175 player and draws is marginal evidence that the 150 player has got better or that the 175 player has got worse. In arithmetic terms, if you know for certain that both players will play (or have counted) exactly 30 games then you could adjust the 150 player to (150*29/30) + 175 * (1/30) = 150.83 with the 175 player dropping by the 0.83 to
174.17.

As an actual formula it's

(last year grade) * (30-n) /30 + (current year performance) * n/30 for n <30 and last year grade based on at least 30

and (current year performance) for n>=30 .

Current year performance is measured in the same way as it would be for a new player namely (sum of opponents grades plus 50 * (win-loss) ) / n


You don't know for certain that 30 games will be counted so you have to perform the calculation using the 175 * 1 bit. This simple arithmetic point seems to be confusing you totally into thinking that the 175 in the calculation has some metaphysical meaning relating to assumed playing standard.

So a player playing one game in a season ( 30 in the previous season) would be rewarded or penalised 0.83 of a grading point for drawing with a player 25 points above or below. The 0.83 result applies up to and including 30 games after which it is scaled down based on the number of games played. So if you play 30 games you get 30 * 0.83 = 25. Beyond that you have to scale it back so as not to overshoot. That basically is your answer has to how an individual game is rated and it demonstrates some degree of similarity in treatment to an Elo system. For a K of 15, an Elo system would only give you 0.47 points ( 0.25* 15 /8) . I think you would award 0.415 for each win. This is closer to the Elo approach and would converge more slowly if at all to the new strength (lag) .


How would you for example rate Carsen's performance in the forthcoming London event.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Fri Dec 04, 2009 5:04 pm

Roger de Coverly wrote:
Robert Jurjevic wrote:In other words one requires that corrected grades match the actual performance of the players
This is an axiom which neither an Elo system nor an ECF system fully accepts. I think of both of them as a bit Bayesian in that they treat the result of a game as evidence that the playing strength may have shifted but not conclusive.
I think that my axiom is the best bet of all I've seen so far. If you find any better axiom, please let me know, and I will accept it at once.
Roger de Coverly wrote:
Robert Jurjevic wrote:ECF requires 30 games to accept evidence that a player has shifted and in Elo it's derived from the K factor.
I hope you see that ECF system 'shifts' the players' grades in every single game for twice the amount than what my axiom requires (what the 30 games have to do with the rule how to allocate grading points in respect of each single game).

The ECF grade (based on at least 30 games) is an average (over single games) of allocated grading points in respect of each game (grade 'shift' is done in every game by allocating grading points in respect of each game).

Or let me put this way, if you would chose to allocate grading points in each game so that each player gets allocated his ECF grade from previous season regardless of the game result, the average of all allocated grading points in respect of each game would equal the ECF grade from previous season for every player, and you will end up with a system with fixed grades forever.

My point is that the ECF grade correction depends on the rule which talks about how to allocate grading points in respect of each game (these rules do not care if 30, 60 or 100 games have to be played in order for the results to be treated statistically significant), and my axiom is a constraint in those rules, rules for grading individual games (or to be more precise, constraint in rules for grading matches between two players in a season), my formulae are in fact a way of presenting a number of those rules using a single mathematical form (the axiom is applied to that formulae, hence it applies to the rules which advise how to allocate grading points in respect of each game).

!
I would be so glad if you could see this. Thanks.
Last edited by Robert Jurjevic on Fri Dec 04, 2009 5:35 pm, edited 9 times in total.
Robert Jurjevic
Vafra

E Michael White
Posts: 1420
Joined: Fri Jun 01, 2007 6:31 pm

Re: GRADING ANOMALIES

Post by E Michael White » Fri Dec 04, 2009 5:12 pm

Roger de Coverly wrote:I chose 31 players playing 30 games to be in the most stable zone of the ECF system and where the game count has its minimum effect. Players playing 1*150 and 28 * 175 would get 1*175 averaged back in from the previous year, so there would be no effect.
Yes with lots of exceptions and variations. eg if one of the 29ers had a recent record going backwards of 10 at 170 180 175

The two biggest problems for the ECF rating system are that it doesnt deal very well with varying activity and heteroscedasticity when 2 players play more than 1 game during a rating period. The latter is likely to be an increasing issue if increased numbers of pensioners play regularly against one another in say a local laegue and tournaments. While this can be adjusted for it makes the calculations more difficult which destroys one attraction for players which is that they can calculate their grades easily and regularly.

Roger de Coverly
Posts: 22608
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Fri Dec 04, 2009 5:37 pm

Robert Jurjevic wrote: my axiom is a constraint in the rules for grading individual games (or to be more precise, constraint in the rules for grading matches between two players in a season)
Quite honestly there's no great use for a system which only grades matches between lighthouse keepers (and if you are dealing with lighthouse keepers Elo methods work as well). What' s needed is a system which ranks players in order of strength when they meet a variety of opposition in leagues and tournaments. You need the rankings to be more or less correct so that seeded pairing systems, board orders, restricted eligibility etc. can perform properly .The added lag and slow convergence of adding 25 points for a win is less satisfactory for this purpose than the ECF system . Take my 100/125/150/175/200 as an example. You could have probable 200 standard players allowed to play in under 180 events and mathematically the 200 standard player can never reach exactly 200 no matter how many games he plays.
We have a player who plays 30 games each season always scoring 50% but he moves up a class every year until he peaks at 200. So if he starts at 100, then his first grade is 100, second 125, third 150, fourth 175 and final 200. Under your system I think his grades are 100 , (100+125)/2 = 112.5, (112.5+150)/2 = 131.25, (131.25+175)/2 = 153.25, (153.25+200)/2 = 176.75.
So as a benchmark of a grading system - If two players score the same results (not against each other) over a sufficiently high number of games, does the system give them the same rank? Elo type systems actually fail this test but have other compensating advantages.

Roger de Coverly
Posts: 22608
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Fri Dec 04, 2009 5:54 pm

E Michael White wrote:
problems for the ECF rating system are that it doesnt deal very well with varying activity
You might be able to deal with varying activity at the expense of lag. Assume you know the date of every game and that results are posted with sufficient speed to enable monthly or quarterly lists. Then the new grade is always the results for the most recent 30 games no matter how far back in time you go (within reason!). So except for players with more than 30 games or players with less than 30 qualifying games, you would always be using (1/30) so the interchange of points would be mostly zerosum. I'd suspect less active improving juniors would still present major issues though.

Mind you if you could upgrade things to get every game dated and reported quickly enough for monthly or quarterly list updates, I'd imagine you would go Elo.

Even with monthly lists you would still get players who played more than 30 games in a period. If you play morning and afternoon at the British, that's 25 in one go. Sean's all play alls could give you another 9 and then there's the bank holiday weekend congresses.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Fri Dec 04, 2009 6:32 pm

Hello Roger,
Roger de Coverly wrote:How would you for example rate Carsen's performance in the forthcoming London event.
I wouldn't grade him unless I have at least 30 games he played in English territory each of which has not been played more than three years ago (from the moment when his grade was meant to be calculated).

!
Roger, I am going to be at the The London Chess Classic on 14th of December, will also play in Korchnoi simultaneous display, I wondered if you will be there on that day and if you might wish to have a quick chat with me? Thanks.

Kind regards,
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22608
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Fri Dec 04, 2009 7:25 pm

Robert Jurjevic wrote:I wouldn't grade him unless I have at least 30 games he played in English territory each of which has not been played more than three years ago (from the moment when his grade was meant to be calculated)
I'll convert the Elo's then and calculate it. We've established using the FIDE rules that he needs to score 5/7 (3 wins 4 draws say) to make a rating gain and stay ahead of Topalov at the top of the rating list.

If we convert to ECF using (FIDE-650)/8 = ECF then

Carlsen 269
Kramnik 265
Nakamura 258
Short 257
Adams 256
Hua 252
McShane 246
Howell 243

Thus his average opposition is 254. A plus 3 result (5/7) will perform as 275 and a plus 2 as 268. This is consistent with the FIDE result. If we use +25 for a win, then firstly I have to adjust the field to include his 269, so the field becomes 261.5 and his performance at plus 3 is 272 and plus 2 is 268.6. This is more or less consistent.

Let's see what happens if he destroys the field with say 6.5/7 ( plus 6). As an ECF performance score this is (254*7 +300)/7 = 297. As a plus 25 score this is (261.5*7 +150) /7 = 283. This is consistent with the lag inherent in the +25 approach.

(edited to clarify that where players score their rating in a plus 25 system it's gives the same results as the ECF system)

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Mon Dec 07, 2009 1:29 pm

Hello Roger,

Conversion formulae:

Code: Select all

ECF x 8 + 650 = FIDE
(FIDE - 650) / 8 = ECF 
The players:
Carlsen (269) 2801
Kramnik (265) 2772
Nakamura (258) 2715
Short (257) 2707
Adams (256) 2698
Hua (252) 2665
McShane (246) 2615
Howell (243) 2597

Let us assume that Carlsen played four rounds in a row against his opposition and that he drew all of the (total of 28) games.

On can use online FIDE rating calculator at... http://ratings.fide.com/calculator_rtd.phtml ...to calculate Carlsen's rating after the 28 games (I've chosen 'K' factor of 15 and did not change ratings of Carlsen's opponents after every game they drawn): 2801 - 4*0.6 - 4*1.8 - 4*1.95 - 4*2.1 - 4*2.7 - 4*3.6 - 4*3.9 = 2734.4 (260.55).

After the 28 draws GS would assign to Carlsen a new grade of 254 and AGS3 a new grade of 261.5.

I know that according to FIDE calculator at... http://ratings.fide.com/calculator_rtd.phtml ... if Carlsen (playing more than 28 games) kept drawing against his opposition (if we do not change ratings of Carlsen's opponents after every game) Carlsen would kept losing points, eventually in the limit reaching the average rating of his opposition.

Nevertheless, if we updated the ratings of Kramnik, Nakamura, etc. after every draw against Carlsen (assuming that in other games Kramnik, Nakamura, etc. neither gained nor lost grading points on average), one would find that Carslen's rating would be: 2801 - 0.6 - 1.8 - 1.95 - 2.1 - 2.7 - 3.6 - 3.9 - 0.3 - 1.35 - 1.5 - 1.8 - 2.4 - 3.3 - 3.6 - 0 - 1.05 - 1.2 - 1.5 - 2.1 - 3 - 3.3 + 0.3 - 0.75 - 0.9 - 1.2 - 1.8 - 2.7 - 3 = 2747.9 (262.23), and if Carlsen kept drawing against his opposition (playing more than 28 games) one could expect that in the limit his rating would be around 2742 (261.5), i.e., around AGS3's prediction.

The point is that FIDE system is capable of telling that either both Carlsen worsened and the opposition improved, or that only Carlsen worsened, etc., because one can calculate live ratings and use live ratings in rating of each game.

ECF does not know if, say, in the fourth encounter between Carlsen and Howell, Howell already drew three times against Carlsen and that Howell is a bit stronger player in that game than what his rating of 2597 would suggest. That is why I think it is wrong to assume that if Calrsen drew against Howell in that fourth game he must have played at the Calrsen's level of 2801 and that Calrsen must have played at the Howell's level of 2597 (in fact for me assuming that they have not played at the same level is a logical mistake). Lacking live ratings or other specific info the best bet should have been to assume that both players played at the average level of 2699 (though for me two other logically possible extreme cases are to assume that both either played at 2801 or 2597 level).

Note that though ECF does adjustments on the basis of individual games it is assumed that a minimum of 30 games has to be taken into account in order to change one's grade, so in the case of Carlsen's fourth game against Howell according to AGS3 Carlsen would be assigned 2699 rating points and he would actually lose for that game (2801 - 2699)/30 = 3.4 rating points, which is more or less in accord with FIDE's adjustment (FIDE penalizes Carlsen for his draws against Howell for (3.9 + 3.6 + 3.3 + 3)/4 = 3.45 on average ). According to GS (current ECF grading system) Carlsen would be assigned 2597 rating points and he would actually lose for that game (2801 - 2597)/30 = 6.8 rating points, which is roughly double than FIDE's adjustment.

Nothing much would change if Calrsen did play 28 different opponents instead (playing Howell in the last game), as (as the time passes) some of those opponents may have improved or worsened and once you are faced with the problem of rating Carlsen's 28th game against Howell, unless you are using live ratings or have some other specific info, you should penalize Carlsen for his draw against Howell for approximately (2801 - 2699)/30 = 3.4 rather than (2801 - 2597)/30 = 6.8 rating points, as the best bet should have been to assume that both players played at the average level of 2699.

Kind regards,
Last edited by Robert Jurjevic on Mon Dec 07, 2009 2:49 pm, edited 10 times in total.
Robert Jurjevic
Vafra

Alex Holowczak
Posts: 9309
Joined: Sat May 30, 2009 5:18 pm
Location: Oldbury, Worcestershire
Contact:

Re: GRADING ANOMALIES

Post by Alex Holowczak » Mon Dec 07, 2009 1:37 pm

Robert Jurjevic wrote: Let us assume that Carlsen played four rounds in a row against his opposition and that he drew all of the (total of 32) games.
28 games - Carlsen can't play himself.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Mon Dec 07, 2009 1:49 pm

Alex Holowczak wrote:
Robert Jurjevic wrote: Let us assume that Carlsen played four rounds in a row against his opposition and that he drew all of the (total of 32) games.
28 games - Carlsen can't play himself.
Sure, thanks, I've updated the post now.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22608
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Mon Dec 07, 2009 2:56 pm

Robert Jurjevic wrote:On can use online FIDE rating calculator at... http://ratings.fide.com/calculator_rtd.phtml ...to calculate Carlsen's rating after the 32 games (I've chosen 'K' factor of 15 and did not change ratings of Carlsen's opponents after every game they drawn): 2801 - 4*0.6 - 4*1.8 - 4*1.95 - 4*2.1 - 4*2.7 - 4*3.6 - 4*3.9 = 2734.4 (260.55).

After the 32 draws GS would assign to Carlsen a new grade of 254 and AGS3 a new grade of 261.5.
Actually K=10 would apply to Carlsen and all his opponents.

I think you are finally beginning to see the point. Both the FIDE system and your proposals are systems which lag behind the ECF system is terms of how rapidly they change a player's rating to react to new information.
Robert Jurjevic wrote:The point is that FIDE system is capable of telling that either both Carlsen worsened and the opposition improved, or that only Carlsen worsened, etc., because one can calculate live ratings and use live ratings in rating of each game.
I think you mean Elo systems if you are talking game by game update as used on the Chess Servers. The FIDE ratings are only updated every two months, so if you play 4 games against the same person inside 2 months, each game has the same calculation.

If the London Classic was played every two months and Carlsen drew every game then in each rating list he would decline in rating and his opponents would increase. Over time and rating periods they would meet in the middle on the International System.

There's a hidden assumption in the ECF system (and the international one) that player strength doesn't change during a rating period - rather it leaps or crashes at the end of the measurement period. So if it says 175 in the most recent grading list, you presume an average playing strength of 175 for the whole of the rating period regardless of individual results. Most of the established rating theory asserts that player strength is a random variable - in other words it fluctuates from one game to the next and the published grade or rating is an attempt to pin down the mean. The "Professional Chess Ratings" series which I don't think is published any more attempted to extend this by adding a measure of variability. In other words how "likely" a player was to have an extremely good or extremely bad result.

I go back to a point I made many, many months ago - If you want a system which "averages" a rating using the results of a game, then just use the tried and tested Elo approach rather than a dubious offshoot of the ECF system.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Mon Dec 07, 2009 3:25 pm

Hello Roger,
Roger de Coverly wrote:I think you are finally beginning to see the point. Both the FIDE system and your proposals are systems which lag behind the ECF system is terms of how rapidly they change a player's rating to react to new information.
;)
What makes you think that I haven't seen the point a long time ago (I was always talking about the rules for grading individual games and 'my' rules for grading individual games are even more general than FIDE's, as I have two factors 'k', each for one player, which could depend on ratings trust, while FIDE assumes that the two factors are in most cases equal, 'my' rules could also be adapted to be used in a live system).

The AGS3 rule for live grading system may read:
Rule AGS3 live: For a win you score average grade plus 25; for a draw, average grade; and for a loss, average grade minus 25. Average grade is half of the sum of your and your opponent's grade. Note that, if your opponent's grade differs from yours by more than 40 points, it is taken to be exactly 40 points above (or below) yours. Your new grade after each game is your grade plus difference between score and your grade divided by 30.

(note that a live system would cope with the "junior problem" better than the current one, let us assume that an improving junior played 60 games in a season performing better than expected, if the junior played well enough a live system could have increased the junior's grade for a larger amount than what would the current system do, that is because a live system changes one's grade for a full amount after every 30 games, while the current system can't change one's grade for more than a full amount once a season, regardless of how many games a player did play in the season)

Roger, please could you let me know if you would agree with the following:
(for the example in my previous post where Carlsen draws all of his 28 games)

1) FIDE penalizes Carlsen for his draws against Howell for (3.9 + 3.6 + 3.3 + 3)/4 = 3.45 rating points in each game on average (one corrects players ratings after each draw)

2) GS (current ECF grading system) would penalize Carlsen for his draws against Howell for (2801 - 2597)/30 = 6.8 rating points in each game

3) AGS3 (amended ECF grading system tree) would penalize Carlsen for his draws against Howell for (2801 - 2699)/30 = 3.4 rating points in each game

4) the AGS3's correction of 3.4 rating points is much closer to the FIDE's average correction of 3.45 than the GS's correction of 6.8

5) if ECF would wish to make rating corrections of the order of magnitude as FIDE they should abandon GS and adopt (at lest) AGS3

Kind regards,
Robert Jurjevic
Vafra

Post Reply