GRADING ANOMALIES

General discussions about ratings.
Brian Valentine
Posts: 626
Joined: Fri Apr 03, 2009 1:30 pm

Re: GRADING ANOMALIES

Post by Brian Valentine » Tue Jun 09, 2009 6:56 pm

Alex Holowczak wrote:In my opinion, a less rigid way of doing it would be to use a system similar to that used in Go, but Chess already has it to an extent. Why not have the top graded players graded 1, then the next down graded 2 etc. down to about 20. Everyone fits on to this list, beginners (no graded games) start at grade b (for beginner). Then, you have to achieve a series of norms to progress up the system to get to the next grade. English GMs would be graded 1, IMs 2 and FMs 3, then the rest would hopefully fall into place behind them. At tournaments, the sections wouldn't be U130 or something, they'd be U10. You could jump up as many tiers at once as you like, too, so long as you meet the norms. Since you couldn't lose your grade, having low graded players wouldn't annoy people in higher sections (I read elsewhere on here that that was an issue with some people).
I don't know the Go system but it sounds like the Bridge system. I understand the Bridge players are considering a chess style system! The norms system tends to suffer from inflation as more people get norms. To review the inflation syatem you might need a rating system underneath.

Looking at the "improving adults you might want to look at the 200up rule in Scotland: http://www.chessscotland.com/grading/gradeexp.htm However the number of amendments suggest this idea is not without problems. It might also cross Robert's stretching and equality rules

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Jun 09, 2009 7:05 pm

Hello Brian, thanks a lot for reading through my (not so short) post. :)
Brian Valentine wrote:1. There is no evidence that stretching is taking place. There is evidence that high graded players tend to under perform and lower graded players outperform. This is a weakness in all grading systems and may even have been constant over a long period, if not since inception. Hence I think you must define precisely what YOU mean by stretching. I think your mathematical proofs hide the real problems in player universes, rather than a simple match situation, but I can't be sure.

Mathematically, a system stretches the grades if 'ka + kb > 1', and since for current grading system it holds that 'ka + kb = 2', it stretches the grades for '|p - q|' (which is "a lot"). Non-mathematically, evidence of stretching is shown in the example you have mentioned in point 3 below, where two players graded 100 play 30 games during a course of a season and one of the players scores 80% (current grading system stretches the grades in that example for 30 grading points, i.e. for '|p - q| = |50 - 20| = 30').
Brian Valentine wrote:2. It would help if you could append summaries of all your alternative grading system.
They are summarized in the Replacing GS section...

Code: Select all

--------------------------------------------------------------
grading   stretches  uses FIDE   changes less     preserves 
system    grades     'p = f(d)'  trusted grades   total system
          ('k')      (yellow)    more rapidly     grade
--------------------------------------------------------------
GS        yes        no          no               yes
AGS3      no         no          no               yes 
ÉGS5      no         yes         no               yes
ÉGS6      no         yes         yes              no
--------------------------------------------------------------
Table 1: Main differences between GS (current Grading System), AGS3 (Amended Grading System three), ÉGS5 (Élo Grading System five) and ÉGS6 (Élo Grading System six).

and in the The formulae section.
Brian Valentine wrote:3. In one paragraph you state: "Let us assume that two ungraded players both with estimated grade of 100 play a match of 30 games during a course of a season and that one of the players scores 80%. Then, it follows (from the relationships 'p = f(d)') that (at the end of the season) one of the players should be regarded stronger (than the other) for approximately 30 grading points.". This seems to be the foundation of your stretching argument. I think you are mis-representing the graph which is performance against a 100 player (in your world). As you state the ecf system would rate the diffence as 60 points. However if there were only 2 players then you wouldn't need a rating system as score -to- date would be a sound ranking system.
From all 'p = f(d)' (either of ECF linear or any of Élo) it follows that the grade difference of new grades (in the example) should be approximately 30 grading points, current grading system gives the difference of 60 grading points (i.e., 30 grading points of stretching).
Brian Valentine wrote:Looking at the recent thread I think we need a definition for "equal grade for equal performance" since I suspect yours is different from Roger's idea.
Basically, "equal grade for equal performance" requires that say if you have a 130 player who scores 50% against a pool of 160 players that the 130 player becomes a 160 player (according to the systems which neither stretch nor shrink the grades the 130 player becomes approximately a 145 player). Please see the "Equal grade for equal performance" section.

It looks like (some) people cannot abandon "equal grade for equal performance" rule (which is in my opinion unsound). Unfortunately, requiring the rule to hold is equivalent to asking of the system to stretch the grades for '|p - g|' (which is "a lot").
Robert Jurjevic
Vafra

Alex Holowczak
Posts: 9311
Joined: Sat May 30, 2009 5:18 pm
Location: Oldbury, Worcestershire
Contact:

Re: GRADING ANOMALIES

Post by Alex Holowczak » Tue Jun 09, 2009 7:11 pm

Brian Valentine wrote:I don't know the Go system but it sounds like the Bridge system. I understand the Bridge players are considering a chess style system! The norms system tends to suffer from inflation as more people get norms. To review the inflation syatem you might need a rating system underneath.
This should help explain it better than I can. In most top level tournaments though, dan and kyu are used, not the "modern" Elo system.
Brian Valentine wrote:Looking at the "improving adults you might want to look at the 200up rule in Scotland: http://www.chessscotland.com/grading/gradeexp.htm However the number of amendments suggest this idea is not without problems. It might also cross Robert's stretching and equality rules
I think that article shows my concerns. Every attempt to make the system statistically fairer seems to make it much more complicated to calculate.

Roger de Coverly
Posts: 22610
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Jun 09, 2009 7:47 pm

Let us assume that two ungraded players both with estimated grade of 100 play a match of 30 games during a course of a season and that one of the players scores 80%
I would rather you said:-

Let there be two ungraded players. They have unknown strength because they are ungraded. They each play 30 games against an average (graded) field of 100. One scores 80%, the other 20%. What are their new grades?

The answer according to the ECF system ( and for that matter the new player estimation process of Elo systems) is that the 80% guy gets a grade of 130 and the 20% guy gets a grade of 70. If they play each other (once they've established a grade) then the 40 point rule cuts in to distort matters.

They may or may not be a force called "stretch" operating on the ECF grading system. If there is, it's just one of many and I don't think it is particularly powerful. Just look at relative grades towards the top of the list, say 150 plus and tell me whether you see any real evidence that club and county players 150-190 are drifting away from the IMs and GMs 210-250.
simply because they wouldn't be aware of the problem.
The whole "deflation" issue was produced like a rabbit out of a hat about three or four years ago. You'd think if it had been going on for thirty years, that players whose experience spanned the whole era might have noticed and had been agitating for changes. They hadn't. There was certainly an issue that the 8*ECF + 600 = FIDE no longer worked in the 2000-2200 range because players under 175 were getting FIDEs over 2000. This was because of inflationary issues in the FIDE system (only rating best performances) rather than deflation (or stretch) in the ECF.

There's an issue of players with zero and negative grades and the number of players with grades of well below 100. This probably is down to the extension of the grading system to below a playing standard at which you could expect it to work reliably coupled with effects of the new player estimation process, junior improvement rates, junior increments, 40 point rule memory etc, etc. The right way to deal with these issues would have been to hypothetically change parameters and parallel run, then review the results for sensitivity.

One of the main advantages of a system which doesn't average grades between opponents is that it makes it so much easier to compute estimates for unrated players since (40 point rule permitting) you can compute an end period grade (based on graded opposition) without having to make any prior guess of their strength.
Last edited by Roger de Coverly on Tue Jun 09, 2009 8:20 pm, edited 1 time in total.

Brian Valentine
Posts: 626
Joined: Fri Apr 03, 2009 1:30 pm

Re: GRADING ANOMALIES

Post by Brian Valentine » Tue Jun 09, 2009 7:59 pm

Robert Jurjevic wrote:Mathematically, a system stretches the grades if 'ka + kb > 1', and since for current grading system it holds that 'ka + kb = 2', it stretches the grades for '|p - q|' (which is "a lot"). Non-mathematically, evidence of stretching is shown in the example you have mentioned in point 3 below, where two players graded 100 play 30 games during a course of a season and one of the players scores 80% (current grading system stretches the grades in that example for 30 grading points, i.e. for '|p - q| = |50 - 20| = 30').
.

If this is the definition of system stretches it is not relevant to the system stretching issue within the ECF system. In the ECF system the grading difference of 60 points arises if one player scores 80% against another and those are the only games played and any earlier grades or estimates are discarded. Your graph used to get a 30 point difference is the one outlined above: player A scores 80% against players rated the same; and is consequence although not shown on the graph the other player's grades fall. It is wrong to suggest that the right difference for the situation you actually address, where both player's grade changes, is a new difference of 30 points.
Robert Jurjevic wrote:Basically, "equal grade for equal performance" requires that say if you have a 130 player who scores 50% against a pool of 160 players that the 130 player becomes a 160 player (according to the systems which neither stretch nor shrink the grades the 130 player becomes approximately a 145 player). Please see the "Equal grade for equal performance" section.
I think you need to look again at Roger de Coverly's original definition of equal grade for equal work (posted May 20 at 5:26). This is a good definition of a useful , but not central test of a good system. Yours is a caricature of that concept. The k's are used to allow for degrees of belief on all results over a rating period not to control stretch

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Jun 09, 2009 8:46 pm

Brian Valentine wrote:
Robert Jurjevic wrote:Mathematically, a system stretches the grades if 'ka + kb > 1', and since for current grading system it holds that 'ka + kb = 2', it stretches the grades for '|p - q|' (which is "a lot"). Non-mathematically, evidence of stretching is shown in the example you have mentioned in point 3 below, where two players graded 100 play 30 games during a course of a season and one of the players scores 80% (current grading system stretches the grades in that example for 30 grading points, i.e. for '|p - q| = |50 - 20| = 30').
If this is the definition of system stretches it is not relevant to the system stretching issue within the ECF system.
Why not? Mathematical requirement for a grading system (using the mentioned formulae) not to stretch nor shrink the grades is that '(a2 - a) + (b - b2)' is equal to 'q - p' for any 'a', 'b', 'p' and 'q', where 'a' and 'b' are the grades of players 'A' and 'B', 'p' expected performance of player 'A' (expected performance of player 'B' is then '100 - p'), 'q' actual performance of player 'A' (actual performance of player 'B' is then '100 - q') and 'a2' and 'b2' new grades of players 'A' and 'B'.
Brian Valentine wrote:In the ECF system the grading difference of 60 points arises if one player scores 80% against another and those are the only games played and any earlier grades or estimates are discarded. Your graph used to get a 30 point difference is the one outlined above: player A scores 80% against players rated the same; and is consequence although not shown on the graph the other player's grades fall. It is wrong to suggest that the right difference for the situation you actually address, where both player's grade changes, is a new difference of 30 points.
Let us assume that two players both with grade of 100 play a match of 30 games during a course of a season and that one of the players scores 80%. Then, it follows (from the relationships 'p = f(d)') that (at the end of the season) one of the players should be regarded stronger (than the other) for approximately 30 grading points.

According to GS new grades of the players (in the above example) are 130 and 70.

According to ÉGS5 new grades of the players (in the above example) are 115 and 85.

As '115 - 85 = 30' and '130 - 70 = 60' it is obvious that GS stretched the grades for 30 grading points (the new grade difference calculated by GS is twice as big than it should have been, as according to all 'p = f(d)' the difference between new grades should be 30, not 60 grading points)!

Brian Valentine wrote:
Robert Jurjevic wrote:Basically, "equal grade for equal performance" requires that say if you have a 130 player who scores 50% against a pool of 160 players that the 130 player becomes a 160 player (according to the systems which neither stretch nor shrink the grades the 130 player becomes approximately a 145 player). Please see the "Equal grade for equal performance" section.
I think you need to look again at Roger de Coverly's original definition of equal grade for equal work (posted May 20 at 5:26). This is a good definition of a useful , but not central test of a good system. Yours is a caricature of that concept. The k's are used to allow for degrees of belief on all results over a rating period not to control stretch
Roger de Coverly's post you have mentioned says:
Roger de Coverly wrote:A premise of the ECF grading system is equal grade for equal work. So if two players get the same results in the same period against the same average field, they get the same grade provided they've both played more than the 30 game qualification and provided their previous grade is not so far away from their performance that the 40 point cut off comes into play. So over 30 games, a player graded 175 scores 50% against an average field of 150, a player graded 125 does the same. They will both have the same grade at the next cut off being 150. If the performance is only over 15 games, then they still have different grades of 137.5 v 162.5
so his "definition" is basicall equivalent to mine (I have used 30 and Roger 25 grading points differnce in the example). Unfortunately, I'll have to reiterate myself and say:
Robert Jurjevic wrote:It looks that (some) people cannot abandon "equal grade for equal performance" rule (which is in my opinion unsound). Unfortunately, requiring the rule to hold is equivalent to asking of the system to stretch the grades for '|p - q|' (which is "a lot").
My understandng is that Roger also considered a situation where a player may not have played 30 games in the season when the rule may not hold, depending on the player's results in the previous season, but this is irreleavnt for the point I am making.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22610
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Tue Jun 09, 2009 9:28 pm

Let us assume that two players both with grade of 100 play a match of 30 games during a course of a season
Lighthouse keepers ! It's known that the ECF system is suspect in two player universes. The normal example is players of 100 and 150 who only play each other and score 50%. They swop grades in an ECF system over 30 games but converge towards 125 in an Elo one. Actually they also converge even in an ECF system provided they only play 15 games before a new grading list comes out. The 100 player gets (100 * 15 + 150 *15) / 30 and the 150 player gets (150 * 15 + 100*15) /30. In the second half of their match, they're in a new season, so they are both 125. So the reporting frequency and its interaction with the minimum game count is also a parameter.

I wonder how many players there have been in the nearly 60 year lifetime of the ECF grading system who have played multiple games but only against one opponent. Short v Kasparov in the 93-94 grading year perhaps (if they bothered to include it). For all practical purposes please assume that players play a multitude of opponents, not the same one !
equal grade for equal performance" rule (which is in my opinion unsound)
It also applies in an Elo system but only when both players start unrated. It's a point in favour of the ECF system that it applies there to rated players as well and in its context perfectly sound and historically tested. In many practical circumstances, for example local leagues and tournaments you are trying to rank players who mostly face the same opposition. So you rank them by their "score" in this year's competition - you only bring in last year's "score" if they have played fewer games than your qualification cutoff.

Brian Valentine
Posts: 626
Joined: Fri Apr 03, 2009 1:30 pm

Re: GRADING ANOMALIES

Post by Brian Valentine » Tue Jun 09, 2009 9:43 pm

Robert Jurjevic wrote:Let us assume that two players both with grade of 100 play a match of 30 games during a course of a season and that one of the players scores 80%. Then, it follows (from the relationships 'p = f(d)') that (at the end of the season) one of the players should be regarded stronger (than the other) for approximately 30 grading points.

According to GS new grades of the players (in the above example) are 130 and 70.

According to ÉGS5 new grades of the players (in the above example) are 115 and 85.
Robert,
In your weblink this argument is : "Let us assume that two ungraded players both with estimated grade of 100 play a match of 30 games during a course of a season and that one of the players scores 80%. Then, it follows (from the relationships 'p = f(d)') that (at the end of the season) one of the players should be regarded stronger (than the other) for approximately 30 grading points." The Word "estimated" is an important addition (I think you talked about it somewhere earlier). In the argument uncluding "estimated" one is looking at a problem close to the junior/fast improving adult issue. In the argument quoted here, using actual grades (I can't read it any other way), the issue is about the inherent stability of performance (albeit taken to extremes for illustration).

In the latter case the position is that the difference in strength is (as far as can be assessed) 30 points , but that is not the same as the difference in performance. (In my mathematical model submitted earlier strength is G and performance is P). In your example performance is distorted by poor assessment of current strength - the 100 each player is given. If the system works well, and strength is stable, then the identical grades would not have arisen.

Hence I think your example is not helpful in explaining system stretch, but this discussion has lead to me understanding your issue and I may go back to your paper now I have got your point.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Tue Jun 09, 2009 10:29 pm

Brian Valentine wrote:Hence I think your example is not helpful in explaining system stretch, but this discussion has lead to me understanding your issue and I may go back to your paper now I have got your point.
Thanks a lot for showing the interest.

Let 'a' and 'b' are the grades of players 'A' and 'B', 'p' expected performance of player 'A' (expected performance of player 'B' is then '100 - p'), 'q' actual performance of player 'A' (actual performance of player 'B' is then '100 - q') and 'a2' and 'b2' new grades of players 'A' and 'B'.

'a2' and 'b2' are calculated using the following formulae (holds for any grading system mentioned here, including the current one):

'a2 = a + ka*(q - p)'
'b2 = b + kb*((100 - q) - (100 - p))'

then I examine a term '(a2 - a) + (b - b2)' (which is a measure of how much the grades drift apart due to a difference in actual and expected performance 'q - p').

I then require that it should hold that '(a2 - a) + (b - b2) = 'q - p' for any 'a', 'b', 'p' and 'q'. If '(a2 - a) + (b - b2) > 'q - p' then the system stretches the grades, and if '(a2 - a) + (b - b2) < 'q - p' then the system shrinks the grades.

It can be shown that if 'ka + kb = 2' then '(a2 - a) + (b - b2) = 2*(q - p)' and if 'ka + kb = 1' then '(a2 - a) + (b - b2) = q - p'.

For current grading system it holds that '(a2 - a) + (b - b2) = 2*(q - p)'.

Thanks a lot.
Robert Jurjevic
Vafra

E Michael White
Posts: 1420
Joined: Fri Jun 01, 2007 6:31 pm

Re: GRADING ANOMALIES

Post by E Michael White » Wed Jun 10, 2009 2:22 pm

Roger de Coverly wrote:
A N Other wrote:This means that whereas for a grading difference of 25 points the stronger player should score 75%, his actual score is more like 68%.
I don't know whether I've ever seen a satisfactory explanation of how this was measured. But assuming it to be true, why doesn't this act to counter stretch?
One reason, I guess, is that there are fewer games played between 25 points difference opponents than against other pairings which stretch/spread the system.

A pattern of games played which wont stretch but may spread the system is where everyone plays everyone else in an all play all.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Wed Jun 10, 2009 2:29 pm

I've got it! :-) :-) [new document version 10/06/2009 1.12 should clarify my point, sections "Factors 'k' and grade stretching", "Equal grade for equal performance"" and "Mathematical proof" were updated]

http://www.ecforum.org.uk/viewtopic.php ... 149#p10149

http://www.jurjevic.org.uk/chess/grade/ ... malies.htm

My point in a nutshell is that the 'k' factors (wrongly chosen) in the present grading system are causing the grade stretching and that the amount of stretch (due to the 'k' factors) is larger (in fact it is equal to '|p - q|') than grade fluctuations caused by other anomalies which may be corrected by using FIDE logistic relation for 'p = f(d)', Glickman idea on changing less trusted grades (based on frequency of play) faster than more trusted grades, or even a solution to the "junior problem".
Roger de Coverly wrote:
Let us assume that two players both with grade of 100 play a match of 30 games during a course of a season
Lighthouse keepers ! It's known that the ECF system is suspect in two player universes. The normal example is players of 100 and 150 who only play each other and score 50%. They swop grades in an ECF system over 30 games but converge towards 125 in an Elo one. Actually they also converge even in an ECF system provided they only play 15 games before a new grading list comes out. The 100 player gets (100 * 15 + 150 *15) / 30 and the 150 player gets (150 * 15 + 100*15) /30. In the second half of their match, they're in a new season, so they are both 125. So the reporting frequency and its interaction with the minimum game count is also a parameter.

I wonder how many players there have been in the nearly 60 year lifetime of the ECF grading system who have played multiple games but only against one opponent. Short v Kasparov in the 93-94 grading year perhaps (if they bothered to include it). For all practical purposes please assume that players play a multitude of opponents, not the same one!
Let us assume that two pool of players both with average grade of 100 play each other during a course of a season and that one of the player pools scores 80%. Let us assume that each player in the pool playes only players of the other pool and that each player played exactly 30 games (for simplicity we can assume that in each pool there are 30 players each graded 100 and that each player from one pool plays each player from other pool, totaling in 900 games). Then, it follows (from the relationships 'p = f(d)') that at the end of the season one of the player pools should be regarded stronger (than the other) for approximately 30 grading points (because it scored 30% more).

(Note that it is unlikely that one of the pools would score so high in practice, though in order to please those who might be troubled with that, we could assume that, say, players of the well performing pool are all juniors who had been lucky enough to be coached by Garry Kasparov in the summer break before the start of the season.)

According to GS new grades of the player pools in the above example are 130 and 70 (the pool grades drift apart for '130 - 70 = 60' grading points).

According to ÉGS new grades of the player pools in the above example are 115 and 85 (the pool grades drift apart for '115 - 85 = 30' grading points)..

As the grade drifts are '115 - 85 = 30' and '130 - 70 = 60' it is obvious that (current grading system) GS stretched the grades for 30 grading points (the new grade difference calculated by GS is twice as big than it should have been)!

(You see how ÉGS is fair, it did not assign grades of 130 and 100, as it did not assume that the better pool improved and the other stayed as it was, but it guessed that the result was due to both one of the pools improving and other worsening, though if Kasparov really did coach the juniors in the better pool, the grades of 130 and 100 would be a better guess. GS grades of 130 and 70 makes no sense at all, as if the better pool was given 130 the other pool should have been given 100, not 70, giving 70 to other pool is as if the better pool had scored approximately 94%, according to Élo's logistic 'p = f(d)'.)

(Note that ÉGS2 would assign grades of 130 and 100 if all of the players in the better pool were ungraded and if all of the players in the other pool were graded, that is because ÉGS2 changes less trusted grades more rapidly than more trusted grades, and in this extreme case the grades of graded players remain unaffected by games played against ungraded players. Well, it would be nice if we could take into account if, say, Kasparov was coaching a player, but...)

In my opinion, the main reason for the grade stretching is factor 'k' which is twice as big in GS than ÉGS and ÉGS2 (please note that in the above example we eliminated the differences in 'p = f(d)', so 'p = f(d)' couldn't be a cause of the stretching).

Roger de Coverly wrote:
equal grade for equal performance" rule (which is in my opinion unsound)
It also applies in an Elo system but only when both players start unrated. It's a point in favour of the ECF system that it applies there to rated players as well and in its context perfectly sound and historically tested. In many practical circumstances, for example local leagues and tournaments you are trying to rank players who mostly face the same opposition. So you rank them by their "score" in this year's competition - you only bring in last year's "score" if they have played fewer games than your qualification cutoff.
Systems which neither stretch nor shrink the grades do not obey the rule which is known as "equal grade for equal performance", say if you have a 130 player who scores 50% against a pool of 160 players "equal grade for equal performance" rule requires that the 130 player becomes a 160 player (according to the systems which neither stretch nor shrink the grades the 130 player becomes approximately a 145 player).

So it would seem that one could opt either for a system which obeys "equal grade for equal performance" rule and stretches the grades or a system which does not obey "equal grade for equal performance" rule and neither stretches nor shrinks the grades.

Let us assume that in the above "equal grade for equal performance" example the 130 player played 300 games during a course of a season and that each player in the pool (there are 10 players in the pool) played 30 games against the 130 player. Then, taking into account only the 300 games the pool players played against the 130 player, it follows (from the relationships 'p = f(d)') that at the end of the season the 130 player should be regarded approximately equally strong as the pool of players he played.

"Equal grade for equal performance" rule requires that new grade of the 130 player is 160 (assuming that the pool grade stays approximately 160).

Taking into account only the 300 games the pool players played against the 130 player, a system which neither stretches nor shrinks the grades requires that new grade of both the 130 player and the pool is approximately 145.

If really all the pool payers performed at their level of 160 and the 130 player did improve, the 130 player should become the 160 and the pool players should remain 160. The problem with the current grading system is that even it assigns 160 to the 130 player it panelizes the pool players for the games which they have played against the 130 players, what is causing the grade stretching.

A system which wouldn't stretch the grades, if assigning to the 130 player a grade of 160, when calculating the grade of the pool players, should ignore the games the pool players have played against the 130 player (as it is already assumed that they performed on 160 level and the games they have played against the 130 player should have no effect on their grade), or it can assume that both the pool players worsened and the 130 player improved (splinting it 50/50) and assigning to the 130 player a grade of 145, penalizing the pool players for their games against the 30 player (i.e., if the pool players played only the games against the 130 player the pool grade would lower to 145) not stretching the grades.

The above argument is enough for me to claim that "equal grade for equal performance" rule is unsound and should be abandoned in favour of a system which neither stretches nor shrinks the grades.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22610
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Wed Jun 10, 2009 2:59 pm

If really all the pool payers performed at their level of 160 and the 130 player did improve, the 130 player should become the 160 and the pool players should remain 160. The problem with the current grading system is that even it assigns 160 to the 130 player it panelizes the pool players for the games which they have played against the 130 players, what is causing the grade stretching.
You need to carefully look at this in the historic context of the grading system. It's been known for many years that if a player improves from 130 to 160 that there is a danger that this player will dilute the 160 pool a bit - maybe gaining his 30 points by taking 1 point of each of them. This is deflation surely, not stretch. The 160 pool would have dropped a point down to 159,

In practice various other forces come into play, some natural, some artificial. One of them is players not playing to their previous standard - for example a 190 player only playing to a 160 standard. Another is presuming that juniors are improving players and awarding a junior increment. Yet another is introducing a bit of inflation by over-estimating the start grades of new players. At the top end, the 40 point rule increases the rewards to the top players. At the bottom end it acts as a drag on improving beginners.

As far as I am concerned, the bottom line on stretch and deflation is the published grades on the wall chart. Amongst players that I might expect to score about 50%, I don't think there has been any real movement in the grades over the years. Grading is a complex system, I do not think you can model it by plucking formulae from the air based on unrealistic examples.

I've said it before and I will say it again. If you want a system that contains rating lag, then just use the tried and tested Elo methodology. If you want to measure performance over a 30 game season with new players treated the same way as established ones, then use the current ECF system. Your "improvements" just add rating lag to the ECF system. You would get similar "lag" effects by increasing the 30 game minimum for averaging or by increasing the frequency of publication.

By way of example, lets assume we publish a new rating list half way through the season in which our 130 player is performing at 160. If he's only played 15 games and our qualification standard remains 30 games, then his new rating is 145 which is your averaging effect.

Brian Valentine
Posts: 626
Joined: Fri Apr 03, 2009 1:30 pm

Re: GRADING ANOMALIES

Post by Brian Valentine » Wed Jun 10, 2009 3:19 pm

Robert,
You are barking up the wrong tree on this. For one, your unrealistic assumption of a two player match hides the fact that normally the factor is 1/n where is based on games included in the next list. For two, the system is trying to rank players as though their strength is not changing much. In the situation you give you are trying to address the improving junior issue. At present this is addressed by actually adding points, but in the elo system it is met by a larger k. You should listen to Roger's patient remarks carefully.

User avatar
Robert Jurjevic
Posts: 207
Joined: Wed May 16, 2007 1:31 pm
Location: Surrey

Re: GRADING ANOMALIES

Post by Robert Jurjevic » Wed Jun 10, 2009 3:19 pm

Roger de Coverly, I am sorry and sad that you can't or won't see it, the key point in the "equal grade for equal performance" example is ...The problem with the current grading system is that even it assigns 160 to the 130 player it panelizes the pool players for the games which they have played against the 130 players ...in...

If really all the pool payers performed at their level of 160 and the 130 player did improve, the 130 player should become the 160 and the pool players should remain 160. The problem with the current grading system is that even it assigns 160 to the 130 player it panelizes the pool players for the games which they have played against the 130 players, what is causing the grade stretching.

Are you saving that if I was one of the 160 pool payers and I drew against the 130 player that you would not take that game into account when calculating my grade, as I was a member of the 160 pool and the grade of the 130 player was already corrected as much as he was playing against a player (me) who was performing exactly as a 160 player? If you are taking that game into account when calculating my grade (slightly lowering my grade) you are stretching the grades (if you want to lower my grade you can but then you should not rise the 130 player grade for 30 points but less, the sum of the grade corrections must be 30 in order not to stretch nor shrink the grades).

But, the main thing I would like to address is...

My point in a nutshell is that the 'k' factors (wrongly chosen) in the present grading system are causing the grade stretching and that the amount of stretch (due to the 'k' factors) is larger (in fact it is equal to '|p - q|') than grade fluctuations caused by other anomalies which may be corrected by using FIDE logistic relation for 'p = f(d)', Glickman idea on changing less trusted grades (based on frequency of play) faster than more trusted grades, or even a solution to the "junior problem".

Thanks a lot.
Robert Jurjevic
Vafra

Roger de Coverly
Posts: 22610
Joined: Tue Apr 15, 2008 2:51 pm

Re: GRADING ANOMALIES

Post by Roger de Coverly » Wed Jun 10, 2009 4:21 pm

If really all the pool payers performed at their level of 160 and the 130 player did improve, the 130 player should become the 160 and the pool players should remain 160. The problem with the current grading system is that even it assigns 160 to the 130 player it panelizes the pool players for the games which they have played against the 130 players, what is causing the grade stretching.
As I said before, this is an effect which has been known about for around 40 years and in effect mechanisms exist which compensate for it. Back in the 1960s, they only rated about the top 10% and it became apparent that the numbers of players in the top grades 1a 248-241 etc. was diminishing. In the event at least two changes seem to have combined to add back the lost points - firstly a junior increment to "feed" points into the system and secondly the 40 point rule which "over rewards" the really top players if they play opposition well below their standard. I suspect there was (unofficially) also a minimum grade of 100 used as a first estimate for new players.

I would also say that diluting the 160 pool to 159 is deflation not stretch because exactly the same effect will apply between the FMs at 200 and the GMs at 230. So the distance between GMs and "160s" remains unchanged but the whole distribution moves down a point.

I would further say that there are other ways of getting from 130 to 160 than scoring 50% against the 160s. You could also score much better than 50% against the 130s thereby diluting them down to 129. You could do much better than 75% against 105s.

We are talking about a model in which most players do not change standard much and trying to reduce rating lag for those that do.

Post Reply