I've got it
!
[new document version 10/06/2009 1.12 should clarify my point, sections "Factors 'k' and grade stretching", "Equal grade for equal performance"" and "Mathematical proof" were updated]
http://www.ecforum.org.uk/viewtopic.php ... 149#p10149
http://www.jurjevic.org.uk/chess/grade/ ... malies.htm
My point in a nutshell is that the 'k' factors (wrongly chosen) in the present grading system are causing the grade stretching and that the amount of stretch (due to the 'k' factors) is larger (in fact it is equal to '|p - q|') than grade fluctuations caused by other anomalies which may be corrected by using FIDE logistic relation for 'p = f(d)', Glickman idea on changing less trusted grades (based on frequency of play) faster than more trusted grades, or even a solution to the "junior problem".
Roger de Coverly wrote:Let us assume that two players both with grade of 100 play a match of 30 games during a course of a season
Lighthouse keepers ! It's known that the ECF system is suspect in two player universes. The normal example is players of 100 and 150 who only play each other and score 50%. They swop grades in an ECF system over 30 games but converge towards 125 in an Elo one. Actually they also converge even in an ECF system provided they only play 15 games before a new grading list comes out. The 100 player gets (100 * 15 + 150 *15) / 30 and the 150 player gets (150 * 15 + 100*15) /30. In the second half of their match, they're in a new season, so they are both 125. So the reporting frequency and its interaction with the minimum game count is also a parameter.
I wonder how many players there have been in the nearly 60 year lifetime of the ECF grading system who have played multiple games but only against one opponent. Short v Kasparov in the 93-94 grading year perhaps (if they bothered to include it). For all practical purposes please assume that players play a multitude of opponents, not the same one!
Let us assume that two pool of players both with average grade of 100 play each other during a course of a season and that one of the player pools scores 80%. Let us assume that each player in the pool playes only players of the other pool and that each player played exactly 30 games (for simplicity we can assume that in each pool there are 30 players each graded 100 and that each player from one pool plays each player from other pool, totaling in 900 games). Then, it follows (from the relationships 'p = f(d)') that at the end of the season one of the player pools should be regarded stronger (than the other) for approximately 30 grading points (because it scored 30% more).
(Note that it is unlikely that one of the pools would score so high in practice, though in order to please those who might be troubled with that, we could assume that, say, players of the well performing pool are all juniors who had been lucky enough to be coached by Garry Kasparov in the summer break before the start of the season.)
According to GS new grades of the player pools in the above example are 130 and 70 (the pool grades drift apart for '130 - 70 = 60' grading points).
According to ÉGS new grades of the player pools in the above example are 115 and 85 (the pool grades drift apart for '115 - 85 = 30' grading points)..
As the grade drifts are '115 - 85 = 30' and '130 - 70 = 60' it is obvious that (current grading system) GS
stretched the grades for
30 grading points (the new grade difference calculated by GS is twice as big than it should have been)
!
(You see how ÉGS is fair, it did not assign grades of 130 and 100, as it did not assume that the better pool improved and the other stayed as it was, but it guessed that the result was due to both one of the pools improving and other worsening, though if Kasparov really did coach the juniors in the better pool, the grades of 130 and 100 would be a better guess. GS grades of 130 and 70 makes no sense at all, as if the better pool was given 130 the other pool should have been given 100, not 70, giving 70 to other pool is as if the better pool had scored approximately 94%, according to Élo's logistic 'p = f(d)'.)
(Note that ÉGS2 would assign grades of 130 and 100 if all of the players in the better pool were ungraded and if all of the players in the other pool were graded, that is because ÉGS2 changes less trusted grades more rapidly than more trusted grades, and in this extreme case the grades of graded players remain unaffected by games played against ungraded players. Well, it would be nice if we could take into account if, say, Kasparov was coaching a player, but...)
In my opinion, the main reason for the grade stretching is factor 'k' which is twice as big in GS than ÉGS and ÉGS2 (please note that in the above example we eliminated the differences in 'p = f(d)', so 'p = f(d)' couldn't be a cause of the stretching).
Roger de Coverly wrote:equal grade for equal performance" rule (which is in my opinion unsound)
It also applies in an Elo system but only when both players start unrated. It's a point in favour of the ECF system that it applies there to rated players as well and in its context perfectly sound and historically tested. In many practical circumstances, for example local leagues and tournaments you are trying to rank players who mostly face the same opposition. So you rank them by their "score" in this year's competition - you only bring in last year's "score" if they have played fewer games than your qualification cutoff.
Systems which neither stretch nor shrink the grades do not obey the rule which is known as "equal grade for equal performance", say if you have a 130 player who scores 50% against a pool of 160 players "equal grade for equal performance" rule requires that the 130 player becomes a 160 player (according to the systems which neither stretch nor shrink the grades the 130 player becomes approximately a 145 player).
So it would seem that one could opt either for a system which obeys "equal grade for equal performance" rule and stretches the grades or a system which does not obey "equal grade for equal performance" rule and neither stretches nor shrinks the grades.
Let us assume that in the above "equal grade for equal performance" example the 130 player played 300 games during a course of a season and that each player in the pool (there are 10 players in the pool) played 30 games against the 130 player. Then, taking into account only the 300 games the pool players played against the 130 player, it follows (from the relationships 'p = f(d)') that at the end of the season the 130 player should be regarded approximately equally strong as the pool of players he played.
"Equal grade for equal performance" rule requires that new grade of the 130 player is 160 (assuming that the pool grade stays approximately 160).
Taking into account only the 300 games the pool players played against the 130 player, a system which neither stretches nor shrinks the grades requires that new grade of both the 130 player and the pool is approximately 145.
If really all the pool payers performed at their level of 160 and the 130 player did improve, the 130 player should become the 160 and the pool players should remain 160. The problem with the current grading system is that even it assigns 160 to the 130 player it panelizes the pool players for the games which they have played against the 130 players, what is causing the grade stretching.
A system which wouldn't stretch the grades, if assigning to the 130 player a grade of 160, when calculating the grade of the pool players, should ignore the games the pool players have played against the 130 player (as it is already assumed that they performed on 160 level and the games they have played against the 130 player should have no effect on their grade), or it can assume that both the pool players worsened and the 130 player improved (splinting it 50/50) and assigning to the 130 player a grade of 145, penalizing the pool players for their games against the 30 player (i.e., if the pool players played only the games against the 130 player the pool grade would lower to 145) not stretching the grades.
The above argument is enough for me to claim that "equal grade for equal performance" rule is unsound and should be abandoned in favour of a system which neither stretches nor shrinks the grades.