Chess Player Strip Searched

The very latest International round up of English news.
User avatar
Paolo Casaschi
Posts: 1194
Joined: Thu Jan 08, 2009 6:46 am

Re: Chess Player Strip Searched

Post by Paolo Casaschi » Sun May 18, 2014 9:34 pm

NickFaulks wrote:It does say that, but I don't believe that is really Ken Regan's intention, and it's his baby.
Next time you talk with him you should double-check about this. Few years ago, in a paper still available online, he made the case to reject the possibility of convicting a cheater based solely on statistical evidence. Recently however he changed his mind, with a lengthy dissertation about the reasons for this turnaround and coming up with the one-in-a-million threshold. Following this trend in his papers, his mind seems going in the direction stated clearly in the proposal.
As a confirmation, the paper allows the ACC to start an investigation after an event is concluded even without any formal complain raised by other players or arbiters. How do they intend to collect evidence in such a situation, unless the only evidence they need is statistical analysis of the games?

In my opinion, the worse solution would be leaving everything to the discretion of the ACC commission, for them to determine if the next Ivanov should be banned for life and if the game of a super GM is the occasional false positive. Comparing to a very different situation, I always wondered what would have happened if a player from, let's say, Bermuda, rather than Ivanchuk, had refused to provide an anti-doping sample at the chess Olympics few years ago.

David Sedgwick
Posts: 5252
Joined: Mon Apr 09, 2007 5:56 pm
Location: Croydon
Contact:

Re: Chess Player Strip Searched

Post by David Sedgwick » Sun May 18, 2014 10:39 pm

Paolo Casaschi wrote:Comparing to a very different situation, I always wondered what would have happened if a player from, let's say, Bermuda, rather than Ivanchuk, had refused to provide an anti-doping sample at the chess Olympics few years ago.
Substitute Papua New Guinea for Bermuda and I can tell you the answer. He goes on to become Secretary of the FIDE Anti-Cheating Commission.

I should make clear that I'm quite satisfied that Shaun Press had not taken any prohibited substances.

There were enough question marks about that quasi-judicial process to lead me to believe that Sean will understand better than most the need for proper procedures with the one now under discussion.

NickFaulks
Posts: 9255
Joined: Sat Jan 02, 2010 1:28 pm

Re: Chess Player Strip Searched

Post by NickFaulks » Sun May 18, 2014 10:51 pm

Paolo Casaschi wrote: Comparing to a very different situation, I always wondered what would have happened if a player from, let's say, Bermuda, rather than Ivanchuk, had refused to provide an anti-doping sample at the chess Olympics few years ago.
Funny you should say that.

http://en.wikipedia.org/wiki/36th_Chess ... ug_testing

"Drug testing

Having been formally recognized by the International Olympic Committee in 1999, in preparation for prospective inclusion in future iterations of the Olympic Games, FIDE implemented (in 2001) doping restrictions consistent with those adopted by the World Anti-Doping Agency (WADA). Two players, Shaun Press of Papua New Guinea and Bobby Miller of Bermuda, refused, for various reasons, to submit urine samples for analysis. Both players appeared before a FIDE disciplinary panel, which decided to cancel the players' performances (Press had scored 7½ points in 14 games, while Miller had scored 3½ points in 9 games), reducing the final score of Papua New Guinea to 15½ (from 23) and that of Bermuda to 18½ (from 22)."

Both players received full and enthusiastic support from their federations. There was no effect on the medals table, but it did mean that we finished above PNG despite Rupert Jones mysteriously scoring 10/13 and getting an FM title.

The players received warnings. Thanks for this reasonable treatment were due to the very sensible Jana Bellin, and it also didn't hurt that Jon Speelman, who was bemused by the whole process, was on the disciplinary panel.
If you want a picture of the future, imagine a QR code stamped on a human face — forever.

User avatar
Paolo Casaschi
Posts: 1194
Joined: Thu Jan 08, 2009 6:46 am

Re: Chess Player Strip Searched

Post by Paolo Casaschi » Sun May 18, 2014 10:58 pm

NickFaulks wrote:
Paolo Casaschi wrote: Comparing to a very different situation, I always wondered what would have happened if a player from, let's say, Bermuda, rather than Ivanchuk, had refused to provide an anti-doping sample at the chess Olympics few years ago.
Funny you should say that.

http://en.wikipedia.org/wiki/36th_Chess ... ug_testing
I most likely had heard of something like that but completely forgot the details. The Ivanchuk incident is much easier to remember. It appears that my memory also applies the kind of double standards I would not want from any ACC committee. :oops:

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: Chess Player Strip Searched

Post by Roger de Coverly » Sun May 18, 2014 11:28 pm

Paolo Casaschi wrote: Comparing to a very different situation, I always wondered what would have happened if a player from, let's say, Bermuda, rather than Ivanchuk, had refused to provide an anti-doping sample at the chess Olympics few years ago.
It did actually happen at Calvia in 2006. I think the players involved were from Bermuda and Papua New Guinea. They were banned for two years, which didn't actually matter as the only rated events they played in were the Olympiads.

There's no belief that stimulants etc. are any advantage in chess, so the anti-doping committee is only there for the demonstration of "compliance".

With bans for supposedly using computers liable to legal challenge, it isn't a good start for a statement in the first report that players can be banned purely on computer evidence to be hand-waved away that they didn't really mean that. Some proof-reading and a sign-off by all the contributors to the report would seem desirable.

(edit) Bermuda (/edit)
(edit)
If that incident demonstrated anything, it was that trying to test amateur players was a potential disaster area. No, or hardly any, attempts have been made since.(/edit)

None of the "media" websites have yet picked up on this report. If they are remotely following a role of challenging FIDE, expect a lot of questions on similar lines.

NickFaulks
Posts: 9255
Joined: Sat Jan 02, 2010 1:28 pm

Re: Chess Player Strip Searched

Post by NickFaulks » Mon May 19, 2014 12:26 am

Roger de Coverly wrote: They were banned for two years, which didn't actually matter as the only rated events they played in were the Olympiads.
I don't recall any ban, and during the subsequent two years Bobby played in five rated events, including the Turin Olympiad, so presumably that part was an unlucky guess.
If you want a picture of the future, imagine a QR code stamped on a human face — forever.

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: Chess Player Strip Searched

Post by Roger de Coverly » Mon May 19, 2014 1:23 am

NickFaulks wrote: I don't recall any ban, and during the subsequent two years Bobby played in five rated events, including the Turin Olympiad, so presumably that part was an unlucky guess.
It was certainly mentioned, but according to Ian Rogers, not actually applied.

http://www.inforchess.com/columnis/rog060.asp

The precedent as I see it, is that you need to know before you play in an event whether there are any stupid or unacceptable regulations in place so that if necessary you decline to play.

NickFaulks
Posts: 9255
Joined: Sat Jan 02, 2010 1:28 pm

Re: Chess Player Strip Searched

Post by NickFaulks » Mon May 19, 2014 2:15 am

I suppose it is natural that Ian would portray Shaun in the lead role, but I felt we were equal partners in Calvia, and Nick deFirmian made it his personal crusade. I think Ian exaggerates the devastation felt by the two teams at the points deduction - a ban would actually have meant something. I have just had to check Olimpbase to confrim that history shows that we really were bumped down the table.
If you want a picture of the future, imagine a QR code stamped on a human face — forever.

KWRegan
Posts: 9
Joined: Sun May 18, 2014 5:42 pm

Re: Chess Player Strip Searched

Post by KWRegan » Wed May 21, 2014 6:20 am

I've just registered for this forum, and can answer various questions about my model. A few things in recent pages:

1. The 20-year estimate for a 1M-to-1 false positive (using 4.75 sigma as criterion) is based on about 200,000 tested games per year, not 50,000, because the unit of testing is player-performance-in-one-event, not individual games. One can figure 8 games per event on average, but as there are 2 sides to a game it's a multiplier of 4. The way I've described it before is based on 1,000 player-performances per week of TWIC, which amounts to a million per 20 years. Since many events are split across 2 weeks of TWIC, this number leaves some growing room for getting all games and from more events. Another commenter is right to note the biased-selection factor of getting only the top-board games from many events, but I allow for it and getting all the games from events goes into that allowance. In any event, the section is written to allow principals to prefer the 5.00, "3M-to-one", 60-year standard as safer; one reason for proposing 4.75 is that a couple 4.8's have arisen in practice.


2. Regarding
"Check your assumptions, the paper recommends the statistical analysis (with the given risk of false positive) to be used as sole evidence for a conviction. In other words, in Little Britain style, you are banned for 5 years because "computer says so" :-)
and
It does say that, but I don't believe that is really Ken Regan's intention, and it's his baby. The wording needs tightening up to confirm that the test will only be used to convict when there is independent supporting evidence, but in reality I don't think this is a frightening part of the paper.
---and regarding "turnaround" and "changed his mind"---

That is correct. At the time of a March 2012 NY Times story, I wrote a page "Parable of the Golfers" with the statement, "There just aren't enough moves and games and players to get even a sniff of five-sigma, but 3.5 sigma, that happens..." Prior to then I'd had no z-score above 4.00, not even in the Feller case, though one from August 2011 would qualify now that I have a better handle on ratings below 2200. Part of what I meant by "not enough games and players" is: how would you empirically test my 1-in-a-million projection? For a true field test, you need 1M player-performances, but per above that needs about 4M games---a large fraction of all the games ever recorded. And at 4-6 CPU-hours per game for my full test, it would take a long time. What I've done is run simulations on multiple 10K-size sets drawn randomly from my training data, which falls under accepted "bootstrapping" methods and directly warrants the projections out to the 3.50--4.00 range. Above 4.00 it's extrapolation, and a good defense lawyer could try to knock that down. Once I finish converting my model to the the Houdini-Komodo-Stockfish troika, getting it up to multiple 100K bootstrap trials will be possible.

I was shocked when numbers over 5.00 (even over 6.00 upon excluding round 8 and moves past 70) tumbled out of Zadar. I'd thought I'd only get them in cases of "consecutive events", either by combining the games or using Stouffer's Rule to combine the z-scores. The primary question I see is: what near-term actions can be justified from so high a readout alone? Banning is far-term, but IMHO the potential to ban is necessary in principle to open a process near-term, and for me a fair central process would have been much preferable to what we saw extend through last December.

Besides publications on my website I have also written articles for the Maths/CS blog I co-manage, whose standing in my field I analogize to Sergey Shipov's Crestbook, and article titles suffice for Google: "Thirteen Sigma" (about the 2010 Azov Don Cup) and "Littlewood's Law" explain perceptions about statistical results, while "The Crown Game Affair" laid out the argument on Ivanov (with campiness intended to mitigate the legal risk).


3. Among points earlier in this thread, let me say the "recipe" is in materials on my website and FIDE has no intent to keep it secret. I do regard full-test data as private, while screening tests (which have no z-score judgment value) are already essentially being done publicly by chess-db.com and some other sites. The most accessible version of the recipe may be the part of my Tallinn talk http://www.cse.buffalo.edu/~regan/Talks ... ssTalk.pdf from slide-overlay 47 onward. What happened is that I realized from the preliminary ACC meeting in Paris and from talking with people the night of my arrival in Tallinn and at breakfast that what people needed most to hear were the fundamentals of evidentiary statistics stated in the chess context. So I hurriedly wrote overlays 1--46 that day, finishing just before the 5pm meeting, and wound up stopping my talk at slide 48 where I'd intended to begin :) . The attitude of saying "Analytics" in the title is that it's not just statistics; two others of my private reports last year went into game-move detail. In Buffalo my professional colleagues duly tempered expectations about statistics, as did I in revising the document---but one must allow that in a dozen positive cases I count the statistics were or would-have-been effective, plus there are several dozen cases (4 so far this year) where the model turns aside accusations/loud-whispers I think are unfounded. I don't really disagree with points raised here about "smarter cheating"; what I hope is that "smarter science" together with smarter prevention will combine to make the risk/reward and complexity/detectability curves more adverse for potential cheaters.

The math in the "recipe" is not deep, and unlike online systems it has no sub-surface information such as move-timing and player-profiling. Once the main principle and a few "big-data-style" regularities are digested, there are only a couple choices beyond doing the simplest thing. I analogize the latter to realizing in the Marshall Ruy that 11...c6 is better than the original 11...Nf6 idea, and realizing that Black need not fear certain endgames. The current ramifications of updating the model are like all the "book" that's developed since, including taking d2-d3 and/or Re5-e2 seriously, but the basic ideas and effectiveness points of the Marshall have been pretty consistent.

[edit: added "directly" before "warrants" in point 2---I do think the extrapolation is warranted too, just not empirically directly.]

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: Chess Player Strip Searched

Post by Roger de Coverly » Wed May 21, 2014 8:31 am

KWRegan wrote: The way I've described it before is based on 1,000 player-performances per week of TWIC, which amounts to a million per 20 years. Since many events are split across 2 weeks of TWIC, this number leaves some growing room for getting all games and from more events.
TWIC is selective and always has been, deliberately only containing supposedly top class games. If there's some intention to monitor all chess rated by FIDE, the numbers of games played in a year are many times greater. The FIDE ratings server will know how many games are played. One in a million false positive would be quite frequent once you tested or notionally tested every game.

But what are you trying to detect? Is it someone who cannot really play chess, in which case a statistical method may or perhaps will catch them, or is it someone who can play above their rating by sometimes taking advice. In the latter case, without physical evidence, how do you know they consulted an engine during the game, given it is legal to consult before the game?

In some countries, the Federations are up to a point controlled by the players. Based on likely attitudes in the UK, I see no acceptance of the notion that you will be defaulted if you have a device in in your possession if it's in your pocket. We've already rubbished the idea up thread. Not all games are FIDE rated and if FIDE are going to require ridiculous rules that regard players as cheats unless demonstrated otherwise, events may withdraw from being rated or rule out ever becoming rated.

MartinCarpenter
Posts: 3180
Joined: Tue May 24, 2011 10:58 am

Re: Chess Player Strip Searched

Post by MartinCarpenter » Wed May 21, 2014 12:01 pm

Well he's doing it by event which makes a very big difference to the numbers :)

The approach in those slides looks exceedingly sensible overall.
(Have an agreed on test, trigger potential warnings to check for at a reasonable level of probability and so on upwards.).

I suspect you might find the stats for execptional performances falling apart rather if you get comfortably below 2200, because there's lots of very badly underrated juniors down there nowadays.

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: Chess Player Strip Searched

Post by Roger de Coverly » Wed May 21, 2014 12:36 pm

MartinCarpenter wrote:Well he's doing it by event which makes a very big difference to the numbers :)
Weekend tournaments are five rounds, four with a bye. League weekends are rated independently, so average of games per "event" will become lower than the implied eight. I would distrust an approach which drew extravagant conclusions using data collected from TWIC, as TWIC is itself a biased subset of the total of rated games. Indeed TWIC probably under-reports rated weekend Congress games. There's a bias being Mark's editorial choice of what games to include and whether he collects or receives the games in time for his publication deadlines. The FIDE ratings server now stores games in pgn, again that's biased by whether the organiser collects, inputs them and sends them in with the rating report.

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: Chess Player Strip Searched

Post by Roger de Coverly » Wed May 21, 2014 2:23 pm

MartinCarpenter wrote: I suspect you might find the stats for execptional performances falling apart rather if you get comfortably below 2200, because there's lots of very badly underrated juniors down there nowadays.
The standard historic interpretation of an out or under performance against an Elo rating was that the rating was obsolete and thus you apply a K factor to the difference in actual score v expected score to change the rating. It's also obvious that to get a higher rating, you need a higher standard of play. So when someone out-performs, they've got better at chess. It doesn't follow in the absence of physical evidence that their higher standard of play was because they were consulting an engine.

KWRegan
Posts: 9
Joined: Sun May 18, 2014 5:42 pm

Re: Chess Player Strip Searched

Post by KWRegan » Wed May 21, 2014 2:42 pm

In reply to Mr. de Coverly, I would very much like to see someone do an independent assessment of sources and numbers of games and "biases", including looking at actual issues of TWIC and other sources. Also please bear in mind the distinction between screening tests (which carry no z-scores) and full tests (which only arbiters and ACC may call for). I've run full tests on my top several dozen screening-test results, and most turn out to be minor deviations in the full test, mainly because the games were more tactical/forcing and so have a higher matching expectation to begin with. Whereas I believe that many full-test outliers would come (say) from people in the bottom quarter of the rating chart who have a great day and manage an even score, or even 5.0/9 in a bigger event, with nobody else caring.

As for weekenders, the small sample size of games is a natural obstacle to statistical conclusions of any kind. For league play, whether it is even appropriate to sub-divide a weekend from an "event" has been a large consideration in two of my reports already. Not unduly burdening organizers of these events was an uppermost purpose, as was reassuring professionals in the larger and longer prize events. We considered lots of arguments about cellphones, including their tendency to sprout wings and fly away when left in the open :o . The time-current intrinsic rating of junior chess is a major topic on-tap. Let me just say that in my own rising-junior experience there were bumps both ways, while measurement of rating-inflation effects will have to go down to their level as it does not seem to have percolated up to 2600 level, see http://www.cse.buffalo.edu/~regan/chess ... reg4yr.jpg as well as my papers in 2011.

Roger de Coverly
Posts: 22607
Joined: Tue Apr 15, 2008 2:51 pm

Re: Chess Player Strip Searched

Post by Roger de Coverly » Wed May 21, 2014 3:05 pm

KWRegan wrote:In reply to Mr. de Coverly, I would very much like to see someone do an independent assessment of sources and numbers of games and "biases", including looking at actual issues of TWIC and other sources.
Anyone who uses TWIC knows full well that it's not a complete source of information and never has been. If you play lower rated players, which in this context means sub 2200, and want information on what and how they play, a necessary approach can be to build up your own collection of games from relevant organiser's websites. That said, you might get lucky by checking your opponent on the FIDE rating site as for some events, pgns are stored there. This will give games not in TWIC and thus not in the annual commercial ChessBase collections either.

In
http://www.chessprofessionals.org/conte ... g-proposal
I read
The Committee recommends the implementation of a FIDE Internet-based Game Screening Tool for pre-scanning games and identifying potential instances of cheating, together with the adoption of a full-testing procedure in cases of complaints. Together they shall meet the highest academic and judicial standards, in that they have been subject to publication and peer review, have a limited and documented error rate, have undergone vast empirical testing, are continuously maintained, and are generally accepted by the scientific community. Once in place, the Internet-based Game Screening Tool will be accessible to arbiters and chess officials and will be a useful instrument to prevent fraud, while the full test procedure will adhere to greater privacy as managed by FIDE and ACC.


Is this saying that if you can collect and input the games, they can be screened? This potentially applies to games by near beginners, not just the elite who might expect to appear in TWIC.

It's all very well making computer only cheating accusations against players sitting at home behind a computer screen, since there are no witnesses. It's different or at least should be different for games played in the presence of an opponent, other players, spectators and arbiters.

Here's the details of a UK event in the near future for which every game will potentially be FIDE rated including the lowest section.
http://www.e2e4.org.uk/sunningdale/May2014/index.htm

Post Reply