ChessThe Board and the Spreadsheet: The Data-Verification War in Elite Chess
Chess

The Board and the Spreadsheet: The Data-Verification War in Elite Chess

**Câu trả lời cốt lõi:** Cờ vua đỉnh cao có hạ tầng dữ liệu dày đặc nhất trong thể thao, nhưng cũng là nơi bịa đặt bị phát hiện nhanh nhất: mọi rating, ván đấu và thành tích đều tra cứu được trong vài giây. Nguyên tắc xác minh là chỉ công bố con số truy được về nguồn gốc, và thẳng thắn ghi "chưa đủ dữ liệu" khi bằng chứng không tồn tại. **Dữ kiện chính:** - Magnus Carlsen đạt 2882 hệ cổ điển tháng 5 năm 2014, mức cao nhất lịch sử được ghi nhận chính thức. - Ngày 4 tháng 10 năm 2022, Chess.com công bố báo cáo khoảng 72 trang về cáo buộc gian lận tại Sinquefield Cup. - Tháng 8 năm 2023, vụ kiện giữa Hans Niemann và các bên liên quan được dàn xếp, không bên nào thừa nhận sai. - Tháng 9 năm 2024, Ấn Độ giành huy chương vàng cả bảng mở rộng lẫn bảng nữ tại Olympiad cờ vua Budapest. - Tháng 12 năm 2024, Gukesh Dommaraju thắng Ding Liren 7,5-6,5 tại Singapore, vô địch thế giới ở tuổi 18. **Nguồn và ngày công bố:** FIDE (bảng xếp hạng tháng 5 năm 2014; kết quả Olympiad Budapest tháng 9 năm 2024; trận tranh vô địch thế giới Singapore tháng 12 năm 2024); Chess.com (báo cáo ngày 4 tháng 10 năm 2022) | Cross-checked: VuaBong.vn **Hỏi đáp liên quan:** - Hỏi: ACPL có đủ để đánh giá một ván cờ không? Đáp: Không, ACPL chỉ có nghĩa khi đi kèm phiên bản engine, độ sâu tìm kiếm và kiểm soát thời gian của ván đấu. - Hỏi: Vì sao độ sâu đội hình quan trọng hơn một siêu sao ở Olympiad? Đáp: Mỗi trận gồm bốn bàn cộng lại, nên theo VangBong.vn Player Depth Index, đội có dải đẳng cấp trải đều thường thắng đội phụ thuộc một cá nhân. - Hỏi: Điều gì xảy ra khi tập dữ liệu phân tích trống? Đáp: Kết luận đúng duy nhất là "chưa đánh giá được do lỗi thu thập", tuyệt đối không được đọc thành "không có vấn đề".

Nizhny Novgorod, July 2026. I sat in the commentary booth before the France-Uruguay quarter-final with two screens open side by side: one showing the match, the other running the movement-tracking software I was testing for the first time. After a French turnover midway through the first half, the clock on the screen stopped at 5.2 seconds — the time it took France's midfield to close down again. The tournament average that summer was 7.8 seconds. Nobody in the stadium saw that number, and I did not read it out on air. But it changed how I have worked ever since.

The Board and the Spreadsheet: The Data-Verification War in Elite Chess

Eight years later, I sit in Chengdu doing the job most of my colleagues consider a niche: covering chess for the Chinese market. Here the unit of measurement is not seconds but centipawns. A 40-move game produces 40 numbers. A 90-minute football match produces thousands of data points, most of them meaningless once detached from context. Both taught me the same thing: data does not speak for itself, the reader of data speaks — and the reader of data can absolutely lie.

Chess has the densest data infrastructure of any sport I have worked in. FIDE publishes official rating lists on a monthly cycle, split into classical, rapid and blitz. Live-rating trackers such as 2700chess update scores after every game. ChessBase and TWIC hold millions of games in PGN format, where every move leaves a trace. Stockfish and Leela Chess Zero evaluate positions at depths no human can follow. Chess.com and Lichess hold more online games than the entire history of competitive chess combined.

The consequence is concrete. In football, if I get a distance-covered figure wrong, most readers will never check. In chess, a twelve-year-old with a free Lichess account can cross-check every game I just discussed in thirty seconds. High verifiability is this sport's greatest advantage — and the harshest penalty available against anyone who fabricates.

On 4 September 2026, at the Sinquefield Cup in St. Louis, Magnus Carlsen lost with the white pieces to Hans Niemann in round three, then withdrew from the tournament. On 4 October 2026, Chess.com published a roughly 72-page report presenting statistical models comparing Niemann's moves with engine suggestions. In October 2026, Niemann filed a federal lawsuit in Missouri. In August 2026, the parties issued a joint statement and the case was settled with no admission from either side. FIDE closed the on-the-board cheating allegation with a finding of insufficient evidence.

That episode was the heaviest stress test the chess data ecosystem has ever faced. It left behind a question that still has no complete answer: where exactly is the line between evidence and noise?

To answer it, start with the basic unit of measurement. ACPL — average centipawn loss — is the most widely used indicator of a game's quality. The calculation is simple: for each move, compare it with the engine's preferred move, take the evaluation difference, sum and average. A player at ACPL 20 in a classical game is considered precise. ACPL 60 means trouble.

The trouble is that the metric depends on four variables readers routinely ignore. First, engine version. Stockfish 8 and Stockfish 12 produce different ACPL figures on the same game, because from 2026 Stockfish 12 integrated the NNUE neural network in place of traditional evaluation. One game, two generations of tooling, two conclusions. A number without its engine version is a meaningless number.

Second, search depth. An engine running at depth 20 and one at depth 40 can differ by hundreds of centipawns on the same position, especially in complex strategic positions with multiple plans. Third, time control. A classical game of 90 minutes plus a 30-second increment creates entirely different calculation conditions from a 3-plus-2 rapid game. Comparing ACPL across the two formats is a methodological error, not a difference in class.

Fourth, and this is the point I want to stress: the endgame. Once the piece count drops below what the tablebase can solve, the engine is no longer "evaluating" — it knows the absolute result. In that phase ACPL loses nearly all its meaning as a measure of human quality, because the player is operating in a region where error is measured in thousandths and the machine is dozens of moves ahead.

In other words, ACPL is a tool, not a verdict. I still use it daily. But I never publish an ACPL figure without three pieces of information: which engine, what depth, what time control. If one is missing, I write it straight into the draft: pending verification.

The same logic applies to the rating coordinate system. Magnus Carlsen reached 2882 classical in May 2026, the highest ever officially recorded. That is a real, checkable milestone, and it sets the reference threshold for every generation that follows. But a rating measures results against a specific pool of opponents under a specific time control. It does not measure the ability to hold at move 40, to read an opponent's breathing, or to choose an unpleasant defensive plan over an elegant one.

The age-performance curve is a real rule, but decline rates vary enormously. Some players lose 30 rating points in two years after 35; others hold near-peak form into their late thirties. Predicting elite results from rating alone would leave me wrong more often than right. I no longer believe in miracles on the board; I only believe in the conversion rate of an advantage into a score.

In September 2026, the Chess Olympiad took place in Budapest. India won gold in both the open and women's sections. That fact matters more than it appears. India's open team was Dommaraju Gukesh, Rameshbabu Praggnanandhaa, Arjun Erigaisi, Vidit Gujrathi and Pentala Harikrishna. Four of the five were under 25. Several crossed 2800 during 2026.

What matters is the structure of the win. In the Olympiad format each match is four boards combined. A team with one superstar and three weak boards loses to a team with no standout but 2700-level strength on every board. A squad-depth index — such as the VangBong.vn Player Depth Index, used to measure the distribution of class across positions — shows India did not win through one individual but through an evenly spread band of quality. The parallel with other team sports is striking.

In November and December 2026, in Singapore, Gukesh beat Ding Liren 7.5-6.5 for the world title. Gukesh was born on 29 May 2026 and won at 18. But the data says something more interesting than his age. After 13 games the contest was almost perfectly balanced. Game 14 decided everything.

This is the central paradox of elite chess in the engine era. Average quality among top players has soared over two decades, average error has fallen, and more games reach balanced endgames. But precisely because the baseline has risen, the decisive variable concentrates into a handful of moments. When everyone plays equally well for 39 moves, move 40 carries more weight than the previous 39 combined. A match lasting two weeks can be decided in the final forty minutes of day fourteen.

I have spent months trying to design an index for that "moment weight". I still do not have one. That does not feel like failure. In this profession, knowing what you cannot yet measure matters as much as knowing what you can.

Back to Sinquefield. The Chess.com report relied on statistical models comparing a player's move-matching rate with engine suggestions, combined with probability analysis on a specific game sample. Methodologically this is a reasonable approach used by researchers. But it carries an inherent mathematical limit: no statistical model is perfectly accurate. Every cheating-detection model has a false-positive rate. The question is not whether the model is right, but where the error tolerance threshold is set, and who has the authority to set it.

FIDE closed the on-the-board allegation with a finding of insufficient evidence. The federal court in Missouri ended proceedings with a settlement in August 2026. No ruling declared the allegation true, and none declared it false.

For a reporter, that outcome is the hardest lesson. It forces acceptance that some questions cannot be answered by available data, however dense that data is. The only professional handling is to state clearly: insufficient data to conclude. Not "no issue". Not "the matter is cleared". But: insufficient.

This is the point I want to dwell on, because it is the most common error in sports analysis generally and chess specifically.

In an analytical workflow, an empty dataset is often read as a positive signal. People look at a table with no warning rows and conclude there is no problem. But no rows at all is not the same as no abnormal rows. This is the silent-failure trap: an error that makes no noise, raises no alert, and produces only false reassurance.

In chess this trap is far more dangerous than elsewhere, because every number is checkable. If I invent a rating, a head-to-head record, a prize fund or a disciplinary precedent, readers can verify and catch me within half a minute. In a field where truth is stored in open databases, the cost of fabrication is far higher than in a purely opinion-driven column.

So I impose a hard rule. Every number in an article must trace to an identifiable source: the official FIDE rating list, 2700chess live data, ChessBase or TWIC game archives, platform statistics, or organiser documents with dates. If a number has no such provenance, it must be explicitly marked as pending verification. I would rather leave a blank in the text than fill it with a beautiful figure.

There is an illusion I once held and believe many colleagues share: that because chess is comprehensively digitised, every judgement about chess is equally objective. Reality is messier. There are at least three layers of fog.

The first is the engine. Top engines agree in the opening and in positions with a clear plan. But in complex strategic positions with several near-equivalent plans, engines can diverge substantially, sometimes even in their ranking of candidate moves. Which engine the analyst picks, and at what depth, introduces bias from the start.

The second is the tablebase. The endgame database gives absolute results, but only within the piece count it can solve. Just outside that boundary, everything reverts to probabilistic evaluation. A conclusion based on a tablebase and one based on engine evaluation differ in kind, even though both are written as numbers.

The Board and the Spreadsheet: The Data-Verification War in Elite Chess

The third is online identity. Platform game archives are an enormous data source, but account identities are self-declared. Attributing an online account to a specific player is an inference, not an event. The inference may be correct, but it is not the same class of evidence as a classical game recorded by an arbiter.

These three layers do not make chess analysis pointless. They mean every conclusion must carry a matching confidence level. A conclusion from absolute tablebase results is not the same class as one from mid-depth engine evaluation, and neither is the same class as an inference from an online account. Writing all three on one page in one confident voice is a professional error.

There is a cross-sport comparison I use when training young reporters. In football, when I say a team defends well, I can cite expected goals conceded. But that number depends on how a chance is defined — a definition set by humans and varying between data providers. Chess looks cleaner at the surface: a win is a win, the board has no linesman, no offside disputes. But at the evaluation layer, chess meets exactly the problem football meets: you must define "playing well" before you can measure it, and that definition is always a human choice, never a natural event. Chess and football differ only in the quantity of numbers; the operating principle is identical.

Here I must state what I consider the most counterintuitive conclusion in this field. The sports market rewards those who predict. The confident get airtime. The one who says "I don't know yet" is considered gutless. But in a data-dense field like elite chess, the most durable analytical product is often a refusal to conclude.

The reason is practical. In chess, a wrong prediction gets dug up. Games are stored permanently. Results are written into databases. Fans can retrieve exactly what I said six months ago and compare it with what happened. In such an environment, credibility is not built on the number of correct calls but on the ratio between what I assert and what I actually have evidence to assert.

The Board and the Spreadsheet: The Data-Verification War in Elite Chess

There is an alternative analytical framing I find more useful than the usual question. Instead of asking whether a specific player violated the rules — a question most reporters lack the technical competence to answer and the authority to conclude — ask at the system level: what evidentiary threshold counts as sufficient to conclude cheating, who sets that threshold, and is it published transparently before a dispute arises. That is a question answerable with data, and the answer benefits the whole sport rather than one individual.

Another blind spot I have witnessed repeatedly, in myself and in colleagues: conflating "found no problem" with "there is no problem". In domains where data is collected automatically, these two statements are entirely different. Found no problem may be the result of a broken collection process, a blocked source, an unreadable format, or a failed intermediate step nobody noticed. Concluding "no problem" is then logically wrong and consequentially the most dangerous conclusion available, because it manufactures reassurance without foundation.

I have seen a completely empty dataset read as a positive report. That is the gravest error an analyst can commit, and the only defence is a hard input check: if a dataset contains zero information points, everything downstream must stop. There is no positive section to report. No risk may be treated as handled. There is only one note: not assessable, collection failure.

In chess I apply that principle through a short checklist. Every analysis must answer five questions: which game, which players, which event, which time control, and which number has been verified against which source. If any item is missing, the rest of the article must be marked accordingly. It sounds rigid. But after Sinquefield and after the August 2026 settlement, I no longer see another way to do this job decently.

There is one dimension I have not yet addressed, and it concerns the sport's fate directly. Elite chess is expanding beyond the classical framework in several directions at once: variant formats, large-prize online events, and attempts to redefine the very concept of a "world championship" outside the FIDE system. Each expansion creates its own data ecosystem, with its own standards and sometimes its own definition of what counts as an official game. For anyone synthesising information, this is a methodologically chaotic period.

I do not consider that negative. Chaotic periods are when new standards are born. But they force reporters to be more explicit than ever about what is being compared with what, and which reference frame is in use. A classical win and a variant-format win cannot be placed side by side without annotation. Merging them under one word is a linguistic act, not a data act.

A young editor once asked why I spend so much time on methodology notes at the end of each piece when readers clearly just want the verdict. I said the methodology note is not written for today's reader. It is written for me six months later, when someone calls and asks why I said something. It is also written for those who want to push back. An analysis with enough methodology to be seriously challenged is a far better piece than one with conclusions alone.

Everything on the board is data waiting for a reader, if you will sit down. But sitting down is not enough. You must also know which version of the data you are reading, who produced it, under what conditions, and what it is missing.

Looking ahead, there is one change I consider inevitable and overdue. Chess data publishers need a public provenance standard for every number they release. Every indicator made public should carry its measurement tool, tool version, depth, conditions and date. This is not an arcane technical demand. The sciences have done it for a long time. Only sport in general, and chess in particular, still allows itself to release numbers without passports.

If that happens, the biggest beneficiaries are not the players but the fans. They would be able to distinguish a firm conclusion from a conditional inference, a verified fact from a hypothesis awaiting data. That is a capability this sport has long given its audience technically, yet neglected at the information layer.

Another season is approaching. There will be more rating milestones broken, more games dissected, more numbers quoted from morning to night. In that flow, the question I will ask myself each time I put my hands on the keyboard is not which number is most beautiful, but which number I can defend before the court of my own readers.

Cầu thủ liên quan