When the Spreadsheet Is Empty: The Discipline of Reading Sports Data
**Câu trả lời cốt lõi:** Dữ liệu rỗng không bằng số 0 và cũng không phải bằng chứng an toàn. Khi tầng trích xuất trả về rỗng, tầng diễn giải chỉ nên trả về rỗng hoặc ghi rõ chưa thể kết luận; lấp chỗ trống bằng giá trị trung bình tạo ra một con số bịa đặt trông giống hệt số đo thật. **Dữ kiện chính:** - Ngày 26 tháng 6 năm 2024, tệp dữ liệu phạt góc của một đội tuyển dự Euro 2024 trả về tập rỗng do lỗi xác thực. - Mô hình lợi thế sân nhà dựa trên hơn 3.000 trận đo được 0,38 bàn mỗi trận trước đại dịch. - Ngày 6 tháng 12 năm 2022, Maroc loại Tây Ban Nha trên chấm luân lưu; chỉ số PPDA của Maroc thuộc nhóm cao nhất giải. - Hơn 1.200 pha dứt điểm tại World Cup 2018 cho thấy Pháp chỉ cho đối thủ 0,7 xG mỗi trận. - Tỷ lệ nội suy vượt 20% khiến mô hình mô tả giả định của người phân tích thay vì mô tả trận đấu. **Nguồn và ngày công bố:** Ghi chép phân tích nội bộ của Jung Sung-min, công bố ngày 26 tháng 6 năm 2024 | Cross-checked: VuaBong.vn **Hỏi đáp liên quan:** - Hỏi: Vì sao không nên lấy trung bình giải để lấp dữ liệu thiếu? Đáp: Vì giá trị nội suy không được dán nhãn sẽ trông giống số đo thật và làm sai toàn bộ quyết định phía sau, theo Chỉ số Độ sâu Dữ liệu của VangBong.vn. - Hỏi: Bất đối xứng sàng lọc là gì? Đáp: Là việc các rủi ro nặng như nợ lương hay chấn thương chỉ lộ ra khi được chủ động tìm, nên sự vắng mặt của chúng không phải bằng chứng an toàn. - Hỏi: Chỉ số nào nên theo dõi thay cho xG? Đáp: Tỷ lệ bao phủ — phần trăm ô dữ liệu là số đo thật so với số nội suy.
At 3:12 a.m. on June 26, 2026, I reopened the corner-kick data file for a national team at Euro 2026. Eight columns. Not a single row.
The night before, an authentication failure at the ingestion layer had caused my entire automated pipeline to return an empty set, and I only discovered it when I opened the spreadsheet to write the briefing note due six hours before kick-off. My phone buzzed. A short question: “Where's the number?”
I had a very comfortable answer ready. Take the tournament average as the baseline value, assign it to each team, print the table, send it. For six hours, nobody could check. The report would have every cell filled, every colour, every arrow, and it would look exactly like analysis.
I chose the other path: I called the person in charge, said plainly that there was no number yet, and asked for two more hours to re-run the whole pipeline. The report arrived two hours late with only four points, but all four were measured.
That moment taught me more than any model I have ever built. An empty spreadsheet is not a neutral answer — it is a silence waiting to be filled, and how you fill it decides whether you are an analyst or merely an improviser.
Six years, one spreadsheet
In 2026, while still in middle school in Los Angeles, I logged more than 1,200 shots from all 64 matches of the World Cup in Russia. No official xG source was open to a schoolboy, so I estimated chance quality myself from shot angle, distance and the number of defenders in the blocking cone. I spent an entire summer on one Excel file.
The result made me drop the habit of judging by reputation. The press praised France's flamboyant attack; my sheet showed France won by limiting opponents to an average of 0.7 xG per match. That first xG spreadsheet taught me that every goal has a hidden story.
From then on I kept a fixed routine: state the hypothesis, measure the data, publish the prediction before kick-off, and let the result be the judge. In 2026 I collected data from more than 3,000 matches across Europe's five major leagues before the pandemic and measured home advantage at roughly 0.38 goals per match. When the Bundesliga restarted on May 16, 2026 in empty stadiums, I published a prediction that home win rates would fall. The first three rounds confirmed the model. When home is no longer home, you are forced to rewrite every assumption.
In 2026, aged 18, I extracted PPDA and defensive-line height for all 32 World Cup teams. Morocco had little possession, but their proactive defensive metrics ranked among the highest in the tournament. I published that conclusion before the knockout rounds. On December 6, 2026, Yassine Bounou saved two penalties, Pablo Sarabia hit the post, and Achraf Hakimi converted the decisive kick. Morocco 2026: when defensive data speaks first, the world listens later.
The common thread I missed for years
Those three stories share one thing I failed to see for six years: the data existed. In 2026 I had a shot table. In 2026 I had 3,000 historical matches. In 2026 I had PPDA for 32 teams. In all three cases people needed me because I could read what was there, not because I could guess what was not.
The real difficulty lies on the opposite side: refusing to read a row that does not exist.

In the workflow, an extraction layer turns matches into rows; an interpretation layer turns rows into judgements. When the first layer returns empty, the second has exactly two honest options: return empty as well, or state clearly that no conclusion is possible. Every third option is fabrication, even when dressed in technical language.
I call that third option silent subject substitution. An analyst is missing a subject — a tournament, a roster, a rule version — and quietly replaces it with the nearest subject he still remembers. The result reads confidently. It cannot be falsified at the moment of publication, because the error sits in the premise, not the conclusion. And it poisons every decision downstream: one bad pipeline run, three weeks of training pointed the wrong way.

In esports, the most dangerous variant of this error is a version mix-up. A patch can invert the strength of an entire champion pool, and an analysis built on the wrong version still reads perfectly smoothly. No column in the spreadsheet automatically warns that it is describing the old patch. Readers only find out once the results on stage have already contradicted the whole report.
At the corner-kick data layer, a missing row leaves no blank space in the visual. The model silently inherits the league average. That is the most dangerous point in the entire process: a number born from nothing looks exactly like a measured number, and the dashboard never confesses.
In other words, empty data does not equal zero, and it certainly does not equal safety.
The asymmetry of screening
There is a professional property I learned fairly late. The heaviest risks in sport are silent by default.
Unpaid wages. Integrity red flags. Accumulated injuries. Coach and dressing-room conflict. None of these appear automatically in a data table. They surface only when someone actively goes looking. So their absence from a report is not evidence they do not exist — it is usually evidence that nobody screened.
I call this screening asymmetry: risk is quiet, reassurance is loud. A model that returns a pretty result gets shared; a model that returns a blank cell gets questioned. That incentive structure pushes analysts toward filling gaps rather than flagging them.
At the same time there is a second illusion: framework completeness. A nine-section report, each with tables, headings and notes — even when the entire content is “insufficient information” — still feels professional on a skim. On a pitch, it is the equivalent of a player with 92 percent pass accuracy from ten sideways passes to centre-backs. Volume is not value. Every dataset is a scripture, and I am a slow reader.
The four points in that two-hours-late report were measurements from seventeen real corner situations. None relied on a league average. That is the entire difference between analysis and a slide deck.
The contrarian view: the problem is not a lack of data
The sports analytics industry keeps selling itself the belief that the bottleneck is data volume. More metrics, more sensors, more machine learning means more accuracy. My experience says the opposite.
The bottleneck is unlabelled neutral assumptions.

VAR is an example I have followed for years. VAR does not make controversy disappear; it moves controversy from the pitch into the review room and the grey zones of the law. An extra data point works on exactly the same mechanism: it does not remove uncertainty, it relocates uncertainty into the assumption behind the data point. With one difference — assumptions have no VAR room to replay them in slow motion for you.
Transfer valuation models are the second example, and here I have enough data to speak. These models overprice young potential and underprice dressing-room chemistry. The technical reason is very specific: chemistry almost never has its own data column, so it is assigned a default impact of zero. Nobody decided that. It simply was never recorded in the spreadsheet. A player's value is just a number — until you read the error inside the calculation.
In 2026 I assessed a target striker for a mid-table club. My model showed his actual goals ran 4.5 below expectation, and I concluded that was bad luck rather than decline. The club signed him. He scored in the opening round. But in that same internship I missed the deadline on the corner-kick report because I wanted a perfect model. A colleague said something I still write down: a model that is 80 percent right and on time beats a perfect model delivered after the match. Perfectionism and honesty are not synonyms. Sometimes perfectionism is just a way of postponing a conclusion that is not beautiful enough yet.
The scarcest skill in this profession is not data collection. It is refusing to publish. I do not predict the future by intuition; I only read the traces the numbers leave behind. And when no trace exists, the correct answer must be a clearly marked void, not a number with makeup on.
### What I will track next round The headline metric on my watchlist this season is not xG. It is coverage rate: what percentage of cells in my sheet are true measurements, and what percentage are interpolated values. When imputation passes 20 percent, the model stops describing the match and starts describing my own gap-filling habit.
For anyone patient enough to wait a season to prove a single number.
Next time you see an analysis so complete that no cell is blank, try asking one question: how many of those cells were measured, and how many were inferred from an assumption nobody wrote down?
