HomeWorld CricketEmpty Rows Are Not Zeros: How Cricket's Silent Data Gaps Distort Its Baselines

Empty Rows Are Not Zeros: How Cricket's Silent Data Gaps Distort Its Baselines

প্রশ্ন: ক্রিকেটের তথ্যভাণ্ডারে খালি সারি আর শূন্যের পার্থক্য কী? মূল উত্তর: খালি সারি মানে তথ্য অনুপস্থিত, শূন্য মানে ঘটনা সত্যিই ঘটেনি। এই দুটোকে এক করলে Average, Economy ও হোম-উইন বেসলাইন বিকৃত হয়। সঠিক পদ্ধতি হলো প্রতিটি ফাঁকাকে সত্য-শূন্য, এলোমেলো-অনুপস্থিত, নাকি অপর্যবেক্ষিত হিসেবে শ্রেণিবদ্ধ করা। মূল তথ্য: - ২০২০ সালের মার্চে করোনাভাইরাসের কারণে বাংলাদেশ প্রিমিয়ার League স্থগিত হয়। - দর্শক-উপস্থিতিসহ হোম-উইন বেসলাইন ৪৩.৭%; দর্শকশূন্য মৌসুমে তা ৩৭.৯%। - চোদ্দো মাসে আগের চার মৌসুমের ৪৬২টি ম্যাচ পুনরায় কোড করা হয়। - একটি ঘরোয়া সিজনে ২২টি ম্যাচের মধ্যে ৪টি সারি সম্পূর্ণ ফাঁকা পাওয়া গেছে। - টেস্ট, ওয়ানডে ও টি-টোয়েন্টির সিদ্ধান্ত কখনো মেশানো উচিত নয়। সূত্র: লেখকের হাতে-কোড করা ঘরোয়া সিজন ডেটাবেস ও পদ্ধতি-পরিশিষ্ট, প্রকাশিত ২০২১ | Cross-checked: cricsultan.com সম্পর্কিত প্রশ্নোত্তর: প্রশ্ন: কেন ফাঁকা সারি বিশ্লেষণের আগে শ্রেণিবদ্ধ করা জরুরি? উত্তর: কারণ অজানাকে শূন্য ভাবলে ভুল Average তৈরি হয়, আর ম্যাচ বাদ দিলে নমুনা ছোট হয়ে যায়। প্রশ্ন: এই ডেটা-ফাঁক কি নির্দিষ্ট League বা দেশে বেশি? উত্তর: হ্যাঁ, কম সম্পদশালী Leagueে ফাঁকা সারি বেশি, যা গ্লোবাল বেসলাইনকে ধনী Leagueের দিকে ঝুঁকিয়ে দেয়; cricsultan.com Player Depth Index-এ এই প্রবণতা লক্ষণীয়। প্রশ্ন: একটি বেসলাইন কখন বিশ্বাসযোগ্য? উত্তর: যখন তার পাশে নমুনার আকার (n), পদ্ধতি ও টাইমস্ট্যাম্প প্রকাশ্যে লেখা থাকে।

Last month I opened the hand-coded season file from 2026 again, and the margins disagreed. The database was supposed to contain 22 matches from that season; on inspection, four of those match rows were entirely blank — no ball count, no runs, no wickets, no toss result. Yet the 'total matches' column still returned 22, and the 'total runs' column still handed back a number without any error. A blank cell and a zero are not the same thing — but our scorebooks, spreadsheets and online databases all treat them as one. That silent error is the centre of today's discussion.

I have spent sixteen years in broadcast work from Chattogram, and alongside it I have counted shots in a notebook nobody asked for. In 2026, when Facebook Live and YouTube highlights began displacing Bangladesh's evening TV wrap, I did not chase the new format; instead I hand-tagged every shot of one domestic side's 22 matches — location, body part, defensive pressure. That habit fixed a rule I no longer break: I do not write a claim without a denominator. Every piece begins with a count — shots, presses, metres — and only then allows itself one sentence of judgement. Today's subject is the exact opposite of that rule: when the denominator itself is blank.

Context

Cricket's statistical history is old — Wisden's scorebooks, the BBC archive, local newspaper results. But the age of a statistic and the reliability of a database are not the same thing. There is a fundamental difference between a scorecard and a database: a scorecard is a description, a database is a decision. On a scorecard, a blank space means 'nothing was written here' — the reader understands. In a database, that same blank cell is read by the machine as zero, and that zero flows into averages, strike rates and home-win baselines.

In Bangladesh's context the problem is sharper. Here domestic cricket data lives largely in handwritten scoresheets, club files and local journalists' notebooks. Before and after the Bangladesh Premier League was suspended in March 2026 because of the coronavirus, some matches each season were washed out by rain, some were postponed a day and played at another venue, and some scorecards never reached a digital archive at all. When someone builds a 'season summary', those missing matches quietly become zeros and enter the totals — when the truth is that those balls were never bowled. Understanding the difference between absent and achieved-zero is now compulsory for me.

Core Analysis

From 2026 to 2026, when the Bangladesh Premier League had stopped and stadiums were empty, I wrote no opinion. For fourteen months I re-coded 462 matches from the previous four seasons — logging shot location, game state and attendance for each. That work taught me that silence is itself a dataset; you only have to know how to read it.

The attendance baseline was a 43.7% home-win rate. But before that I had to answer a question: matches that were abandoned, or whose attendance was never recorded — how should I count them? If I count them as zero, the average drops, and I wrongly conclude that 'crowds were low so home advantage was low'. But the truth is that where there is no information, home advantage may not exist, or it may — we do not know. Treating the unknown as zero means handing the unknown a false measurement.

And here is my favourite piece of work: splitting the blank rows into three classes. First, the true zero — the ball genuinely did not happen, so 'wickets = 0' is correct when no wicket fell. Second, missing-at-random — the scorecard was lost, but the match happened. Third, unobserved — the match never started, or happened outside our camera. These are three different things, and collapsing all three into one is the database's greatest crime.

There is a real example behind this classification that I have seen many times. Say a spinner's domestic economy is being calculated. Two of his 12 matches were washed out by rain, matches in which he bowled not a single ball. If those two matches are counted as '0 runs, 0 overs', the economy falls falsely and he becomes artificially the best bowler of all. Conversely, if the matches are dropped entirely, the sample shrinks, and someone declares a 'form' based on a small sample. Both paths are wrong, and both errors come from the same place — failing to check the nature of the blank row.

I found four blank rows among the 22 matches of a domestic season. A small number, right? But drop four matches and you cannot reliably state any win-average-policy from the remaining eighteen, because eighteen matches do not represent a full season. And if you add the four as zeros, what you get is entirely fictional.

I follow one method for this verification. First I write a metadata line for each blank row: did the match happen? was the scorecard found? was the data lost, or never written at all? Then I place each gap into a class — zero, missing-at-random, or unobserved. Only then do I calculate, and beside the calculation I write the sample size (n) clearly. I do not print a number without its n.

Why is this habit so important? Because most of cricket's baselines — home advantage, fast bowlers' economy, openers' averages — are built from a mixture of memory and database, and we assume memory and database agree. But my experience says they often do not. When I reopen a handwritten season, the margins often disagree. And that disagreement is the real information — not the number, but the gap in the number.

There is another subtle trap here, one of format. Test, ODI and T20 have fundamentally different tactical logic, and conclusions from one format cannot be mixed into another. In domestic databases this format label is often blank. So a T20 strike rate from one place and a Test average from another are joined together, and an impossible 'greatest of all time' list is produced. Without a format label, a run figure is to me like unlabelled medicine — it may work, it may harm, but nobody knows which.

The Lesson of the Timestamp

Another dimension matters here, which I learned in 2026. That year, during a World Cup knockout match, I was tagging pressing off a 720p feed from Chattogram. The match was in one team's control until the break; but after the break I saw the pressing intensity change — PPDA (passes per defensive action) was 11.8 before the break and fell to 6.9 after it. Parity arrived in the 68th minute. I filed the chart at the 90th minute, before extra time began, and the outlet published it while the match was still being decided. That experience taught me: publish from inside the event, and stamp every chart with a time — so nobody can look back and say this is hindsight.

Minute sixty — that is where the match stopped obeying its script.

The same rule applies to domestic cricket databases. An innings average is meaningful only when I know how many balls it rests on, in which over the match changed pace, and when the data was recorded. A number without a timestamp is to me like an undated witness — saying much, proving little.

Testing the Press Against the Heat

In 2026, during Euro 2026 and the Tokyo Olympics, the whole industry celebrated gegenpressing as the new meta. I did not repeat it, I tested it. Across 51 Euro matches, teams with a PPDA under 8.0 won 12 of the 20 knockout-relevant games. But in the Tokyo men's tournament, at 33°C and 70% humidity, teams in the same PPDA band won only 3 of 11. Same tactic, different environment, different result. This is where my baseline-custodian self wakes up: I do not interpret a season without the previous season's control, and I fix the environment — temperature, pitch, humidity — as context variables in advance.

These environment variables matter for domestic cricket too. Chattogram's humidity, the time of day, the character of the pitch — without knowing them, a spinner's economy or a fast bowler's strike rate is just a number, not a story. And these variables are often blank in the database, because nobody recorded them. So we analyse on incomplete information while projecting the confidence of completeness.

Let me make one example clear. In the Bangladesh Premier League, the matches of stars such as Shakib Al Hasan, Mushfiqur Rahim, Mahmudullah Riyad and Tamim Iqbal also live in this same domestic database. The reliability of their records depends on the blank rows of that database — how many matches were recorded, and how many are silent. There is no question about their achievement; the question is about our record.

Contrarian Angle

Now let me say something uncomfortable, which runs against the tone of this whole piece. Not every blank row is a fault. As data analysis has grown more popular this decade, a tendency has grown with it: to find someone to blame for whatever cannot be found. But not every absence is an administrative failure. Some information genuinely was never created — because the match did not happen, the ball was not bowled, the decision was not taken. To call that 'data lost' is to pin blame on a thing that never existed.

So I classify each gap first, then judge it. Without understanding the difference between true zero, missing-at-random and unobserved, if someone declares 'Bangladesh's domestic cricket data is corrupt', they are building a narrative instead of evidence. And narrative spreads faster than evidence, especially when the number is blank and the claim is loud.

The second contrarian point: correlation and causation are different. Domestic league crowds fell and home wins fell — the two happened together, but that does not mean crowds create home advantage. There may be a third cause behind both that we did not measure. When I say 'if crowds return, home wins will return', I am saying more than the evidence supports. The honest answer is: we have an association, not a causation.

Empty Rows Are Not Zeros: How Cricket's Silent Data Gaps Distort Its Baselines

And a third point, which I make against myself: my favourite line — 'silence is itself a dataset' — is dangerously beautiful. It is true, but it has a trap. If someone romanticises silence too much, they miss the ordinary cause hiding behind blank data — often it is not a thrilling mystery, but ordinary neglect, fatigue, or lack of budget. Calling silence data does not mean inventing a story for it; it means looking for its cause.

The Politics of Incompleteness

One point must be added here, because without it the picture is incomplete. Data gaps are never neutral. A league or country with more resources has more of its matches caught on camera, more scorecards digitised, more analysts hired. A league with fewer resources has more blank rows — because there are fewer people to keep records, fewer tools, less time. So when someone builds a global baseline, they are in fact making the rich leagues' data the standard and pushing the poor leagues' blank data to the margins. That inequality is the quietest truth of data.

I stay conscious of my own role here. I was born in Australia, I work in Chattogram, I am comfortable in technical language. From this position it is easy to slide into a 'come from outside and lecture' register. I want to avoid that. Bangladesh's domestic cricket data is being built by local scorers, statisticians and journalists — often unpaid, with scarce equipment. Before speaking of flaws in their work, I acknowledge their constraints. The problem is not their skill; the problem is the system.

There is another layer, which touches the very beginning of cricket's history. Many records from the first decades of Bangladesh cricket still travel by word of mouth, not in a database. Someone remembers a match's score, but it was never written on paper — this 'oral history' is valuable, but it is never zero, and it is not verifiable either. When someone collapses these two layers — oral memory and digital record — into one, a third, imaginary baseline is created.

Takeaway

So what will my eye look for in the next round? I will watch three signals. First, whether the number of blank rows in domestic cricket databases is coming into the open — written as a hidden '0', or honestly as 'no data'. Second, whether any baseline declaration is printed beside its sample size and method. Third, whether the habit of recording match timestamps and environment variables is growing, or whether analysis still rests on memory as before.

I know the blank cells will not fill themselves. Someone must give them a class — zero, missing, or unobserved. The season I hand-coded taught me this lesson: the number does not tell us the truth, the gaps in the number do. And before the next match begins, my first job will be to know — what the denominator actually is.

Related Players