The Discipline of the Null Result: Swimming Analysis and the Threshold Where You Must Say 'Insufficient Data'
Phân tích bơi lội chỉ có giá trị khi mọi kết luận truy nguyên được về một điểm dữ liệu nguồn; khi đầu vào rỗng, câu trả lời đúng duy nhất là "không đủ thông tin". - Khung phân tích gồm chín chiều: kỹ thuật, thành tích, hệ thống thi đấu, cục diện thế giới, luật và doping, sự nghiệp, rủi ro, truyền thông, lan tỏa ngành. - Tầng phân tích chuyên sâu phụ thuộc hoàn toàn vào tầng bóc tách: không có điểm thông tin thì không có kết luận nào truy nguyên được. - Kỷ lục 100m ếch nam của Adam Peaty là 56,88 giây, xác lập ngày 21 tháng 7 năm 2019 tại Gwangju. - Kỷ lục 100m tự do nam của Pan Zhanle là 46,40 giây, xác lập ngày 31 tháng 7 năm 2024 tại Paris. - Sự im lặng của nguồn về doping không phải là bằng chứng của hành vi sai trái; suy diễn từ đó là tạo tin giả. Nguồn: Phân tích chuyên sâu Stage-2, lĩnh vực bơi lội, ngày 13 tháng 8 năm 2026 | Cross-checked: VuaBong.vn Hỏi: Khi nào một kết quả rỗng có giá trị sử dụng? Đáp: Khi nó chỉ ra chính xác chỗ rò rỉ trong đường ống dữ liệu và đề xuất cách sửa. Hỏi: Dữ liệu tối thiểu để đánh giá kỹ thuật một kình ngư là gì? Đáp: Thời gian phản xạ, split từng 50 mét, nhịp tay, quãng đường mỗi sải, thời gian dưới nước và thời gian xoay bể. Hỏi: Vì sao không được quy đổi thành tích bể ngắn sang bể dài? Đáp: Hai hệ thống này đo trên nền khác nhau, xoay bể chiếm tỷ trọng khác nhau, nên quy đổi thẳng sẽ bóp méo kết luận.
THE DISCIPLINE OF THE NULL RESULT: SWIMMING ANALYSIS AND THE THRESHOLD WHERE YOU MUST SAY 'INSUFFICIENT DATA'
- TWO IN THE MORNING IN NHA TRANG, AND AN EMPTY COLUMN
Two in the morning. I reopen my swimming spreadsheet — a file four years deep, named by date rather than by inspiration — and the column labelled "source information points" is empty. No article title. No meet name. No athlete name. Not a single figure for reaction time off the blocks, no 50-metre splits, no stroke rate, no distance per stroke. An analysis file entered the deep processing layer with nothing inside it to process.

At the same moment, on a sports television channel, a commentator said of a regional female swimmer that "her start technique has improved dramatically compared with last season". He offered no reaction time. No time to breakout. No distance to surfacing. Not one percentage. Only a voice, and a belief.
That coincidence kept me at the desk another forty minutes. Because those two situations — an empty data file and a compliment with no number behind it — are in fact the same problem: people drawing conclusions about swimming without the material to draw them.
Data never lies, but it knows how to hide. And there is another possibility very few people in this trade will admit: sometimes data is not hiding. It simply does not exist. When that happens, the only correct answer for an analyst is to say so plainly.
- TWO ANALYSIS LAYERS AND THE COST OF THE DEPENDENT ONE
I work on a two-layer process. Layer one deconstructs the raw article: title, source, type, list of information points, the author's core positions, the article's purpose, named entities, time sensitivity, source quality. Layer two performs the deep nine-dimension analysis: technical, performance and data, competition system, world landscape, rules and anti-doping, athlete career and team system, risk profile, public narrative, and industry ripple.
The problem is that layer two is wholly dependent. It does not generate data. It only distils data that already exists in layer one. With no information points, no conclusion can be traced back to a source. And in this trade, a conclusion that cannot be traced to a source is not a conclusion — it is a guess dressed in terminology.
I know this from a long road. In 2026 I began my career at Thanh Nien newspaper as a swimming reporter. Back then I wrote by feel: great race, brave swimmer, spectacular finish. Only when I moved into data did I understand that every elegant adjective in a sports article must be paid for with a unit of measurement the reader can verify.
In swimming, the minimum material needed to say anything serious includes: reaction time off the blocks, splits per 50 metres, stroke rate per minute, distance per stroke, underwater time after the start and after each turn, turn time measured separately, and the difference between a 50-metre long course and a 25-metre short course. Without one of those, you can still write beautifully, but you cannot write correctly.
In the Vietnamese market the problem is worse. Domestic meets rarely publish complete splits. SEA Games or Asian Games reports on the Vietnamese team usually give only the final time and the placing. An article saying "Vietnamese swimmer finished with a time of X" says nothing about how that athlete distributed effort, how much was lost over the first 200 metres, or whether the final 50 was faster or slower than average. Readers have often asked me why I did not analyse a particular race. The answer is usually: because the source gave me no split at all.
- NINE DIMENSIONS, AND HOW THEY COLLAPSE WHEN MATERIAL IS MISSING
3.1. Technical dimension
A technical assessment of swimming needs four pillars: start and underwater segment, turns and finish, swim efficiency, and venue adaptability. Without reaction time and time to surfacing, you cannot know how much a swimmer gained or lost over the first 15 metres — the threshold World Aquatics allows for underwater travel after the start and after each turn.
In breaststroke the story is harsher. The rules specify the permitted number of kicks within a cycle, and extra kicks usually surface only if someone has a camera or a referee has the right angle. At domestic meets, kick counts are rarely recorded as data. Which means a technical dispute can occur with no one holding evidence to adjudicate it afterwards.
In butterfly and freestyle, what usually decides the race is rhythm management: a swimmer who raises stroke rate too early at 150 metres while distance per stroke collapses will break in the final 50. Without stroke-rate data per 50, we see the collapse but not the process of collapse. And when we see only the outcome, people tend to attribute it to mental strength — an explanation that cannot be tested and cannot be corrected.
3.2. Performance and data dimension
This is the dimension outsiders assume is easiest: you have a time, you compare it to a record, done. Reality is more complex. A number only means something in the right coordinates: world record, all-time list, and current-season ranking. Those three do not substitute for one another.
A concrete, citable example: Adam Peaty's 100-metre breaststroke world record is 56.88 seconds, set on 21 July 2026 in Gwangju, South Korea, according to World Aquatics records. Another, in freestyle: Pan Zhanle set the 100-metre freestyle world record at 46.40 seconds on 31 July 2026 in Paris, swimming the lead-off leg of a relay. And in the distance events, Katie Ledecky set the 1500-metre freestyle world record at 15 minutes 20.48 seconds on 16 May 2026 in Indianapolis.
Those three figures share one feature: each comes with a date, a place and a meet. Without those three, a number is only a rumour with a unit attached.
Pool type, meanwhile, is almost always omitted in Vietnamese-language reports. Long course (50 metres) and short course (25 metres) produce two different performance systems that cannot be converted directly. A swimmer who excels short course thanks to turn ability does not automatically excel long course. If an article does not state the pool type, the reader has no way to judge the number in front of them.
3.3. Competition-system dimension
A result is worth exactly as much as the tier of competition that produced it. Olympic Games, long-course World Championships, short-course World Championships, World Cup, continental meets, national meets — each tier carries different rigour and different predictive value. Position in the four-year cycle matters too: Olympic year, adjustment year, build-up year, sprint year.
I made this mistake early in my data years. In 2026, building an expected-goals model for V.League in a spreadsheet, I concluded Long An would be relegated because 13 goals scored corresponded to only 8.6 expected. They finished bottom with 18 points. But I must be honest: that call was right partly because of the model and partly because I had not yet separated the competition-tier factor. I later added a hard rule: every conclusion must state its competition tier and cycle position, or its confidence level must be dropped one notch.
In swimming this is even clearer. Qualifying pressure and final pressure are different things. A swimmer who breaks a national record in a morning heat but finishes eighth in the evening final is telling a story about energy distribution, not about character.
3.4. World landscape dimension
World swimming operates as a map by event. Some events have a titleholder so stable that the threshold to overthrow sits within rounding error. Others have an empty throne and wide volatility.
What matters most for a Vietnamese analyst is the talent supply chain. The US collegiate system develops through scholarships and year-round competition. The state-centralised system develops around major-meet cycles. The club system develops around commercial revenue. These three produce different athletes in terms of peak age, load tolerance and retirement age.
For Vietnamese swimming, the supply-chain question matters more than any other. A swimming nation is only durable when at least three cohorts across three age brackets are improving at the same time. If there is only one, every achievement depends on one person, and every injury to that person becomes a national crisis.
3.5. Rules and anti-doping dimension
Here I hold one absolute principle: the silence of a source is not evidence of wrongdoing. If an article does not mention doping, then for me to infer doping is to manufacture fake news, regardless of whether I phrase it as a question or as an open observation.
The checklist has four items: anti-doping, competition and officiating rules, equipment rules, and eligibility. In swimming, equipment rules have historically caused upheaval when high-tech suit lines were restricted. Those changes disrupted an entire generation of performances, and any cross-era analysis must account for them.
On officiating, risk points cluster around false starts, illegal turns, and breaststroke kick counts. These errors can change placings without changing times. In my tracking sheet I always keep a separate line for disqualified results, because they never appear in the time column.
3.6. Athlete career and team-system dimension
A swimmer's career curve has three sensitive segments. The first is puberty, when the body changes faster than technique. The second runs from 18 to 22, when training volume rises sharply and shoulder injuries begin to appear among freestyle, butterfly and medley swimmers, while knee injuries appear among breaststrokers. The third is after 26, when recovery capacity declines faster than the ability to hold strength.
An effective training model must answer three questions: domestic centralisation, overseas training camps, or a hybrid. And it must answer a fourth: who is on the sports-science and rehabilitation staff, how many are there, and is there weekly monitoring data.
In Vietnam I once issued a data warning about a footballer whose acceleration speed had dropped 38 per cent year on year, and a club official retorted that I was sitting in Nha Trang talking about the pitch. When football returned, that player transferred, played 11 matches and lost his starting place. I retell this not to boast. I retell it to say the same logic applies to swimming: acceleration data and recovery data are two early-warning tools, and they are usually ignored because they do not sit in the results table.
Luck is something I do not have. I have probability and sufficiently thick data. In swimming, sufficiently thick data means tracking the same athlete across at least three seasons, with the same metric set, measured the same way.
3.7. Risk-profile dimension
I split risk into six groups: competitive, career and system, anti-doping, rules, psychological and public opinion, and systemic. Each group needs three parameters: level, probability and impact.
The most dangerous group for Vietnamese swimming is career and system, because it connects directly to the supply chain. A top athlete retiring early without an equivalent successor leaves a gap that results cannot fill within two SEA Games cycles.
The second is psychological and public opinion. When media expectations run far ahead of real data, the athlete carries the mental load of a target built out of emotion. This risk appears in no results table, but it is present in every post-defeat interview.
3.8. Public-narrative dimension
Every athlete has a story being told about them. The analyst's job is to check whether fundamentals support that story, whether the sample size is large enough, and how long the story can run.
One type of story makes me especially wary: the story built from a single moment. A final-50 surge, a touch of the wall one hundredth of a second ahead. The moment is real, but the conclusion drawn from it is usually wrong. A team does not collapse overnight. It collapses when its indicators stop connecting to one another. And a swimmer does not rise to world class overnight either; he rises when several indicators move in the same direction across several consecutive seasons.
The gap between market expectation and objective assessment is where I work. People look at the price tag; I look at the curve. Many deals die before they are announced. In swimming, a "deal" is an investment slot, a scholarship, a national-team place — and the performance curve usually foreshadows the administrative decision by several months.
3.9. Industry-ripple dimension
A swimming result ripples across three tiers. Upstream is the youth training market and talent supply. Midstream is athletes and events. Downstream is broadcasting, sponsorship, equipment and derivative markets.
In Vietnam the strongest ripple is usually upstream: children's swimming enrolments rise after every SEA Games with strong results. That effect has a lag of roughly six to twelve months, and it appears in no federation's financial report.
One warning: a ripple is only credible when the causal path is clear. A medal and a wave of swimming enrolments happening at the same time proves nothing on its own. You must separate how much is due to the medal and how much to a drowning-prevention swimming campaign. If you cannot separate them, you are reading a correlation and calling it a cause.
- 'INSUFFICIENT INFORMATION' IS A PROFESSIONAL ANSWER
This is the hardest part of the trade, and the least taught.
When an analysis file reaches the deep processing layer with no information points inside it, there are three responses. The first is to fabricate. The second is to infer from context. The third is to return a null result with a complete framework, marking "insufficient information" at every position requiring substance.
The first is professional fraud. The second is educated fraud. The third is discipline.
The trouble with the third is that it is not rewarded. An analysis stuffed with conclusions will always be shared more widely than one saying I do not know. But the analytical trade survives because some people accept being called boring in order to stop the system poisoning itself with fake data.
I derived three rules from my own mistakes.
Rule one: set the statistical significance threshold in advance, then look at the data. If you look first and pick the threshold afterwards, you will always find a threshold that makes the story you want to tell "significant". That is mistaking statistical noise for signal, and it was my most frequent early error.
Rule two: publish your wrong predictions with the same frequency as your right ones. If over four years I only ever cite Long An 2026, Germany 2026, or a striker with a low expected-goals figure, then I am building an idol, not a model. A model is judged by its hit rate across all predictions, not by the ones that shone.
Rule three: end every piece with a qualitative check on the athlete's physical condition and state. Swimming is not only numbers. A swimmer can lose three per cent of performance to a mild infection three weeks earlier, or to an unhealed shoulder. Those things do not appear in a split. But if you ignore them, you will misdiagnose the cause of a change you measured very precisely.
There is another temptation I must block: using a null result as rhetorical weaponry. It is easy to write that domestic media data is so weak that analysis is impossible. That is a judgement, not an analysis. The dry truth is simpler: most domestic swimming meets have no automatic split-capture system, no budget for one, and no one assigned to the task. This is an infrastructure problem, not a moral one.
I must also state my own limits. In the deep processing layer, every conclusion must trace back to a source information point. When the source is empty, inference is not permitted and guessing is not permitted. That means that in that particular case, I made no conclusion whatsoever about any swimmer, team, meet or organisation. And it also means that if someone reads such a document and attributes a conclusion to it, the conclusion belongs to the attributor, not to the document.
This is where I differ from most sports writers. I will hand an editor a twenty-page file in which every cell reads "insufficient information", with a cover page stating that the problem lies in the data pipeline, plus three suggested fixes. Such a file is more useful than a commentary full of adjectives. It tells the system operators exactly where the leak is.
- SIGNALS TO WATCH IN THE NEXT ROUND
If you have read this far hoping for a list of swimmers about to shine, I must disappoint you. My tracking sheet currently holds no names, because the raw material has not arrived. But there are three signals I am waiting for.
The first is the return of a complete source dataset: article title, meet name, at least one named athlete, and three or more information points. When that happens, the nine-dimension framework can run again immediately, and I can return a genuine technical assessment instead of a null result.
The second is the source field. Not all sources are equal. A federation press release, a bylined report from a journalist at the venue, and an unsourced aggregation are three different levels of reliability. Once the source field is filled in clearly, I can grade the confidence of each information point rather than grading the whole article.
The third is time sensitivity. A number with no date cannot be placed in a cycle position, and therefore cannot take a tier-based discount. In swimming, where the four-year cycle shapes all training strategy, missing a date means missing nearly everything.
There is one question I leave open, not to answer now but to work on over the coming months: if a swimming nation is known only through televised moments, is that nation being measured by the audience's memory or by its own data? And if those two measures keep diverging, which one drifts first — the memory or the data?
I will answer that with a table, not with an essay. When there is enough material. Until then, the note in my file stays as it is: insufficient information, no conclusion.
