Why More Data Doesn't Necessarily Lead to Better Decisions
Why better decisions depend on understanding evidence, uncertainty and context, not simply accumulating more data.
We are extraordinarily good at collecting data.
We have data warehouses, lakes and lakehouses; streaming platforms; search engines; knowledge graphs; machine-learning pipelines; dashboards. Organisations gather records, logs, documents, sensor readings, customer interactions and countless other observations. Storage is cheap and computing power is plentiful. From this perspective, things have never been better.
Yet we still struggle to make sense of what we collect. The problem is increasingly not how to obtain data, but how to understand what the available data means for the decisions we actually face.
It is easy, perhaps even natural, to assume that more data must lead to better decisions. Sometimes it does. Additional data can reduce uncertainty, reveal previously invisible patterns and provide evidence for or against competing explanations.
But additional data can also introduce contradictions, duplicate what we already know, amplify biases, obscure important evidence beneath irrelevant detail or simply make us more confident without making us more correct.
What we need, therefore, are systems that help us reason about evidence rather than merely accumulate data. They should show where evidence came from, distinguish corroboration from repetition, expose contradictions, represent uncertainty, compare plausible interpretations and explain why a conclusion was reached.
Such systems should also be capable of admitting, “We cannot make a reliable choice,” or, more helpfully, “Here’s what we need to know to become more certain.”
The objective, after all, is rarely to own the most data; the objective is to make good, well-justified decisions.
Data is evidence, not an answer
Suppose we are trying to decide whether two records describe the same person.
One dataset contains a name and date of birth. Another contains a similar name, an address and an email address. A third contains a different spelling of the name and an organisation with which the person is associated.
The important question is not simply how many records we have, but what those records tell us. Do the names disagree because they refer to different people, or because of transliteration, abbreviation or error? Is the address current? How reliable is the source? Is the email address uniquely identifying?
Answering those questions requires us to interpret the records as evidence.
This distinction becomes increasingly important as datasets grow because observations are not necessarily independent. Ten sources repeating the same claim do not make it ten times more likely to be true if all ten obtained the claim from an unseen eleventh source.
The value of a dataset therefore depends not only on how much information it contains, but on the meaning and evidential value of the individual observations within it.
The question determines what matters
There is no universally “good” dataset. Evidence is useful in relation to a particular question.
Imagine that we have an enormous collection of information about a ship: ownership records, historical movements, port visits, company registrations, photographs, inspection reports and years of position data.
Whether that collection is sufficient depends entirely on what we are trying to decide:
- Where is the vessel now?
- Who ultimately controls it?
- Has it behaved unusually?
- Is this the same vessel mentioned in another dataset?
- Does it need further attention?
Each question requires different evidence, and each answer will tolerate different levels of uncertainty.
This sounds obvious, but data projects often proceed in the opposite direction. Organisations first assemble a large collection of data and only afterwards ask what can we do with all this?
A better starting point is more often, what decision are we trying to make, and what evidence would help us make it?
That change of direction — from data-first to decision-first thinking — can dramatically alter what needs to be collected and analysed.
Not all evidence deserves equal weight
Once we start thinking in terms of evidence rather than simply data, another distinction becomes important: not all evidence deserves equal weight.
A measurement from a calibrated instrument is different from somebody’s recollection. An official company filing is different from an automatically scraped web page. A contemporary observation may deserve more weight than one recorded ten years ago.
Even apparently objective datasets embody choices about what was measured, how it was categorised, what was omitted and how errors were handled.
Evidence therefore needs context.
In practical systems, that context might include provenance, recency, confidence, known error rates or an assessment of source reliability. In human analysis, it may simply mean being explicit about why one piece of evidence deserves more attention than another. Police and intelligence analysts, for example, routinely distinguish the reliability of a source from the credibility of the information that source provides.
The goal is not to pretend that every input is equally trustworthy. It is to make the differences visible enough that they can inform our reasoning.
More evidence can create more uncertainty
Even when we understand the relevance and quality of our evidence, gathering more of it does not necessarily make the answer clearer.
Imagine that three reliable sources agree about something. We might initially have quite high confidence in their conclusion.
Then a fourth reliable source contradicts them. We now have more evidence than we had before, but less certainty about what is true.
This is not a failure of the analytical process. The additional evidence has revealed uncertainty that was previously hidden.
Real-world data is full of such complications. Sources differ in reliability. Measurements contain errors. Information becomes stale. Categories and attributes do not line up neatly between systems. People change names and addresses. Organisations merge and divide. Sources that appear independent turn out to share a common origin.
A useful analytical system therefore needs to do more than accumulate evidence. It needs to preserve and represent uncertainty, disagreement and provenance.
Without that capability, additional data can create an illusion of certainty rather than genuine knowledge.
Sometimes the right answer is to collect more data
None of this means that gathering more data is undesirable. Sometimes more data genuinely does provide valuable additional evidence.
The important question is whether additional evidence is likely to change the decision.
Suppose two possible courses of action have very different consequences, and one relatively inexpensive measurement would strongly distinguish between them. Gathering that information may be extremely valuable.
Conversely, if another million records are unlikely to alter what we should do, collecting and processing them may add considerable cost without adding much decision value.
This gives us a more useful way to think about whether we need more data: not simply “Can we collect more?” but “What might we learn from collecting more, and could it change what we should do?”
Analysis can increase confidence without increasing correctness
The problem becomes more difficult still when powerful analytical tools are applied to the evidence.
Given enough data, we can nearly always find something that appears interesting. Machine-learning systems can identify subtle statistical patterns. Graph analysis can uncover complicated networks. Large language models can synthesise enormous quantities of text into remarkably coherent explanations.
These capabilities are genuinely useful, but sophistication of analysis is not the same thing as reliability of conclusion.
A beautifully visualised graph can be based on incorrect entity matches. A machine-learning model can learn a relationship that disappears when circumstances change. An LLM can construct a compelling explanation from contradictory or inadequate evidence.
Sophisticated tools can therefore make weak conclusions more dangerous because they can make those conclusions appear convincing.
Good analytical systems consequently need to ask not only:
What does the analysis suggest?
but also:
Why does it suggest that?
What evidence supports it?
What evidence contradicts it?
What assumptions does the conclusion depend upon?
How uncertain are we?
What would change our mind?
Those questions become more important, not less, as our analytical tools become more powerful.
Some uncertainty cannot be eliminated
Even a system that asks all of these questions will encounter uncertainty that cannot be resolved.
In a messy world, data rarely tells a completely unambiguous story. There will usually be uncertainty somewhere: in measurements, sources, relationships, interpretations or predictions about what happens next.
The realistic question is therefore not “How do we eliminate uncertainty?” but “How do we make a good decision while acknowledging the uncertainty that remains?”
That change of question encourages us to distinguish what we know from what we merely suspect, consider alternative explanations and think explicitly about the consequences of being wrong.
It also changes how we compare possible actions. A decision that remains sensible under several plausible interpretations of the evidence may be preferable to one that is optimal only if our favourite interpretation happens to be correct.
Uncertainty, in other words, is not necessarily something that must be removed before a decision can be made. It is something that the decision-making process itself needs to accommodate.
Better decisions, not bigger datasets
Data has enormous value. So do machine learning, AI, graphs, statistics and the other tools we use to extract information from it.
But they are means rather than ends.
The ultimate measure of an analytical system is not how much information it contains, how sophisticated its model is or how impressive its visualisation looks.
It is whether it helps somebody make a better decision.
Sometimes that requires more data. Sometimes it requires better data. Sometimes it requires recognising that several apparently independent pieces of evidence are actually the same claim repeated many times. Sometimes it requires understanding why sources disagree.
Sometimes it requires accepting that the available evidence simply does not justify the confidence we would like to have.
The difficult part is not accumulating information. It is working out what the evidence means, how certain we can be and, finally, what we should do next.