Blog: Finding signal in the noise - Why some forecasts improve and others don‘t

Humans have always tried to predict the future. During the hunter-gatherer era, people predicted things such as plant and food availability, weather, and animal movement. Today, with the rise of AI, prediction has entered a completely new dimension. We are no longer only making predictions ourselves but we are also teaching machines to make increasingly sophisticated predictions.

Figure A. Popularity of the search term prediction in Google Trends

With the rise of big data, one question becomes central for the art of prediction: How can we find the signal in this vast amount of data and use it to improve our predictions? As I was reading Nate Silver's book "The Signal and the Noise", I became more interested in this question. First, it would be good to define what a prediction is.

A prediction can be defined as a specific, testable claim about a future observable outcome.​ A prediction usually has four important attributes:

  • Target: What observable outcome are we predicting?
  • Horizon: By what date or during what time interval?
  • Probability: How confident are we?
  • Resolution: What evidence will tell us whether the prediction was correct?

Measurability is especially important because predictions improve when outcomes can be measured, scored, and used to update future forecasts.

When does prediction improve?

Prediction works best when the environment teaches the forecaster.

This is one reason why forecasting is much more successful in some areas than in others. Baseball and weather forecasting are good examples.

Baseball has relatively stable rules, many repeated cases, detailed measurements, and fast feedback. These features create a learning loop.

Meteorology also has similar advantages. Weather stations, satellites, radar, and other instruments provide frequent measurements of atmospheric conditions such as temperature, pressure, humidity, wind, and precipitation. Forecasts are produced frequently, and forecasters quickly find out whether they were right or wrong.

These conditions create a learning loop:

observe → model → predict → score → update

This is shown in Figure B. However, more observations do not automatically mean more useful signals. A pattern found in one dataset may simply be noise. Therefore, we also need to ask whether a pattern continues to work outside the sample that produced it.

Figure B. Learning Loop of Forecasting

Prediction versus forecast

Before moving to less successful prediction problems, it is useful to distinguish between a prediction and a forecast. A prediction is any statement about a future outcome.

A forecast is a prediction produced systematically using available data, models, and stated assumptions. Forecasts are also often expressed probabilistically. Therefore, all forecasts are predictions, but not all predictions are forecasts.

For example:

Prediction:

"Candidate A will win."

Forecast:

"Based on current polling data, Candidate A has a 65% chance of winning."

The second statement tells us not only what may happen but also how uncertain the forecaster is.

Why are earthquakes so difficult to predict?

Weather and baseball provide good learning environments. Earthquakes are almost the opposite. Long-term earthquake risk can be estimated probabilistically, but scientists cannot reliably predict the exact time, location, and magnitude of a future earthquake. There are four major reasons to this:

  1. First, large earthquakes are rare events. They happen infrequently on any individual fault, so there are relatively few historical examples available for learning and testing.
  2. Second, there are no reliable precursors. Proposed warning signals such as unusual tremors or changes in radon levels are inconsistent and may also occur without a major earthquake.
  3. Third, earthquakes involve complex fault dynamics. Stress develops across interacting faults, while rupture depends on underground conditions that cannot be completely measured.
  4. Finally, earthquakes have long validation cycles. Major earthquakes on the same fault can be separated by decades or even centuries. This makes models extremely slow and difficult to evaluate.

Nate Silver gives an interesting example to this shown in Figure C.

Figure C. The Precursor Trap of Earthquakes

In L'Aquila, Italy, a swarm of earthquakes occurred before a magnitude 6.3 earthquake in 2009. In Reno, Nevada, a similar-looking earthquake swarm occurred in 2008 but then disappeared without a major earthquake. Both situations initially showed a similar pattern, but they produced very different outcomes. This demonstrates an important forecasting lesson. A visible pattern is not automatically a useful prediction signal.

Why weather forecasting works better

We have had much more success with weather forecasting than earthquake prediction, although weather forecasting also has important limitations. Weather models begin with measurements of current conditions such as temperature, pressure, humidity, and wind. These measurements are highly detailed, but they are never perfectly complete. Because the atmosphere is chaotic, even very small differences in starting conditions can produce increasingly different outcomes as a model looks further into the future. Hence, weather forecasting improves over time, but accuracy still declines with lead time. Short-term forecasts work relatively well because the model begins close to the current state of the atmosphere. As we forecast further into the future, uncertainty grows.

Weather apps can also make good forecasts appear wrong because they compress information about probability, timing, location, and intensity into a single icon. For example, a 10% chance of rain does not mean that rain is impossible. It means that rain should occur in roughly one out of ten comparable cases for that location and time.

Forecasting can improve through feedback

Figure D. Historical NWS data

Figure D provides a good example of how continuous feedback improves forecasting. In the 1970s, the National Weather Service in the United States missed the high temperature by around 6°F on average when making a forecast three days ahead. Later, this average error fell to approximately 3.5°F. The grey points show the yearly forecast error, while the straight line summarizes the long-term improvement. So over several decades, the average error was almost cut in half. This is exactly what we would expect from a good learning environment:

predict → observe the outcome → measure the error → improve the next prediction

But if weather forecasts have improved so much, why do they sometimes still feel wrong? Part of the answer is that people misunderstand probability. Another part is incentives.

Calibration makes probabilities testable

Figure E compares historical rain forecasts from the National Weather Service with forecasts from local television. The x-axis represents the predicted probability of rain, while the y-axis shows how often rain actually occurred. The diagonal line represents perfect calibration. For example, if a forecaster gives a 60% probability of rain many times, rain should occur in approximately 60% of those cases. The historical National Weather Service forecasts were relatively close to the diagonal, meaning that they were approximately calibrated. The local television forecasts were more often below the diagonal, meaning that rain occurred less often than predicted.

This makes some sense from the perspective of incentives. If a TV weather reporter predicts sunshine and people are unexpectedly caught in heavy rain, viewers may be unhappy because they were unprepared. If the reporter predicts rain and it turns out to be sunny, the cost to the viewer is usually smaller. Therefore, historically, some local television forecasts had an incentive to overpredict rain. This example is historical and should not be interpreted as a statement about current local television forecasting. The larger lesson is we should judge probabilities across repeated forecasts, not from one outcome. If someone predicts a 10% event and it happens once, that does not prove that the forecast was bad.

Figure E. Variability of Forecasts

From weather to politics

Political forecasting is another interesting domain, but it creates very different problems. Polls measure current support, while forecasts estimate future election outcomes. A political forecast usually starts with polling data and then adds assumptions about turnout and uncertainty. The process can be simplified into six stages:

collect responses → weight the sample → estimate support → model turnout → simulate uncertainty → calculate win probability

For example, a candidate may receive a 70% probability of winning. If that candidate loses, this does not automatically mean the forecast was wrong. A 70% probability also means that losing remains possible 30% of the time. A well-designed poll can feed a poor forecasting model, and a good forecast can still lose once.

Why political forecasting is difficult

Political forecasting is difficult because both the data and the system change. Elections happen infrequently, and every election takes place under somewhat different conditions. Polling errors can also move together. If several polls miss the same group of voters, they may all be wrong in the same direction. Their errors therefore do not necessarily cancel each other out.

Political forecasts may also influence the system they are trying to predict. Published probabilities can affect media coverage, donations, turnout, and campaign strategy. This creates an important difference between weather and politics. A hurricane does not change direction because it saw a forecast. Voters and campaigns can react to one. Finally, historical patterns can stop working. Political coalitions, institutions, media environments, and information systems change over time. Therefore, unlike weather, the political system can react to the forecast itself.

The media rewards confidence, but forecasting rewards calibration

Political forecasting has another problem which is media incentives. We all know that the media likes sensation and rewards bold claims. But this can be almost the opposite of what good forecasting requires. The media often rewards certainty, simple stories, novelty, and recognizable personalities but a good forecasting needs probabilities, base rates, updating, and public track records. This creates an important paradox.

The qualities that make someone good television are not necessarily the qualities that make someone a good forecaster.

Nate Silver evaluated 733 predictions made on The McLaughlin Group, a long-running American political discussion show where journalists and commentators debated political issues and frequently made confident predictions. Of the 733 evaluated predictions, 338 were mostly or completely right, 338 were mostly or completely wrong, and 57 were mixed (shown in Figure F). Of course, The McLaughlin Group was also designed partly as entertainment. It came from an era when opposing political commentators often argued directly with one another on television. Today, the situation is slightly different. Instead of arguing on the same program, different political groups often have their own media spaces, and people may mainly hear opinions they already agree with.

Figure F. Analysis of McLaughlin Group's Predictions

Foxes and hedgehogs

This does not mean that experts are useless. Philip Tetlock, a psychologist and political scientist, studied whether experts could reliably forecast political and economic events. Beginning in the late 1980s, he collected predictions from academics and government officials on topics including international conflicts, economics, and politics. He found that experts were often overconfident and that forecasting performance depended less on credentials than on how people thought. This led to the distinction between foxes and hedgehogs.

Foxes are adaptable, self-critical, probabilistic, empirical, and comfortable with complexity. They are willing to use several different ideas and change their views when the evidence changes.

Hedgehogs are more specialized, confident, ideological, stubborn, and slow to update. They tend to explain the world through one major idea.

In the studies summarized by Silver, foxes were generally better forecasters, while hedgehogs were often more confident and more media-friendly.

The important lesson is good forecasters are not necessarily the people who know the most. They are often the people who change their minds best.

Can groups forecast better than individuals?

Tetlock's research focuses partly on individual reasoning. But another question is whether we can improve predictions by combining the judgments of many people. Aggregating diverse and independent forecasts can produce more accurate and better-calibrated probabilities. One example is a prediction market. A prediction market allows people to trade contracts based on whether a future event will occur. The resulting market price can be interpreted as a collective estimate of probability. Prediction markets therefore move us from individual judgment toward collective forecasting. However, aggregation works best when participants bring different information and have incentives to reveal what they genuinely believe.

Prediction markets can become less useful when participants all rely on the same information, copy one another, or when conflicts of interest distort incentives. So simply adding more opinions is not enough.

How can we become better forecasters?

So what can we do ourselves? First, we should evaluate probability forecasts across many cases, not judge them from a single outcome. A useful forecasting process is:

1. Define the outcome

What observable event will happen, and by when?

2. Check the base rate

How often has this happened in comparable situations?

3. Assign a probability

Express uncertainty numerically before knowing the outcome.

4. Update

Revise the probability when important new evidence appears.

5. Score the forecast

Compare the probability with what actually happened.

6. Check calibration

Among all events you predicted with 70% probability, did roughly 70% actually happen?

One unexpected outcome does not prove that a forecast was bad. Forecasting skill becomes visible across many recorded and scored predictions. Better forecasting is also trainable

Useful methods include probability and base-rate training, team collaboration, forecast aggregation, and performance tracking (shown in Figure G). The goal is not to eliminate uncertainty but to become better at recognizing it, measuring it, and responding to it. We cannot predict everything, and some systems are much harder to forecast than others. But we can improve our predictions by admitting uncertainty, measuring our mistakes, comparing ourselves with reality, and updating when the evidence changes.

Figure G. How to be better at forecasting

Forecast skill is trainable. The strongest supported interventions here are probability/base-rate training, collaborative teams that share rationales, aggregation, and performance tracking also shown in Figure G.​

REFERENCES

1- Can you predict earthquakes? | U.S. Geological Survey

2- Met Office, public forecast accuracy: https://www.metoffice.gov.uk/about-us/who-we-are/accuracy​

3- NWS, probability of precipitation: https://www.weather.gov/ffc/pop​

4- AAPOR, Polling Accuracy: https://aapor.org/polling-accuracy/​

5- The New Yorker, Why Political Pundits Are Becoming More Wrong: https://www.newyorker.com/news/daily-comment/why-political-pundits-are-becoming-more-wrong​

6- Philip E. Tetlock, Expert Political Judgment (2005).​

7- Mellers et al. (2014), Psychological Science 25(5): https://journals.sagepub.com/doi/10.1177/0956797614524255​

8- Mellers et al. (2015), Perspectives on Psychological Science 10(3): https://doi.org/10.1177/1745691615577794​

  • Written by Ful Belin Korukoglu

*This blog post is polished by ChatGPT to help refine the language, structure, and flow of a draft written by Ful Belin Korukoglu.*

Popular posts from this blog

AI Girlfriend or AI Boyfriend? Social Determinants of Human-AI Relationships

Showcase from the AI in the Global South seminar: Project #2 - Malam: A Low-Bandwidth AI Math Tutor for Afghan Girls

Job: Student Research Assistants (m/w/d)