Identifying emerging trends early can give businesses a significant advantage.
Predictive analytics uses historical data to identify patterns and estimate what may happen next. Today, sophisticated machine-learning models can automate much of this process, but some of the most useful predictive techniques remain surprisingly accessible.
Regression, time-series forecasting and decision trees can all help analysts move from describing what has already happened to exploring what might happen next.
The real challenge, however, is not simply choosing a model.
It is understanding the assumptions behind it.
Forecasting Website Traffic with Historical Data
Earlier in my career, I was asked to predict how many visits a group of websites could expect to receive by the end of a quarter.
I had three years of historical behavioural data available.
Website traffic was not constant throughout the year, so simply calculating an average growth rate would have ignored an important part of the data: seasonality.
I therefore built a trend model adjusted for seasonal patterns.
At a high level, the process involved:
- Organising historical traffic into consistent time periods
- Identifying the underlying trend
- Separating recurring seasonal patterns from longer-term movement
- Calculating seasonal indices
- Applying those seasonal effects to the underlying trend
- Using the resulting model to forecast future traffic
The model performed well enough to provide a useful view of expected end-of-period performance.
But perhaps the most valuable lesson was not the specific forecasting technique.
It was learning to distinguish between signal and noise.
Trend Is Not the Same as Seasonality
Imagine that website traffic increases every December.
A forecasting model that ignores seasonality might interpret that increase as evidence of permanent growth.
But if traffic falls again in January, the December increase is not necessarily indicative of a change in the underlying trend. It may simply be a recurring seasonal pattern.
A useful time-series model, therefore,e tries to understand different components of the data:
Trend: the longer-term direction of movement.
Seasonality: patterns that repeat at predictable intervals.
Noise: variation that cannot easily be explained by the model.
Separating these elements can produce a more realistic forecast.
In my original analysis, I used a quadratic trend combined with seasonal indices in Excel. Today, analysts have access to far more sophisticated forecasting tools, but the fundamental questions remain the same.
What pattern exists in the data?
Is that pattern likely to continue?
And what assumptions are we making when we extend it into the future?
When Forecasting Becomes Guessing
Not every forecast deserves the same level of confidence.
I have encountered another approach that I jokingly called “gut modelling”.
The process looked something like this:
Start with the planned advertising spend.
Estimate how much paid traffic that investment might generate.
Estimate how many sales that traffic might produce.
Then use historical relationships to infer the expected organic traffic and additional sales.
The resulting model might look impressively detailed.
But the number of assumptions underneath it can make the forecast extremely fragile.
Marketing spend does not necessarily have a stable relationship with traffic.
Traffic does not necessarily have a stable relationship with conversion.
Organic demand may be influenced by brand awareness, competitor activity, economic conditions, seasonality and countless other factors.
If the model requires constant manual adjustment to remain plausible, the problem may not be that the seasonal index needs another tweak.
The underlying assumptions may simply be weak.
A Model Is Only as Good as Its Assumptions
This remains one of the most important lessons in analytics.
A sophisticated model built on poor assumptions does not become reliable because the mathematics is complex.
When evaluating a predictive model, ask:
- What assumptions does the model make?
- How stable are the relationships between the variables?
- How much historical data is available?
- Has the environment changed since that data was generated?
- Are important variables missing?
- How sensitive is the forecast to small changes in the assumptions?
- How accurate has the model been when predicting data it has not previously seen?
Forecasts should also communicate uncertainty.
A prediction of exactly 1,253,472 website visits may create an illusion of precision.
A realistic forecast is often better expressed as a range of possible outcomes, accompanied by an explanation of the assumptions that could move the result.
Benchmarking: Looking Beyond Your Own Data
Internal data tells you how your business is performing.
It does not necessarily tell you whether that performance is good.
This is where benchmarking becomes useful.
In one project, I analysed digital activity across several European markets to understand how a client’s online presence compared with competitors.
The analysis combined information including:
- Estimated advertising expenditure
- Advertising impressions
- Marketing-channel activity
- Technology platforms
- Email and newsletter activity
- Sponsorship and affiliate programmes
- Social media presence
- Brand indicators
- Broadband penetration
- E-commerce maturity
- Country-specific market characteristics
- Operational and supply-chain constraints
The objective was not simply to create a competitor ranking.
It was to understand performance in context.
A company with lower digital penetration in one market may not necessarily have a weaker strategy. The market itself may have lower e-commerce adoption, different customer behaviour or operational constraints.
Benchmarking becomes more valuable when it moves beyond:
“Who has the biggest number?”
and asks:
“What factors might explain the difference?”
Decision Trees: Understanding Patterns in Behaviour
Another useful analytical technique is the decision tree.
Decision trees divide a population into increasingly specific groups based on the characteristics that best distinguish different outcomes.
Imagine that we want to understand which visitors are most likely to click on an advertisement.
We might have information about:
- Advertisement format
- Placement
- Message
- Device
- Location
- Previous behaviour
A decision tree might first discover that placement creates the biggest difference in click-through rates.
Within one placement, device type might become the next most important distinction.
Within mobile users, a particular message might perform better than others.
The resulting tree creates a sequence of increasingly specific segments.
This makes decision trees particularly intuitive because the model can be visualised as a series of decisions.
Prediction Is Not Causation
There is, however, an important limitation.
A decision tree might identify that people exposed to a particular advertising format are more likely to click.
That does not necessarily prove that the format caused the increase.
Perhaps the advertisement was shown to a different audience.
Perhaps it appeared on higher-performing pages.
Perhaps another variable influenced both the exposure and the outcome.
Predictive models identify patterns.
Establishing causation usually requires stronger experimental evidence.
This distinction is critical when analytics is used to make business decisions.
From Predictive Analytics to AI
The tools available to analysts have evolved enormously.
What once required manually constructed Excel models can now be performed using statistical software, machine-learning platforms and increasingly sophisticated AI systems.
But better tools do not remove the need for analytical thinking.
In fact, as models become easier to build, understanding their limitations becomes even more important.
The fundamental questions remain remarkably consistent:
What data are we using?
What assumptions are we making?
What patterns has the model learned?
Will those patterns continue?
What uncertainty surrounds the prediction?
And what decision will we make as a result?
The purpose of predictive analytics is not to predict the future with certainty.
It is to use the evidence available to reduce uncertainty enough to make a better decision.
And sometimes, the most valuable insight a model can provide is not a precise prediction.
It is showing us what we still don’t know.
