The Economics of Data Products in the Age of AI: The Costs We See and the Costs We Shift

Iceberg illustrating the hidden costs of data products, with engineering and infrastructure above the waterline and documentation, integration, support and interpretation beneath the surface.

When organisations think about the cost of data products, they usually start with the cost of producing them.

Engineering time. Infrastructure. Data ingestion. Quality checks. Documentation. Governance. Support.

These costs matter, but they are only part of the economics of a data product.

A product that is relatively inexpensive to maintain can still be costly to use. Poor documentation, unclear definitions, duplication and difficult integration all create work, often outside the team responsible for the product.

The cost has not disappeared. It has simply moved.

The internal cost of a data product

Every data product consumes resources throughout its lifecycle.

A new dataset may require engineering effort, quality assurance, documentation, governance and specialist knowledge. Once available, it may continue to require monitoring, updates and support.

These costs are easy to see because they appear in team backlogs, infrastructure bills and resource plans. A new pipeline requiring several months of engineering work attracts scrutiny. An existing dataset requiring occasional maintenance may appear almost free.

But the cost of maintaining a product is not the same as the cost of having it in the ecosystem.

The costs we transfer to users

Imagine a researcher trying to answer a question using several data assets.

One has limited documentation. Another contains similar information but uses different definitions. A third has not been updated for several years, although this is not immediately obvious.

The researcher has to determine which source is appropriate, how the datasets relate to one another and whether their limitations affect the analysis.

None of this effort appears in the cost of running the data platform. It is still a cost.

If one hundred users independently spend time understanding the same poorly documented asset, the organisation has not avoided the cost of documentation. It has distributed the cost among one hundred users.

Investing more in documentation, standardisation, or product design may therefore increase internal costs while reducing a much larger burden elsewhere.

Looking only at the cost of maintaining a data product misses part of the picture. The effort required to understand and use it is part of its cost, too.

AI changes the economics

AI can make data quicker and cheaper to work with. Dataset discovery, documentation, schema mapping and user support can increasingly be assisted by AI.

But easier access does not necessarily make data easier to interpret correctly.

Consider a clinical dataset that records a diagnosis date, but where the term was defined differently across contributing data sources. One recorded the date of first symptoms, another the first clinical coding and another the date of specialist referral. The ambiguity is documented, but only in metadata that many users may never open.

Before AI-assisted tools, a researcher writing a query might consult a colleague, read the data dictionary or pause at an unfamiliar field. That friction could sometimes expose the ambiguity.

An AI assistant may not pause. Asked to calculate time from diagnosis to treatment, it could use the field as if its meaning were consistent and produce a confident, well-structured answer.

The problem with the data was already there. AI does not create the ambiguity, but it can make it easier to miss and much quicker to repeat. The same interpretation might find its way into several queries, reports or models before someone goes back to the source and questions what the field actually means.

That gives organisations another reason to invest in clear definitions, useful metadata and provenance. These may not be the most visible parts of a data product, but they become increasingly important when users are no longer the only ones interpreting the data.

More data does not always mean more value

The first datasets made available for a particular problem may create substantial value. As more are added, however, the additional benefit of each may diminish.

Imagine a platform that already provides several years of hospital activity data. Adding another source covering much of the same population and information may still be useful, but perhaps less useful than improving the quality, coverage or documentation of what is already there.

This is where marginal value matters.

The question is not whether a dataset has some value. Most probably do. What matters is whether the additional value it creates is worth the cost of bringing it into the portfolio and supporting it over time.

Sometimes the better investment is not another dataset. It may be improving documentation, addressing quality problems, reducing duplication or connecting data that users currently have to combine themselves.

Fewer well-designed, well-understood products may create more value than a much larger catalogue that leaves users to make sense of the complexity on their own.

The opportunity cost of maintaining everything

Resources spent on one data product cannot be spent elsewhere.

Engineering time spent maintaining a rarely used asset is unavailable to improve a heavily used one. Specialist expertise spent explaining poorly understood legacy data cannot be used to design better products.

Maintaining the status quo can feel like a neutral decision. It is not.

Every product in a portfolio competes for finite capacity, infrastructure, governance attention and specialist knowledge.

The question is therefore not simply:

Does this data have value?

It is:

Is this the best use of the next unit of investment?

Looking at the total economics

The economics of a data product can be considered through three lenses:

Internal cost: what does it cost to produce, maintain and support?

User cost: how much effort does it require to discover, understand, trust and use?

Opportunity cost: what else could be achieved with the same resources?

AI does not remove these trade-offs. It may reduce the cost of working with data while making it easier for poor assumptions to travel further.

The objective should not be to minimise the internal cost of a data platform or maximise the number of assets it contains.

It should be to invest where the value created is greatest relative to the total system cost.

Sometimes that means adding new data. Sometimes it means improving what we already have. And sometimes it means deciding that maintaining everything is no longer the best use of limited resources.

Comments

Leave a Reply

Discover more from Data With Purpose

Subscribe now to keep reading and get access to the full archive.

Continue reading