Before the Data Contract

bstract navy and teal illustration showing research needs and clinical concepts flowing into a structured data contract.

What clinical cohort discovery taught me about data collection governance

I recently found myself thinking about what should have been a relatively straightforward data requirement.

Once a participant’s sample had been sequenced, we needed to identify and bring into the research environment the primary clinical data that researchers would use alongside the genomic data.

The obvious next step was a data contract between the team supplying the clinical data and the team responsible for processing it. My first instinct was to use the contract to agree on the data structure and the delivery method. But that raised a more basic question: what did researchers actually need the data for? Until we answered that, we could not know what the contract should contain. More specifically, what was the minimum primary clinical information they needed to identify relevant participants and build a potential cohort?

That shifted the problem for me. We were not simply defining how data should pass from one team to another. We first had to decide which clinical information researchers actually needed, why it belonged in the research environment and how we would know whether it was sufficient for cohort discovery.

The source data did not define the need

The source team had already collected a large amount of primary clinical data through various forms, pathways and operational processes. We could have started with those existing tables and decided which ones to move into the research environment.

That would have been the easiest approach to describe in technical terms. It would also have been supply-led.

Starting with the existing fields would have allowed the source system to shape the requirement. Some fields were available because they served an operational purpose, not because researchers needed them. Other important clinical concepts were missing or collected inconsistently. We risked agreeing on the mechanics of the transfer before knowing whether the data could support cohort discovery. These were necessary decisions, but they could not tell us whether the data would meet researchers’ needs.

I needed to turn the problem around and start with the researcher.

The user research reinforced this. Researchers described difficulty knowing where to start, determining which data was relevant and connecting related tables, fields and identifiers. Information was organised around datasets and systems, while the meaning, provenance, codes, limitations and relationships they needed were often found elsewhere. They wanted to approach the data through a research concept or task, rather than having to understand its structure before using it.

What does a researcher need to build a cohort?

Researchers do not begin by asking which source tables are available. They begin with a research question and try to identify participants who meet a set of clinical and genomic criteria.

At a general level, they may need to ask:

  • Which participants have a particular condition or clinical indication?
  • Which participants share a particular phenotype?
  • Which cancer participants have a specific tumour type or disease subtype?
  • At what age, or at what point in time, was the condition identified?
  • Which participants are affected or unaffected within a family?
  • Which potentially relevant participants also have genomic data available?

This gave us a more useful definition of the requirement:

Collect the minimum primary clinical data researchers need to identify and define a potential cohort among participants with genomic data.

The word minimum did not mean selecting the fewest possible fields. It meant identifying the smallest set of reliable, well-defined clinical concepts, together with enough context to interpret them. What counted as the minimum, therefore, depended on the research questions and the clinical pathway. There was no useful minimum in the abstract.

I had to think in concepts, not columns

Once I looked at the problem through the lens of cohort discovery, I saw that we should not move directly from researcher questions to source fields.

We first needed to identify the concepts behind those questions.

Some would be common across different types of research: a pseudonymised participant identifier, the programme or pathway through which someone was recruited, the clinical indication for testing, a diagnosis, an associated date or age, and the provenance and status of the record.

Other concepts would depend on the disease area.

For rare-disease research, cohort discovery may require phenotypic features, affected status, age of onset, proband status and family relationships.

For cancer research, it may require tumour site, morphology, disease subtype, diagnosis date, stage or grade.

This also showed me why a single universal list of minimum clinical fields could be misleading. Data sufficient to identify a broad rare-disease cohort may not be sufficient for a specific cancer study. At the same time, collecting every disease-specific field for every participant would create more data to process, assure and explain without necessarily adding value.

Choosing the right fields was only part of the problem. Researchers also had to know how to interpret them and connect them to other records. That meant keeping definitions, codes, provenance and known limitations close to the data. Otherwise, a field could be available without being usable with confidence.

The comparison with the source exposed the real gaps

Only after identifying the clinical concepts could we assess what the source data could provide.

That comparison was more revealing than a schema review alone. A concept might not have been collected at all, or it might have been available for one clinical pathway but not another. Several fields could represent the same concept in different ways. Information needed for filtering might have been captured as free text, or recorded at recruitment without being updated later.

Missing data presented another problem. The same empty field could mean that the information was unknown, not applicable or never collected. Those differences mattered if researchers were going to use the field to include or exclude participants.

These are sometimes treated as downstream data-quality or mapping problems. In practice, some originate at the point of collection.

We could rename fields and correct data types downstream, but we could not recover information that had never been captured. If “unknown”, “not applicable” and “not collected” were all recorded in the same way, no transformation could reliably separate them. Nor could it tell us whether an older value still reflected the participant’s current clinical position.

This was the point at which the work moved beyond data collection. We had to decide which gaps could be addressed through transformation, which required changes to upstream collection, and which should remain visible limitations of the research data.

Once we had separated the research product’s needs from the source’s limitations, the data contract had a much clearer job to do.

The data contract came next

The data contract still had an important role, but it came after the purpose, user questions and minimum concepts had been established.

It could then turn those decisions into an agreement between the team producing the primary clinical data and the team processing it for research.

The contract needed to say more than which fields would be delivered and in what format. It had to define the purpose of the data, the participants and clinical pathways in scope, and how the source fields represented the clinical concepts researchers needed. It also had to make the less visible details explicit: how codes should be interpreted, what different kinds of missing values meant, what level of completeness we expected and how corrections, withdrawals and changes to the collection process would be handled.

This distinction between a clinical concept and its physical representation became particularly important to me.

A researcher may need to identify participants with a particular diagnosis. The source may represent that diagnosis through several fields, forms or codes. A contract that documents only the physical fields transfers the interpretive work downstream. A useful contract needs to preserve both the meaning and the structure.

Quality expectations also need to reflect how the data will be used. Fields required to identify and link a participant may need strict validation. A more detailed clinical characteristic might be optional, but researchers still need to understand its coverage before using it to exclude participants from a cohort.

Absence needed special attention too. Phenotype data made this particularly clear. A blank field did not tell us whether a feature was absent, unknown or simply never recorded. Treating all three as negative findings could exclude relevant participants before the research even began. Quality requirements therefore need to reflect the decisions researchers will make with the data, not only whether a field passes technical validation.

A contract cannot provide all the governance

Working through this example helped me see the boundary of a data contract more clearly.

The contract can govern the interface between the producer and the consumer. It can make the expected data testable and changes more controlled. But it cannot replace the decisions that surround it.

Some of those decisions had to be made before the data arrived. We needed to be clear about the research needs we were supporting, what information was necessary and proportionate, and what “minimum” meant for each clinical pathway. We also had to decide whether gaps could be documented as limitations or were significant enough to justify changes to the collection process.

The decisions continued after ingestion. Passing validation did not tell us whether the quality was changing over time, whether researchers could see the limitations or what should happen when a clinical definition changed. We also needed to know whether the fields were still being used for the purpose that had justified their inclusion in the research environment.

Without these decisions, a data contract can keep a data flow technically stable while the data becomes less useful, or remains unsuitable, for the problem it is meant to solve.

The pattern extends beyond clinical data

I encountered this problem through clinical and genomic data, but it is not specific to healthcare. In any domain, it is easy to mistake what a source system contains for what a user actually needs.

Customer, financial and operational data products face the same risk. The requirement should begin with the decision or task the data needs to support, followed by the concepts and context required to support it. Only then should the available source fields determine how to deliver that requirement.

What I would do now

The sequence I would use is:

  1. Define the questions users need to answer.
  2. Translate those questions into the minimum set of concepts and supporting context.
  3. Assess how well the source captures those concepts.
  4. Agree what should be provided and what limitations must remain visible.
  5. Formalise the exchange through a data contract.
  6. Validate and monitor the data against its intended use.
  7. Refine the collection based on user feedback and changing needs.

This does not make schemas, pipelines or data contracts less important. It gives them a clear purpose.

My starting assumption was that we needed a contract to bring primary clinical data into the research environment once genomic data became available. What I came to understand was that the contract could only be meaningful after we had defined the research outcome it was there to support.

The goal was not to move as much clinical data as possible. It was to provide enough trustworthy, well-defined information for researchers to find the right participants, while being honest about what the data could and could not tell them.

For me, that is the difference between collecting data and governing its collection. The data contract mattered, but the question it was meant to support had to come first.

Comments

Leave a Reply

Discover more from Data with Purpose

Subscribe now to keep reading and get access to the full archive.

Continue reading