Blog

In Life Sciences, data integrity is non-negotiable

Mark Hahnel

by Mark Hahnel
VP Open Research | Digital Science

Life Sciences data integrity blog - concept image

In the age of AI, research intelligence has to begin with trusted data. The Inside Our Data series explores the foundational data infrastructure that makes trust possible. 

The cost of an answer no one can explain

Enterprises rely on research intelligence to set and meet strategic objectives; intelligence is what makes an enterprise competitive. Research intelligence, as a concept, isn’t new. What is new is the enterprise’s necessary reliance on huge volumes of data—and on AI-driven insights and analysis which take that data as truth. 

Enterprises are moving fast to embed AI into research, analytics, and strategic planning. Far from the zeitgeisty pilots which were largely based on frontier model usage, AI-driven intelligence now comprises foundational infrastructure that informs how decisions are made. 

But for the advances and benefits this technology has already brought about, it has also brought risk. Research and AI-driven intelligence are only as trustworthy and defensible as the data they use, and not all data can withstand the necessary scrutiny of independent review.

This can pose an existential threat to research enterprises operating in regulated industries such as Life Sciences. As enterprises continue to evolve and embrace powerful new technologies, it’s more important than ever that their underlying data can stand up to audit. 

In this article, we’ll look at what happens when enterprises lack data integrity, what “trustworthy” data actually means in Life Sciences, and how the right infrastructure can fortify enterprise data in the world of AI.

What happens when the data underneath a scientific conclusion can’t be checked?

In 2020, two COVID-19 studies were published—one in the Lancet and one in the New England Journal of Medicine—using data from Surgisphere, a little-known analytics firm. 

But soon, there was a problem: Surgisphere refused to release its data for an independent audit. Both studies were retracted nine days apart.

The Lancet study claimed hydroxychloroquine increased mortality risk in COVID-19 patients—a finding that prompted the WHO to briefly pause a hydroxychloroquine arm of its global Solidarity trial before the retraction. The NEJM study was also retracted, but kept being cited long after: a Journal of the American Medical Association Internal Medicine analysis found 652 verified citations, with more than half of them occurring at least three months after the retractions took place. 

This incident illustrates what is at risk when data can’t be audited. We don’t know why Surgisphere wouldn’t release the data. Maybe it was all fabricated, maybe it wasn’t. There’s no way to know. But it doesn’t really matter. Data that cannot be audited is contagious. Bad data doesn’t stay where it started. It moves into papers, then models. Bad data has always been contagious. AI gives it a much higher reproduction rate.

It’s likely that the initial retractions were costly and frustrating for the firms who carried out the studies. But this incident also contributed to a wave of inaccuracy in critical research areas. Data that can’t be audited can halt clinical trials, knock percentage points off a stock valuation, cause lasting reputational damage, and most seriously, negatively impact the lives of real people. 

What “trusted data” really means

When it comes to defining what makes good data, enterprises aren’t starting from square one—Life Sciences has already formalized what “trustworthy” data means.

Regulators have relied on ALCOA—Attributable, Legible, Contemporaneous, Original, and Accurate—since the 1990s to assess data integrity in clinical and manufacturing contexts. More recently, guidance from bodies like the Medicines and Healthcare products Regulatory Agency and the World Health Organization extended this into ALCOA+, adding four further requirements: data should also be Complete, Consistent, Enduring, and Available. It’s a checklist built around whether a record can be verified after the fact, and it’s still required for trustworthy data today.

The related framework, FAIR—Findable, Accessible, Interoperable, and Reusable—addresses a different but equally consequential point of failure. Introduced in 2016, FAIR has had a substantial impact on how research-generating organizations think about their data: not just whether it exists, but whether it can be found, retrieved under clear terms, and reused with confidence in its provenance. FAIR doesn’t require data to be open to all—a dataset behind a paywall or access agreement can still be fully FAIR-compliant. However, in order to satisfy the requirements of the framework, it needs to be made available, in a FAIR manner, to the people who would make assertions on that data. 

The Surgisphere retractions occurred because the underlying datasets were non-compliant with these frameworks. The data wasn’t available or accessible for audit, which meant we also couldn’t know if it exemplified the necessary integrity that made it suitable for use in research.

These frameworks comprise a non-negotiable baseline for Life Sciences research enterprises, but there remain grey areas which can have unintended effects on the quality of research datasets. For example, an open dataset which is seemingly FAIR and ALCOA+-compliant could be skewed toward whichever countries or funders proactively volunteer their data. In a regulated environment, this isn’t enough; passing an ALCOA+ or FAIR checklist doesn’t tell you whether a dataset is representative—curation, applied on top of these frameworks, can correct for that skew. 

Data infrastructure designed to accommodate investigation

Data curation refers to the ongoing process of ensuring that data is complete and representative—a process that requires human judgment and relationships to execute. This is a foundational tenet of the datasets which comprise Dimensions by Digital Science, one of the world’s largest research and funding data repositories.

Dimensions was built around the idea that research intelligence is only useful if it can be traced across the full lifecycle it describes—not only publications, but the funding, trials, patents, and policy activity that surround them. Dimensions datasets span six linked content types: more than 165 million publications, 8.2 million grants, 74 million research datasets, 180 million patents, 976,000 clinical trials, and 2.5 million policy documents, all cross-referenced with the others. 

A model surfaces a promising area of research. Don’t just take the answer. Ask:

  • Who funded it?
  • Which researchers produced it?
  • What publications followed?
  • What datasets underpin them?
  • Were patents filed?
  • Did it progress into clinical trials?
  • Did it influence policy?

This structure enables attributability and originality under ALCOA+: publication records are enriched through full-text indexing and linked back to direct publisher partnerships, Crossref, PubMed, and other authoritative sources. The grant data comes from more than 700 funders worldwide, sourced by data experts directly from funder organizations wherever possible. Clinical trial records are pulled directly from official registries spanning every major region, so status, sponsors, and outcomes reflect the authoritative record rather than a secondhand summary. Patent data is provided by IFI Claims, curated and normalized by Digital Science teams.

Every record carries a persistent identifier and a link back to its original source, so a grant, publication, or patent is Findable and its provenance is never in question. Records are Accessible under clear, documented terms—whether that’s open data or a governed connection through a licensed platform, so users always know what they’re looking at and where it came from. And the cross-referencing between content types is what makes the data Interoperable and Reusable in practice: a grant can be traced through to the publications it funded, the datasets and patents those publications generated, and the clinical trials or policy documents that followed. Research across more than 100 countries and every major discipline reduces the blind spots that come from a literature-only view or a single-region dataset. This is data that can be audited—and that enterprises can trust to drive the decisions they make. 

The final step to unlocking truly powerful and trustworthy intelligence is ensuring this data infrastructure is in sync with enterprise-specific ontologies. Pairing trusted data with semantic definitions lays the groundwork for life sciences enterprises to more safely rely on AI-driven research intelligence in the years to come.  

Building trusted foundations for future AI implementations

The Surgisphere debacle exemplifies the failures that AI-assisted workflows now risk automating at scale: fluent, confident outputs based on data that doesn’t meet industry standards. Today, AI-assisted workflows are increasingly embedded in how R&D and Medical Affairs teams triage literature, surface signals, and make decisions. The efficiencies and insights to be gained from this technology are unprecedented, but this also raises the stakes: an AI working from ungoverned data doesn’t just produce a bad answer, it can introduce existential risk. 

The best way to guard against such a failure—and set your company up for long-term success—is to take a two-pronged approach, pairing an enterprise-specific semantic layer, such as a knowledge graph, with data that is FAIR and ALCOA+-compliant. Digital Science offers knowledge graph infrastructure designed to grow with an enterprise via its proprietary technology, metaphacts

By defining a semantic layer, an enterprise sets the scope for the data that AI is able to access and defines the logical relations between defined entities. This means the model can only interpret the data it is given access to in the context of an approved series of rules. This mitigates the risk of hallucinations or logical failures, and makes it simple for auditors to interrogate the pathways that led to a certain output. This is how to ensure trusted data is treated predictably by trusted models.

The result is accurate intelligence with an in-built audit trail that enterprises can trust to stand up to independent audit. 

The bar for data integrity will keep rising

As AI becomes more embedded in R&D and Medical Affairs decision-making, so too will audits by regulators and internal stakeholders. Trusted intelligence starts with trusted data, and trusted data is best used in sync with foundational enterprise infrastructure.

Digital Science provides one of the world’s broadest collections of connected research intelligence—combining Dimensions, Altmetric, and IFI Claims to help enterprise organizations support analytics, strategic decision-making, innovation, and AI workflows that can be explained, audited, and defended.