Blog

Grounding AI Agents: literature vs. structured databases in the biopharma data stack

Sharing perspectives from the hubXchange 2026 Roundtable: “Assembling the data stack: what pharma needs from external knowledge in the age of AI Agents”

As biopharmaceutical R&D transitions toward autonomous AI agents, the industry is forced to re-examine the core substrate upon which these agents reason. A critical roundtable discussion hosted by Digital Science at the AI in Drug Discovery hubXchange 2026 addressed a foundational bottleneck: How do we balance unstructured literature with specialized structured databases to build an AI-ready data stack?

The consensus was clear: while LLMs excel at parsing the vast sea of academic literature, they frequently struggle with structured data, exposing a deep division between literature consumption and structured database integration. In this article, we recap key insights from the roundtable discussion, and propose what this means for biopharma companies moving forward. 

Unstructured literature: the promise and pitfalls of unstructured text

For many drug discovery teams, the primary value of Large Language Models (LLMs) today lies in their ability to conquer the sheer volume of academic literature. LLMs have dramatically optimized literature reviews by embedding text and measuring similar prompts across databases, removing the manual burden of reading endless papers.

However, relying strictly on LLMs introduces several acute pain points:

  • The ingestion bottleneck: While LLMs are excellent at reading text, they struggle to extract data from tables, images, and graphs. Roundtable participants highlighted that automated, validated methods to draw conclusions from tables—for example, answering, “Here are the top things you need to know in this table”—remain a critical visual gap.
  • The challenge of novelty: Experiments like Datasetpapers.com demonstrate that LLMs struggle to discover ‘unknown unknowns.’ True discovery requires explaining why a paper is groundbreaking or identifying novelty—a cognitive task where generative AI still struggles.
  • The content quality spectrum: Literature quality varies wildly. Curated papers from established publishers represent a far more reliable substrate than unvetted public repositories.
  • The missing negative results: A systemic bias exists where researchers are disincentivized from publishing negative results. In the AI age, dumping and organizing negative data must be simplified and incentivized (perhaps via shifts in H-index or S-index metrics) because machine learning models require negative data as much as positive data to build accurate classifiers.

‘Other’ databases: the cost and fragility of structured data

When moving beyond text consumption into computational biology and machine learning, biopharma relies on ‘other’ databases, such as structured sequence, mutation, chemical and clinical databases (e.g., UK Biobank, Gene PC, ADDI and PPMI for biomarker discovery validation).

These databases present a completely different set of challenges:

  • A scarcity of domain databases: There is a severe lack of structured databases in highly specific areas, such as amino acid mutations. Traditional structural models versus sequence-based models remain a continuous pain point.
  • The curation and funding crisis: Unlike the commercial publisher ecosystem, specialized public databases suffer from a chronic lack of ongoing funding. Representing this data, maintaining standards, and keeping it up to date is highly expensive. While large pharmaceutical corporations can absorb these costs, smaller biotech startups are effectively locked out.
  • The funding shift: Government and national lab funding is increasingly shifting away from pharmaceutical R&D toward materials and closed-loop discovery, leaving national standards (such as those from NIST) updated frequently but lacking deep biological curation resources.
  • Fouled public repositories: Without consistent curation and funding, public repositories easily become fouled with incorrect or mislabeled data. Because machine learning models are incredibly picky with data quality, many biopharma organizations are moving toward training models strictly with internal, highly-vetted data or avoiding public repositories altogether.

Harmonization: bridging the divide

The ultimate breakdown in the biopharma data stack occurs when attempting to harmonize unstructured literature findings with structured external databases and internal R&D data.

To make external data truly AI-ready, members of the roundtable discussed addressing several key operational requirements for their specific use cases and data needs:

  1. Defining clean data extraction standards: Clean data must go deeper than prose. In assay standards, data in the ‘methods’ section must be pristine and explicitly represent both positive and negative results.
  2. The metadata connection: Databases that rely on vague terms like ‘sample’ are highly problematic. Every raw read must be rigorously connected to downstream metadata, including omics data, patient profiles and study protocols.
  3. The missing donor ledger: A glaring gap in the current data stack is the lack of an interconnected, harmonized database of human blood and tissue donor identifiers.
  4. Re-identification risks: As AI agents cross-reference and harmonize independent datasets (e.g., matching blood donor IDs across omics databases), the risk of patient re-identification escalates, introducing complex legal and compliance hurdles.

Governance: build vs. buy and the role of knowledge graphs

Because external data sources are frequently unstructured, inconsistent, or lack verified provenance, biopharma organizations generally refuse to use external data for GxP-level decision-making and reporting. Instead, its use is confined to secondary exploratory reporting.

This reality forces organizations to ask: Is it worth purchasing external data and investing massive effort to prove its lineage, or is it more efficient to build it internally?

To navigate this landscape, organizations are leveraging two core architectural strategies:

  • Knowledge graphs: Tools like metaphactory, a Digital Science solution, are critical to building semantic ontologies, mapping disparate terminologies, and pointing autonomous AI agents directly to trusted data sources. They handle licensing, rights management, and legal boundaries.
  • Provenance machines: Tools like metaphactory’s metis act as black-box provenance machines, ensuring that every claim, entity, and relationship extracted by an agent is traceable back to its origin.

Systemic collaboration is needed

Ultimately, biopharma cannot rely on literature or databases in isolation. The future of AI-driven drug discovery depends on a composable substrate where unstructured literature discoveries are programmatically harmonized with highly structured external and internal databases. Building and maintaining this substrate is not a task for any single organization—it requires deep, systemic collaboration between biopharma, biotech, academic publishers, and AI tool developers to ensure the data powering tomorrow’s agents is reliable, traceable, and GxP-compliant.

Building an AI-ready data stack is a team effort. Talk to Digital Science about how metaphactory and metis can help.