Blog

The next era of AI in drug discovery

As AI moves from summarizing papers to generating scientific claims, biopharma faces a hard question: how do you trust an answer with no way to trace it back to the truth? Digital Science’s Mark Hahnel unpacks the shift at hubXchange 2026.

Keynotes & insights from AI in Drug Discovery hubXchange 2026

On September 9, 2026, leaders across the biopharmaceutical and artificial intelligence sectors gathered in San Francisco for the AI in Drug Discovery hubXchange. Digital Science’s Mark Hahnel delivered a keynote address examining how AI is altering scientific inquiry and biopharma R&D. Moving beyond basic task automation, Hahnel articulated a future centered on data provenance, agentic workflows and the assembly of an industry-wide data substrate.

Read on for a recap and key takeaways from Hahnel’s keynote.

Moving beyond ‘AI as a Tool’ to agentic workflows

Scientific research has entered the ‘4th paradigm’, where massive volumes of data must be readily available and actionable. However, the current velocity of AI adoption is accelerating at a rate that introduces friction into traditional organizational workflows. Rapid theoretical milestones highlight this momentum, such as OpenAI publishing proofs for the complex Navier-Stokes fluid mechanics equations on Twitter. The tension between independent mathematicians leveraging Codex to tackle similar challenges, and claims of unethical scooping, demonstrates the need for tools that protect IP while supporting research advancement.

Large corporations are embracing AI, transitioning from basic productivity aids to deploying AI for generating new scientific knowledge and discovering novel drugs. Room consensus at hubXchange aligned with statistics indicating over 50% enterprise adoption across major organizations. The biopharma industry has officially progressed past using isolated AI tools and entered the next phase: managing siloed data alongside specialized domain models and autonomous agentic workflows.

Offline & local models: safeguarding intellectual property

As enterprise adoption deepens, maintaining strict data security and protecting early-stage IP remain paramount. To prevent proprietary research from being inadvertently exposed or ‘scooped’ through public cloud chat windows, biopharma companies are prioritizing local models and offline data architectures.

Dedicated tools like Digital Science’s Papers AI address this demand by keeping enterprise data, local models, and analytical routines fully offline, ensuring researchers can leverage modern AI capabilities without compromising security.

The Provenance Crisis & Data Trust

As generative systems output claims at scale, biopharmaceutical organizations face a foundational challenge: Where did this data originate, and how was this specific claim validated? Grounding claims in verifiable truth is critical because core human facts reside outside the latent weights of Large Language Models (LLMs).

To achieve full traceability, organizations must back up every statement and derived insight. Hahnel highlighted the framework detailed in Digital Science’s FAIR data playbook for Pharma white paper as an essential roadmap for establishing structured, trustworthy data environments. Regulatory compliance necessitates adhering to the FDA + EMA Guiding Principles of Good AI Practice in Drug Development, which require tracking the explicit source, raw underlying data, and precise timestamps for all AI-assisted findings.

Constructing & deconstructing papers for machines

Building robust drug discovery models requires looking beyond high-level literature summaries. While cheap and accessible methods exist—such as using models like Claude to ingest titles and abstracts from PubMed—true drug discovery demands deep full-text extraction. Full text is essential to extract vital scientific nuances, including detailed methods, experimental edge cases, figures, and direct scientific contradictions.

Navigating this domain requires working within a fragmented publisher landscape, where the top 5% of publishers account for approximately 61% of all scientific publications. Existing pharma licensing agreements provide a pathway to deconstruct and reconstruct scientific papers into machine-ready structures, enhancing internal proprietary data.

Data must be structured once across workflows so it can be continuously reused rather than repeatedly extracted. Maintaining these comprehensive global databases requires continuous operational maintenance; for example, maintaining Dimensions‘ global patent database requires a workforce actively liaising with patent offices worldwide to correct inaccuracies and guarantee precision.

Four core techniques for structured data extraction

To derive locally verifiable statements and establish end-to-end data provenance, four primary computational techniques are being actively deployed:

  • Mapping: Leveraging LLM-driven ontology mapping to harmonize disparate scientific terminologies across domains.
  • Graphs: Building dynamic knowledge graphs that represent evolving biological relationships and entities.
  • Triage: Implementing just-in-time triage and filtering to parse incoming streams of scientific literature efficiently.
  • Extractors: Deploying agentic extractors designed to pull out claims, experimental methods, biological entities, and explicit relationships directly from full text.

These techniques allow organizations to extract claims and harmonize them so they are composable with internal proprietary extensions, Electronic Lab Notebooks (ELN), and existing R&D workflows.

Conclusion: the substrate is the work

The primary takeaway from Mark Hahnel’s presentation is clear: while foundational models and agent frameworks will continuously improve, the ultimate value lies in the data substrate beneath them. Grounded scholarly inference requires generating answers built upon verified external data seamlessly combined with internal enterprise assets.

Building this substrate requires deep collaboration across biopharma, biotech, academic publishers, and AI tooling providers. Models and agent frameworks will continue to evolve, but establishing the underlying, composable data substrate is the foundational work that the entire field must build together.

Ready to build the data substrate your AI strategy depends on? Explore how Digital Science’s enterprise solutions help biopharma organizations turn siloed data into trusted, structured, AI-ready assets.