Episode

đŸ”¬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)

Podcast
Latent Space: The AI Engineer Podcast
Published
Jul 21, 2026
Duration seconds
5387
Processing state
processed
Canonical source
https://www.latent.space/p/xaira
Audio
https://api.substack.com/feed/podcast/207941607/06ca3e4bd112d187109d3b25b2c9d8fc.mp3
JSON
/v1/public/podcasts/latent-space-ai-engineer/episodes/causal-models-need-causal-data-xaira-s-x-cell-model-for-drug-discovery-bo-wang-ci-chu-chief-discovery-officer-chief-ai-scientist
Markdown
/podcast/latent-space-ai-engineer/causal-models-need-causal-data-xaira-s-x-cell-model-for-drug-discovery-bo-wang-ci-chu-chief-discovery-officer-chief-ai-scientist.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/latent-space-ai-engineer/episodes/causal-models-need-causal-data-xaira-s-x-cell-model-for-drug-discovery-bo-wang-ci-chu-chief-discovery-officer-chief-ai-scientist/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/latent-space-ai-engineer/causal-models-need-causal-data-xaira-s-x-cell-model-for-drug-discovery-bo-wang-ci-chu-chief-discovery-officer-chief-ai-scientist.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Scaling AI for drug discovery requires moving beyond parameter count to information-rich causal data. Xaira Therapeutics' X-Cell model demonstrates that integrating high-throughput perturbation data allows models to predict how unseen cell lines respond to interventions.

Topics

  • Drug Discovery
  • Diffusion Language Models
  • Gene Expression
  • Single-cell RNA sequencing
  • Causal Inference
  • Bioinformatics
  • Xaira Therapeutics
  • Virtual Cell Models

Highlights

  • Main idea: Scaling laws for biological models are limited by data information content, not just compute or parameters
  • Technical shift: Moving from auto-regressive models (like scGPT) to diffusion-based architectures allows for better modeling of high-dimensional gene expression as an 'editing' process
  • Practical takeaway: High-throughput experimentation is essential to generate the causal datasets needed to predict gene expression changes after perturbations
  • Failure mode: Training on static datasets like CELLxGENE can capture cell states but fails to predict the dynamics of cellular interventions
  • Future frontier: The next breakthrough in 'virtual cell' modeling depends on longitudinal sequencing technology that can track the same cell over multiple time points

Chapters

  1. 1:00 Introduction to Xaira Therapeutics: An introduction to Bo Wang and Ci Chu and their mission to build an AI-driven drug discovery platform.
  2. 8:00 The Power of Predictive Models: Discussing the 'wow moment' when models successfully predict responses to unseen cell line perturbations.
  3. 15:00 The Shift to Data-Driven Interventions: Exploring how the rise of LLMs has influenced the approach to modeling biological interventions.
  4. 21:00 The Necessity of Causal Data: Why large-scale datasets like CELLxGENE are foundational but require causal context for true predictive power.
  5. 28:00 Scaling Perturbation Response: The challenges and opportunities in using high-throughput techniques to scale biological data collection.
  6. 42:00 Diffusion vs. Auto-regressive Models: A technical deep dive into why diffusion models are superior for modeling gene expression as an iterative refinement process.
  7. 55:00 Validating New Biology: Discussing the recent results from the X-Cell preprint and the discovery of new biological insights.