Introduction
Over the past several years, we've positioned ourselves at the intersection of machine learning and biology to answer questions that data alone could not. We asked whether a protein sequence will be biologically active. Why some tumors resist treatment. What a patient’s histological slide reveals about prognosis. How regulatory DNA encodes what a cell does. These questions had data but no answers, and getting them right could change how we detect and treat disease.
Learning the Language of Proteins
Most known proteins have not had their functions experimentally verified, and measuring whether a protein folds, binds, or catalyzes (fitness) requires slow, expensive wet-lab experiments. Protein engineers face a vast space of possible sequences but can test only a tiny fraction. If we can computationally predict which sequences are worth making, we can engineer new therapeutics, enzymes, and materials.
The core challenge is predicting fitness. However, existing benchmarks were limited in scope and consistency. Thus, FLIP focused specifically on fitness prediction across diverse proteins and FLIP2 expanded tended FLIP to more proteins, tasks, and criteria. On the modeling side, Transformers dominated protein language modeling by 2022. We showed that for fitness prediction, convolutions matched their performance at a fraction of the cost, suggesting architecture mattered less than the field assumed.
These foundational results opened two paths for computational protein design: optimization and generation.
For protein engineers iterating on an existing design, active learning methods cut required experiments dramatically, prioritizing variants so each wet-lab round yields maximum information. Rigorous uncertainty quantification showed that no single method dominates across fitness landscapes, and that the best approach depends on how far you're extrapolating from known sequences. For researchers who wanted to design new proteins from scratch, we developed diffusion models that offered a way to generate candidates that did not exist before. FoldingDiff generates backbone structures using a geometric representation that respects how proteins fold in nature. EvoDiff works directly in sequence space, enabling both traditionally structure-driven tasks like scaffolding and the design of disordered proteins that structure-based methods miss.
These threads came together in the Dayhoff Atlas: 3.34 billion natural sequences, 83,000 novel synthetic folds, and language models trained on evolutionary context. We validated the models experimentally and found that data composition matters as much as architecture. Structure-based augmentation nearly doubled the fraction of designed proteins that cells successfully produce (expression yield). CleaveNet showed that our sequence-driven approach to protein language modeling works for targeted applications, designing efficient, selective protease substrates that we validated in the wet lab.
Together our past and current work lays the computational foundation to model protein function directly – from sequence to function, and from function to sequence.
Decoding the Regulatory Genome
The genome encodes instructions for when, where, and how genes are made but interpreting that code remains a fundamental challenge. Most regulatory elements lack functional annotations, and experimental characterization is slow. We've developed computational approaches that use biological signals to learn what the genome encodes to ultimately design new regulatory sequences and control downstream function.
Standard genome language models are pretrained with sequence reconstruction objectives borrowed from natural language processing but often fail to capture biological signal. We addressed this by adding evolutionary-rate prediction as a pretraining task, producing representations that outperform sequence-only training and making even relatively small architectures competitive with much larger existing genome language models. This insight, that evolution encodes function, also applies to sequence design. EnhancAR uses evolutionary conditioning on homologous enhancer sequences to generate functionally diverse enhancers across 1,888 cell types without requiring labeled data from massively parallel reporter assays. These results point toward a general principle: evolutionary history is an underused training signal that encodes functional constraints the genome has already solved. Prior biological knowledge can ground models in other ways as well.
Our work suggests that building biological knowledge into model structure can help models support hypothesis testing rather than prediction alone. Clinical interpretation depends on the same capability. Interpreting which variant may explain a patient’s condition requires evidence scattered across thousands of publications, and genetics professionals identified assembling that evidence, not calling the variant, as the slowest step. We built the Evidence Aggregator to help genetics professionals assemble evidence for rare-disease interpretation more efficiently, extracting variant information and associated clinical features from the literature for any human gene. Evidence also changes over time, creating a need to revisit genomic data as gene-disease and variant knowledge develops. We built Talos to automate this repeated reanalysis. In a large-scale research evaluation, Talos surfaced findings that contributed to hundreds of diagnoses.
Together, this work reflects a broader view of genomic modeling: understanding the genome requires not only learning from sequence, but grounding models in the biological and scientific knowledge through which that sequence acquires meaning.
Understanding Cancer Cell State
Why does the same cancer drug work in one patient and fail in another? In partnership with the Broad Institute, we are investigating whether a cancer cell’s state provides predictive power for precision oncology, beyond only looking at the genetic mutations of tumors. We have profiled metastatic tumors and matched experimental models, called organoids, at single-cell resolution. We found that cell state, shaped by microenvironment, predicts drug sensitivity. That is, even in the background of the same genetic mutation, cancer cells respond differently to treatments depending on where the cells sit and what signals they receive. Tumor mutation status is, on its own, not sufficient to predict outcomes.
We took this insight and launched Ex Vivo, a joint research program with the Broad Institute. To characterize cell states at scale, we had to build our own methods. Off-the-shelf foundation models didn't work for single-cell tasks, and fine-tuning wasn't enough. We developed clustering methods that scale to millions of cells, calibration techniques to prevent overclustering, training algorithms that respect the complexity of biology, and adaptive resampling strategies that improve representation quality for underrepresented cell states.
These computational advances sit alongside an experimental pipeline that enabled drug screening at ten times the usual throughput using pooled perturbations.
In Ex Vivo we bring our computational and experimental approaches together to learn, engineer, and target cancer cell states, going beyond mutations to define a constructionist paradigm for precision oncology.
Reading Cancer in Tissue
When pathologists analyze tissue slides, they can detect cancer, determine its subtype, and assess its severity. But some questions are harder: Which patients will respond to immunotherapy? Which tumors will recur? The answers may be encoded in the tissue or could be revealed when combined with complementary data.
Until recently, computational pathology models were diagnostic-specific and trained on narrow samples. Existing approaches could classify common cancers but struggled with rare ones. Better patient outcomes require methods that generalize, predict genomic biomarkers from images alone, and describe findings in clinical terms.
In collaboration with Paige.ai, we scaled training data and model capacity. Virchow, trained on 1.5 million slides from 100,000 patients, achieved 0.95 AUC in cancer detection across nine common and seven rare cancers, matching specialized clinical-grade models with far less labeled data. Virchow2 expanded to 3.1 million slides; its larger variant, Virchow2G (1.9 billion parameters), set state-of-the-art performance on twelve tile-level benchmarks. Combining information from multiple zoom levels improved predictions for tumor regions, capturing both broad context and cellular detail.
We then closed the gap between how foundation models see images and how pathologists work. PRISM extended models from tiles to whole slides and PRISM2 went further. Trained on 700,000 diagnostic reports and 14 million question-answer pairs, PRISM2 describes findings in clinical language, producing outputs pathologists can interpret directly. Predicting new genomic biomarkers normally requires years of prospective data collection. MARBLE's multimodal pretraining sidesteps this by integrating biomarker knowledge into image representations, enabling prediction of new markers immediately without marker-specific training data. With NOVA, we addressed a different challenge: turning histopathology research questions into executable analysis pipelines. We paired the agentic framework with SlideQuest, a benchmark for evaluating these systems.
Together, these form a pipeline from raw tissue images to clinical-language reports, with the flexibility to predict new biomarkers as they're discovered.
Seeing Cells Clearly
Microscopy-based screening shows promise as a scalable approach to drug discovery. Researchers treat cells with thousands of compounds, image them, and look for visual signatures of drug effects. The challenge is turning pixels into biological understanding. Models are prone to learning shortcuts, patterns that correlate with the answer in one dataset but don't reflect underlying biology and won't generalize to new experiments. Technical noise from experimental batches, contamination from neighboring cells, inconsistent evaluation methods, and ignored chemical context all create opportunities for these shortcuts. We have built approaches aimed at addressing each of these challenges.
Technical differences bring noise and confounders to microscopy datasets. Every experimental plate looks slightly different due to technical variation, and models learn to recognize plates instead of biological phenotypes. A sampling strategy that draws each training mini-batch from the same plate lets normalization layers estimate and remove this variation, achieving state of the art on a standard benchmark without any architectural changes. Neighboring cells create another problem. When you crop a single cell for analysis, other cells in the frame can leak information into the learned representation. Swapping background cells in yeast experiments dropped accuracy by up to 15.8%, showing that models partly learn neighborhood context rather than the target cell's phenotype.
Fragmented evaluation and datasets made methods hard to compare. Previously, each new method was tested on different tasks with custom metrics. MorphoHELM created a unified benchmark for Cell Painting that tests methods across varying levels of batch effects. The results challenged conventional wisdom. No deep learning method consistently beat classic hand-crafted features across all settings. scGeneScope paired 627,000 transcriptomic profiles with 716,000 Cell Painting images under matched drug treatments and showed that combining modalities helps identify mechanisms of action. Multiprofile integration consistently improved performance, but the benchmark also revealed that recent single-cell foundation models, used without fine-tuning, underperformed classical methods trained on the target data.
With reliable evaluation in place, we could build better computational methods for microscopy data and downstream tasks in drug discovery. By treating chemical compounds as treatments that induce phenotypic transformations, rather than ignoring them or simply aligning images with molecular structures, we developed MICON, which performs better across datasets and institutions, outperforming standard contrastive approaches. Vermeer takes a different approach. Given a protein sequence and landmark morphology stains, it synthesizes fluorescent microscopy images predicting where that protein would localize in the cell. Trained on the Human Protein Atlas, Vermeer generalizes to unseen proteins and cell lines with substantially improved biological fidelity, enabling researchers to predict localization for proteins not yet experimentally imaged.
Making Biological Data Learnable
Computational methods typically rely on clean, labeled examples. Biological data is messier. A tumor is a distribution of cell states. Slides from different hospitals look different even when the diagnoses are the same. Human datasets also reflect variation shaped by genetics and environment, which can affect what models learn and how well they generalize. Single-cell datasets can have billions of measurements. Standard approaches that assume more data means better performance often fail in our context.
We found that diversity matters more than raw size. Adding more data from the same source can degrade performance on specific cell types. In a study of 400 pretrained model architectures across 6,400 experiments in single-cell biology, performance plateaued earlier than expected, with no clear scaling laws. Current approaches to predicting cellular perturbations often fail to generalize when the underlying causal mechanisms differ across contexts. We also found that immune-cell responses can be separated into shared donor-perturbation states and cell-type-specific rules that generalize across donors and perturbations.
So much of machine learning is about distributions, and in biology, distributions are often shifting. Techniques like optimal transport let us compare distributions, detect these shifts, and interpretably map between datasets, modalities, or domains. The value of combining modalities is also context dependent: complex biological tasks benefit when models capture synergistic information unavailable from either modality alone.
Across all our work, data choices shaped model behavior as much as architecture choices did.
Our Vision
These individual advances across modalities are converging. Research on cancer and other diseases isn’t accelerated by better classifiers or bigger datasets alone. It requires computational models that reason across scales, from sequences to structures, from cells to tissues, from one patient's tumor to patterns across populations. Our work provides infrastructure, benchmarks, and methods that make this possible.