Why Most Disease Models Miss the Cell-Type Signal
There is a recurring gap between what preclinical disease models show and what happens in the clinic. Late-stage clinical failure rates in complex diseases have remained stubbornly high for decades, and a meaningful fraction of those failures trace back to something that happened at the target selection stage, not at the chemistry or formulation stage. The biology was wrong before the molecule was ever designed.
The question worth asking is structural: what do the most common disease modeling approaches systematically obscure? Not what individual experiments get wrong in specific cases, but what properties of the standard toolkit hide information that would change target selection decisions if it were visible. Cell-type resolution is the answer we keep arriving at. The dominant experimental and computational methods each suppress it in different ways.
Mouse Models: Species and Cell-Type Divergence Compound
Mouse disease models are the workhorse of preclinical target validation. They allow controlled genetic perturbation and clean endpoint measurement that is not possible in human samples. The problems are well known in the field: murine physiology differs from human physiology in important ways, and many diseases do not model faithfully in standard inbred strains.
Less often discussed is the cell-type dimension of the translation problem. Human tissues and mouse tissues contain broadly analogous cell types, but the detailed composition and the specific subtype distributions differ. In the human lung, for example, the myeloid compartment includes cell populations with expression programs that do not have direct one-to-one counterparts in standard mouse models. A target that drives pathology through a specific myeloid subtype in human disease may be expressed quite differently, or in a functionally distinct cell population, in the mouse model used to validate it.
When a mouse model experiment shows target relevance, the signal is real. It just may be indexing a different cell population in the mouse than the one you intended to target in humans. That ambiguity does not surface until you go back to human tissue and ask whether the validated target is expressed in the cell type you hypothesized. In a disease where the relevant human cell population was identified by bulk gene expression and not at cell-type resolution, you often do not know the answer to that question.
Cell Lines: Dedifferentiation Erases Subtype Identity
In vitro cell lines offer experimental tractability and unlimited scalability. The cellular biology problem is well understood: continuous passaging selects for proliferative cells, which progressively dedifferentiate from the original cell type. A cancer cell line that began as a specific epithelial subtype may have lost most of the transcriptional programs that defined that subtype by the time it is widely used in the field.
For target discovery, this creates a specific problem. If a target is expressed in a disease-relevant cell subtype precisely because of that subtype's specialized transcriptional program, a dedifferentiated cell line will not faithfully recapitulate that expression. Experiments run in the cell line will not capture the biology you are trying to model. More precisely, they will capture some biology, but not the biology relevant to the cell-type-specific disease mechanism.
This is not an argument against all in vitro work. It is an argument for being deliberate about which cell systems are appropriate for which questions. A target whose relevance is defined at the cell-type level needs to be interrogated in a system that preserves the relevant cell-type identity, whether that is primary cells, organoids, or carefully characterized cell lines with confirmed retention of the relevant transcriptional programs.
Bulk RNA-seq: The Averaging Problem
Bulk RNA-seq is still the most widely used transcriptomics approach in disease biology, including in large patient cohort studies. A bulk sample contains a heterogeneous mixture of cell types. The expression values you measure are a weighted average of expression across all cells, weighted by cell-type abundance. That average obscures two distinct phenomena: differential expression that is specific to one cell type, and differential cell-type composition between conditions.
Consider a disease where a specific macrophage subtype expands in affected tissue relative to healthy tissue. A bulk comparison between disease and healthy tissue will show differential expression of genes that are markers of that macrophage subtype. But it will also show differential expression of genes that are expressed in many cell types simply because the overall cellular composition has shifted. The two signals are mixed in the bulk output, and separating them requires either deconvolution methods or single-cell data.
Deconvolution helps, but it cannot recover what was never measured. If a disease involves a cell subtype that was not included in the reference panel used for deconvolution, the signal from that subtype will be misattributed to the most similar type in the reference. Rare or novel subtypes, which are often the most biologically interesting from a target discovery perspective, are precisely the ones most likely to be missed.
GWAS and eQTL Analyses: Cell-Type Context Is Lost in Mapping
Genome-wide association studies have identified thousands of loci associated with disease risk. Converting those loci to target hypotheses typically involves looking at what genes are near the associated variants and what tissues those genes are expressed in. Expression quantitative trait loci analyses take this further by asking which variants affect gene expression, providing a molecular mechanism for the genetic association.
The cell-type problem here is that eQTL analyses have historically been done on bulk tissue or on a small number of cell types profiled in large cohorts. A variant that has a strong eQTL effect in a specific cell subtype may appear weak or absent in bulk tissue data if that subtype is rare in the tissue sample. The target gene whose expression is regulated by the disease-associated variant may only be identifiable as such from single-cell eQTL data, where you can ask which cell type shows the strongest regulatory effect.
Single-cell eQTL resources are growing. Studies using pseudobulk approaches on large single-cell datasets have begun to map cell-type-specific genetic regulatory effects for immune cell types and others. This is not yet comprehensive across tissues and disease-relevant cell types, but the direction is toward recovering the cell-type specificity that bulk eQTL analyses lose.
The Common Thread
Mouse models, cell lines, bulk RNA-seq, and bulk eQTL analyses all have genuine value in disease biology. We are not saying any of them are invalid approaches. We are saying they each suppress cell-type resolution in specific ways, and that suppression is not neutral for target discovery.
Targets selected from cell-type-obscured data carry an implicit assumption: that the biology motivating the hypothesis is not cell-type-specific, or that the relevant cell type is approximately represented in the model system used. In diseases where the pathological mechanism is driven by specific cell subtypes, that assumption is frequently wrong, and the wrongness does not become apparent until clinical or late preclinical stages.
The practical question for any discovery team is: at which stage does cell-type resolution need to enter the analysis? Starting with cell-type-resolved human data at the target identification stage changes what hypotheses are generated and which ones look compelling. It does not eliminate the need for in vitro and in vivo validation work. It changes the starting point of that work from a target identified from averaged signals to one identified from the cell population most relevant to the disease mechanism in human tissue.
What Building From Cell-Type Resolution Looks Like
A target hypothesis built from single-cell data starts with the cell population. Which cell type or subtype shows the most disease-specific transcriptional changes? What are the top differentially expressed genes in that population relative to healthy controls of the same cell type? Which of those candidates show expression specificity for the disease-relevant population relative to other cell types in the tissue?
That specificity question is what cell-type resolution enables and bulk approaches cannot answer. A target gene that is differentially expressed in one cell type across multiple independent disease datasets, and that shows restricted expression in that cell type relative to neighboring populations, carries a different evidence profile than one that appears in a bulk differential expression analysis. The cell-type specificity is the signal that allows you to predict which modulation will be selective for the disease mechanism and which will have broad effects in tissue.
This approach produces a smaller set of candidates per indication than bulk analysis, not a larger one. The filtering is more aggressive. That is the point. Starting from cell-type-resolved human biology means fewer candidates, but ones whose evidence base is more precisely tied to the disease mechanism you are trying to address.