The Biology of Disease-Driving Cell Populations
Disease does not happen evenly across a tissue. When we look at diseased organs, whether using histology, molecular profiling, or imaging, we consistently find that specific subsets of cells are doing most of the pathological work: secreting the cytokines that sustain inflammation, remodelling the matrix that leads to fibrosis, acquiring the epigenetic changes that drive aberrant proliferation. Adjacent cell populations of superficially similar type can be largely uninvolved. This cellular heterogeneity within a diseased tissue is not incidental detail: it is the mechanism, and understanding it is a prerequisite for selecting targets that can interrupt it.
The concept of the disease-driving cell population is not new to cell biology. What has changed is our ability to characterise these populations at the molecular level in human tissue at scale, which is what makes cell-type-resolved target discovery possible where it was not before.
The Concept of Pathological Cell State
A cell population can drive disease through two distinct mechanisms. The first is abundance: a cell type that is rare in healthy tissue becomes disproportionately expanded in disease, and its normal function at this elevated frequency causes tissue damage. The second is state: a cell type that is present at normal abundance undergoes transcriptional reprogramming that changes what it does, acquiring a pro-pathological gene expression programme without necessarily expanding numerically.
In practice, both mechanisms operate together in many diseases, but the proportion varies. Some inflammatory conditions are driven primarily by the expansion of pathological immune cell subsets. Fibrotic diseases often involve the emergence of a transcriptionally distinct fibroblast state that is present in normal tissue in very small numbers but becomes dominant in the disease context. Neurodegeneration involves both the loss of specific neuronal populations and the state transition of glial cells toward reactive phenotypes that contribute to the pathological environment.
The distinction matters for target discovery because the two mechanisms require somewhat different approaches. If the pathological cell type is expanded but otherwise normal, you are looking for targets related to its survival, proliferation, or migration signals. If the cell type has undergone a state transition, you are looking for the regulators that maintain the pathological state and the downstream effectors through which it causes damage.
Tissue-Resident versus Recruited Populations
Another axis of variation that matters for target selection is whether the disease-driving population is tissue-resident or recruited from circulation. The distinction affects both the biology of the target and the practical question of drug accessibility.
Tissue-resident macrophages, for example, are a fundamentally different population from monocyte-derived macrophages that infiltrate inflamed tissue. Single-cell profiling of inflammatory lesions consistently shows that these populations have distinct transcriptional programmes and occupy distinct niches within the tissue architecture. A target that is relevant to the pathological function of tissue-resident macrophages may have no activity on the infiltrating population, and vice versa.
Similarly, in fibrotic diseases, the disease-driving mesenchymal population is thought to arise from multiple cell sources in different proportions depending on organ and disease context: tissue-resident fibroblast subsets, pericyte descendants, epithelial cells undergoing mesenchymal transition. Single-cell lineage tracing and transcriptional profiling is clarifying these contributions in ways that are directly relevant to which upstream signals represent valid targets for interrupting fibroblast activation.
How Disease State Differs from Cell Type
One of the more important conceptual shifts in single-cell biology over the last few years is the recognition that "cell type" and "cell state" are distinct levels of description that require distinct analytical approaches.
Cell type refers to the stable identity of a cell: it is a fibroblast, a macrophage, a CD8+ T cell. This identity is maintained across contexts and reflects lineage history and epigenetic programming. Cell state refers to the current transcriptional activity of that cell, which can vary substantially depending on environmental signals, local cytokine milieu, and activation status.
Disease-associated cell states are cell states that emerge specifically in pathological contexts. A macrophage in an atherosclerotic lesion occupies a state different from the same macrophage in healthy adipose tissue: it has adopted a transcriptional programme associated with lipid accumulation and inflammatory signalling that defines its contribution to plaque progression. That programme is partially cell-type-specific (it reflects macrophage biology) and partially disease-context-specific (it would not appear in a macrophage not exposed to the atherosclerotic microenvironment).
For target discovery, identifying disease-associated cell states is particularly valuable because the genes that define those states are directly tied to the pathological mechanism. A gene that is selectively expressed in the disease-associated state of the cell type of interest, and that has been shown in functional studies to regulate the properties of that state, is a much higher-quality target starting point than a gene identified by tissue-level association alone.
Specificity and Its Implications for Safety
When a target gene is concentrated in the disease-driving cell population, the implication is not only better efficacy: it also has a bearing on the likely safety profile of targeting it. A gene with narrow expression to the disease-relevant cell type, with minimal baseline expression in other tissues at the single-cell level, carries a different on-target toxicology risk profile than a gene that is broadly expressed across many cell types.
This is not a guarantee of safety: a gene with selective expression in disease cells can still have important functions in normal biology at the cell-type level, and the detailed characterisation of those functions requires in vitro and in vivo work. But early assessment of expression specificity across human tissues and cell types, using single-cell reference data, gives a first-pass filter that can help prioritise between candidates before significant lab investment.
The worst-case scenario from a target selection standpoint is a gene that is expressed at moderate levels across many cell types, with only modest differential expression in the disease context. Such targets are easy to find using bulk differential expression analysis; they are also the ones most likely to produce on-target side effects and to fail at the selectivity challenge in the clinic. Single-cell expression specificity analysis does not eliminate these candidates, but it makes the risk profile visible earlier.
Quantifying Pathological Contribution
An important question that single-cell analysis can begin to address is: how much of the disease phenotype is attributable to the putative disease-driving population, versus other cell types in the tissue?
This question does not have a simple quantitative answer from transcriptomics alone. But compositional analysis (how the proportion of each cell type changes between disease and healthy tissue), combined with the magnitude of state shift within each population, provides a rough ranking of which populations are most changed. When this ranking correlates with known pathological biology in the relevant disease area, it provides a useful check on the hypothesis. When it produces unexpected findings, those findings are worth investigating, because they may reflect cell populations whose contribution to disease has been underweighted in the field.
We are not claiming that this analysis alone defines the disease mechanism. What it does is provide a cell-type-addressed starting point that anchors target hypotheses in human biology, rather than in model systems that have their own biases and limitations. The gap between a well-characterised mouse model and the actual human disease is one of the better-understood sources of late-stage clinical failure. Working from human single-cell data does not close that gap entirely, but it narrows it considerably, and it does so at the stage where the cost of a course correction is lowest.