Why Cell-Type Resolution Matters in Drug Target Discovery
Bulk RNA sequencing changed biology. It gave researchers a quantitative readout of gene expression across thousands of samples, enabling the large-scale discovery of disease-associated genes that had been impossible with earlier methods. We owe a lot to it. But there is a structural limitation in bulk data that matters specifically for drug target selection, and understanding it is the starting point for why single-cell approaches are changing what discovery teams prioritize.
That limitation is averaging. When you sequence bulk tissue, you are measuring the aggregate RNA output of every cell in that sample. The resulting gene expression profile is a weighted mean: each gene's value reflects how much every cell type in the tissue is expressing it, mixed proportionally to how many of each cell type are present. For many analytic purposes, that average is exactly what you want. For target discovery in heterogeneous tissues, it systematically hides the thing you need to see.
The Signal You Cannot See in an Average
Consider a synovial tissue sample from a patient with inflammatory arthritis. That sample contains fibroblasts, macrophages of several activation states, T cell subsets, B cells, endothelial cells, and mast cells. If a gene is strongly upregulated in one specific activated fibroblast subpopulation that makes up roughly 8% of the total cell count, its contribution to the bulk signal is diluted by the other 92%. A 30-fold upregulation in that subpopulation might register as a 2.4-fold change in the bulk readout. Statistically, that might be significant. Biologically, you have now lost the information that the gene is concentrated in one specific cell class, which is likely the one relevant to the disease mechanism you are trying to hit.
This is not a theoretical edge case. It happens routinely in tissues where the pathological cell population is a minority. Fibrotic diseases, tumour microenvironments, most neurological conditions, and many autoimmune contexts all involve minority cell populations driving much of the observable pathology. Bulk data from these tissues tends to reflect the dominant cell type, not the disease-driving one.
How Resolution Shifts the Hypothesis
When a discovery team works from bulk data, the target hypothesis has a particular shape: "gene X is upregulated 3-fold in disease tissue relative to matched controls." That hypothesis is valid as far as it goes. It tells you something changed. It does not tell you where, in cellular terms, the change is happening.
The single-cell version of the same hypothesis looks different: "gene X is upregulated 22-fold in SPARC-positive activated fibroblasts in disease tissue, with no significant change in macrophages, endothelial cells, or resting fibroblasts." Now you know the cell-type address of the change. That information shapes the drug discovery program in concrete ways: which cell type your drug needs to access, what expression and selectivity profile you should be testing, and whether on-target activity in other cell types is likely to be a safety concern.
The difference is not purely about statistical sensitivity. It is about the kind of question you can ask. Bulk data lets you generate hypotheses at the gene level. Single-cell data lets you generate hypotheses at the cell-type-and-gene level. For many disease areas, the second type of hypothesis is the one that actually predicts whether a program will reach efficacy in human tissue.
A Discovery Gap: A Plausible Example
Consider a growing biotech working on an inflammatory bowel disease indication in 2024. Their bulk RNA-seq screen of colonic mucosal biopsies identifies a cytokine receptor with consistent 2.8-fold upregulation in active disease versus remission. The receptor has a known ligand, known pharmacology in cell lines, and looks tractable. It goes on the priority list.
When the same samples are analysed with single-cell resolution, the picture complicates. The bulk upregulation turns out to come from two sources: an expansion of mast cells in inflamed tissue (mast cells express this receptor at baseline, so their increased abundance inflates the bulk average), and genuine upregulation in a specific subset of ILC3-adjacent innate lymphoid cells. The mast cell contribution is largely incidental to the core pathological mechanism. The ILC3 subset contribution is the more interesting signal, but it is narrower and would require a different pharmacological strategy to address effectively.
This kind of deconvolution is not possible from bulk data. It requires single-cell profiles from the same or matched samples, and a computational framework that can assign those profiles to cell types reliably.
Target Specificity Across Cell Types
Beyond identifying which cell type harbours the disease signal, single-cell data also allows you to evaluate how specifically a target gene is expressed across the broader cellular landscape of the relevant tissue. This is important for safety prediction.
A target that is upregulated 15-fold in the disease-driving cell population but also has significant expression in hepatocytes, for example, will have a different risk profile from one that is effectively absent outside the cell type of interest. Bulk data cannot give you that comparison cleanly. Single-cell data from a well-profiled tissue reference, cross-compared to the disease state, gives you an expression specificity map that is actionable before animal studies begin.
This does not eliminate the need for in vivo work. But it allows you to rank target candidates by a criterion that has predictive value for clinical safety: how narrow is the cell-type expression window in human tissue?
Not Every Question Needs Single-Cell Resolution
We are not saying bulk RNA-seq is the wrong tool for every purpose. It is not. For eQTL analysis, population genetics, large-scale cohort studies examining gene expression across hundreds of patient samples, and network analysis at the pathway level, bulk data at scale is still the appropriate choice. It is cheaper, it scales further, and the statistical power from large N often outweighs the resolution loss.
The cases where cell-type resolution changes the answer are specific: tissue heterogeneity is high, the pathological population is a minority, and your target hypothesis depends on knowing which cells are changed rather than just whether the tissue as a whole shows a change. Drug target discovery in inflammatory, fibrotic, neurological, and oncological diseases frequently falls into this category. Cardiovascular and metabolic diseases involving tissues with less cell-type heterogeneity may not need single-cell resolution for every target, though the option is increasingly useful as organ-specific cell-type biology becomes better characterised.
What Resolution Requires Computationally
Single-cell data is not drop-in replacement for bulk. The analytical challenges are different in kind, not just scale. Cell-type annotation is a nontrivial step: assigning cell identities to clusters requires reference data, domain expertise, and in many disease contexts, careful handling of novel or transitional cell states that do not map cleanly to canonical types.
Batch effects across studies are significant. Multiple droplet-based protocols, tissue dissociation methods, and sequencing depths introduce technical variation that can masquerade as biological variation. Harmonisation methods (variational autoencoders, mutual nearest neighbour correction, reference-anchored methods) have matured considerably, but choosing the right approach for a given analysis question still requires care.
The computational demands of analysing large single-cell datasets at the scale needed for meaningful target prioritisation are also substantial. Working with hundreds of thousands of cells across multiple disease states and matched controls, maintaining annotation consistency across studies, and extracting gene-level expression specificity scores per cell type: these are tractable problems, but they require pipelines that are purpose-built for the task rather than repurposed from bulk workflows.
These are the challenges we built Relation Therapeutics to address. The goal is not to replace the biology judgment that discovery teams bring: it is to ensure that judgment is applied to a target hypothesis that accurately represents the cell-type context of the disease mechanism, not an average that obscures it.