Research

The science behind single-cell target discovery

Our research program covers foundation model development for cell biology, single-cell data harmonization, and computational methods for cell-state-specific target scoring. We share methods and data openly where agreements permit.

Core methods

How we approach single-cell target identification

Three methodological pillars shape the Relation platform and distinguish it from traditional bulk-RNA or literature-mining approaches to target discovery.

Foundation model pre-training

We pre-train transformer-based encoders on large single-cell corpora. The objective is a cell-level representation that generalizes across tissues, donors, and disease conditions: the same model is used for healthy and disease queries without task-specific retraining. Training data is curated from Human Cell Atlas, CELLxGENE, and GEO repositories.

Multi-dataset harmonization

Single-cell data from different labs and platforms contains batch effects that, uncorrected, dominate the signal. We apply and evaluate multiple harmonization strategies to separate technical variation from true biological differences. Our internal benchmark assesses cell-type separation and donor-variation structure after correction.

Cell-state-aware target scoring

Traditional specificity metrics score gene expression across bulk tissue samples. Our scoring approach operates at the cell-type and cell-state level, identifying genes that are specific to particular disease-associated cell populations, not just to disease tissue overall. This matters because many genes appear disease-elevated only because of changes in cell-type composition.

Publications and preprints

Our research agenda

We are an early-stage team. Our first preprints are in preparation. We share our thinking on the blog and plan to post formal preprints to bioRxiv as work reaches publication-ready form.

Our current research program spans three areas:

  • Foundation model pre-training for cross-tissue single-cell representation, with a focus on generalizing across sequencing protocols and donor variation
  • Batch harmonization methodology for multi-source scRNA-seq training corpora, distinguishing technical batch from true biological variation
  • Cell-state-specific expression scoring for target candidate ranking, scoring genes against the specific disease-associated population rather than against tissue bulk

To discuss pre-publication collaboration or to request a briefing on our methods, reach us at [email protected].

Open data

Reference datasets and embedding releases

We make reference datasets and pre-computed cell embeddings available through public repositories when data use agreements permit. These releases are intended to support the broader single-cell and computational biology community.

  • Harmonized multi-tissue reference corpus (subset, CELLxGENE-licensed studies)
  • Cell-type annotation transfer benchmark dataset
  • Foundation model checkpoint for human tissue embedding (non-commercial research use)

To request access to released datasets or to discuss research collaboration, contact us at [email protected].