Documentation

From raw reads to a cross-species reference spanning 100 species

This page documents the computational pipeline used to build the BrainStorm analysis layer: quality control, cross-species integration, the analysis matrix, program-level scoring, ortholog-version handling, and cell-type annotation. All parameters cited here mirror the companion manuscript Methods.

Quality Control and Library Filtering

Our in-house profiling produced 109 single-nucleus libraries spanning 52 vertebrate species, generated with two parallel techniques (SeekOne, 90 libraries, and 10x Chromium, 19 libraries). Nuclei were filtered with a uniform, permissive quality threshold, and libraries with low read-mapping rates were set aside before any downstream analysis.

ParameterSetting
Gene detection filter Cells retained when 200 to 6,000 genes were detected per cell.
Doublet removal Scrublet with its automatic threshold.
Mitochondrial filter None. No mitochondrial-percentage filter was applied in QC.
Library exclusion Five libraries with low read-mapping rates were excluded from downstream analyses.
Passing cells 912,605 cells across 104 libraries passed QC.

For species profiled with both techniques, the two platforms captured comparable cell numbers and gene counts, and cell-type composition and expression patterns were reproducible between them, indicating that both profiling methods are reliable and robust.

Batch Correction and Unsupervised Clustering

Batch effects were corrected with Harmony using library identity as the sole batch covariate, and the integrated cells were clustered with Leiden. This produced the atlas-level cluster annotation used throughout the resource.

ParameterSetting
Integration method Harmony, theta 5.6.
Embedding basis Top 50 principal components of 5,000 highly variable genes.
Batch covariate Library identity as sole batch covariate.
Clustering Unsupervised Leiden clustering, resolution 1.
Atlas clusters 43 clusters.

After integration, cells clustered by cell type rather than by species of origin: the median species iLISI was 7.76 versus a cell-type iLISI of 1.07, and the large clusters each spanned a median of 99 of the 100 analysis-layer species, indicating that the same cell type shares similar molecular blueprints across species.

A Uniform Cross-Species Matrix

The full integrated matrix comprises 165 libraries and 4,414,988 cells (including the 109 newly constructed libraries with 945,408 newly profiled cells). From this, a uniformly processed analysis layer was derived for all cross-species analyses, capped at 5,000 cells per species to balance representation across the tree.

PropertyValue
Analysis matrix 492,121 cells by 11,578 orthologous genes, 100 species.
Per-species cap 5,000 cells per species.
Fish 205,000 cells.
Mammalian 184,583 cells.
Bird 49,052 cells.
Reptile 38,486 cells.
Amphibian 15,000 cells.

The analysis layer is anchored to human homologs retained across species, which defines the gene universe for all cross-species analyses in this resource. For a full species-by-species breakdown, see the species catalogue.

Cell Program Scores and Detection

Across the atlas, cell programs (for example excitatory, inhibitory, astrocyte, or myelin programs) are scored at single-cell resolution, then summarized to species-level calls so that presence and absence can be compared across the phylogeny.

ParameterSetting
Scoring function scanpy.tl.score_genes.
Control size ctrl_size 50 (50 matched control genes per scored program).
Bins n_bins 25.
Data layer use_raw False.
Species-level summary Median of the per-cell scores within each species.
Detection rule
A program is reported as not tested when its gene coverage falls below 0.5. Grey entries in the phylogeny heatmaps indicate programs that could not be tested for a given species, which is a detection limitation and not evidence of absence.

Strict and Inclusive Ortholog Mapping

All cross-species comparisons are evaluated under two ortholog versions that differ in how ambiguously mapped homologs are handled. Reporting both makes the results robust to orthology-mapping stringency.

VersionDefinition
o2o (strict) Strict one-to-one ortholog assignments between human and each species.
o2o+m2o (inclusive) One-to-one plus human-side many-to-one orthologs, retaining assignments covered in at least 70 of the 100 species.

The choice shifts how many genes are carried and how much signal is retained per gene. The gene-age conclusions are stable across both versions: for example, the top neuronal-subtype markers remained 100% Euteleostomi-age under strict one-to-one (n = 261) and one-to-one-plus-many-to-one (n = 314) ortholog sets, and the four cross-species marker programs behaved consistently when assayable orthologs were compared.

Cell-Type Annotation Curation

The 43 atlas clusters were assigned to major cell types by expert manual curation of cluster-level marker expression, and the curated assignments were then executed programmatically as a cluster-to-type assignment table (the full curation table is provided in Supplementary Table 2).

Cluster groupCount (of 43)
Neuronal27
Neuroendocrine1
Astrocytes2
Microglia / CNS-associated macrophages3
Oligodendrocytes3
Oligodendrocyte progenitor cells2
Vascular1
Ambiguous3
Mixed-lineage1

Twelve of the 27 neuronal clusters rest on lower-confidence marker evidence and are retained as curated assignments pending orthogonal support. Unassigned clusters were annotated as co-expressing lineages or left undefined. In total, 483,837 of the 492,121 analysis-layer cells (98.3%) received a unified cell-type annotation. Annotation robustness was stress-tested by a full-pipeline rerun with an alternative highly-variable-gene set: cluster-level concordance was on par with or below the primary run, leaving all main conclusions unchanged.

Core identity rests on a set of 46 canonical markers selected from an extensive literature review for the major brain cell types, including neuronal (pan-neuronal, inhibitory and excitatory) markers, and all 46 core cell-type markers trace to the oldest phylostratum (Euteleostomi-age) within the cross-species-retained ortholog space.

Curated lineage

Neurons

Excitatory and inhibitory programs resolved across most species, with 24 lineage-resolved subclusters and neuronal identity genes as old as the brain-expressed background.

Read the neuron documentation →
Curated lineage

Microglia

Two superimposed layers of microglial identity, an ancient layer and an amniote-strengthened module, resolved across the 100-species analysis layer.

Read the microglia documentation →
Curated lineage

Astrocytes

A conserved yet progressively diversified set of astrocyte subclusters, from progenitor-like basal states to mammal-enriched differentiating states.

Read the astrocyte documentation →
Curated lineage

Oligodendrocytes

The oligodendrocyte lineage and its myelin programs, from a basic program in early jawed vertebrates to lineage-specific remodeling of state composition.

Read the oligodendrocyte documentation →

Explore and Download

Browse the integrated atlas, review available datasets, and connect with the team for questions about the pipeline documented here.