The BrainStorm Data Resources

A uniformly processed, cross-species single-nucleus resource

BrainStorm is distributed as a set of openly accessible data resources: a full integrated expression matrix, a uniformly processed cross-species analysis layer built for comparative work, per-species matrices, complete cell metadata and annotations, and a series of curated supplementary tables covering sampling, curation and evolutionary inference. All newly generated data and derived resources will be released together with the manuscript through public repositories whose accessions will be activated at publication.

165
Libraries
across the full integrated matrix
4,414,988
Cells, Full Matrix
single-nucleus transcriptomes
492,121 × 11,578
Analysis Layer
cells × orthologous genes, 100 species
52
Newly Profiled Species
109 libraries, 945,408 cells

Two Cohesive Layers, One Atlas

The resource is organized around two complementary matrices. The full integrated matrix combines all 165 libraries and 4,414,988 single-nucleus transcriptomes, including the 945,408 cells from 109 libraries generated for the 52 newly profiled species, together with 50 public datasets reanalyzed under a uniform processing pipeline. A uniformly processed analysis layer of 492,121 cells by 11,578 cross-species orthologous genes, capped at 5,000 cells per species, was derived from this full matrix and underlies every cross-species analysis in the project.

New & Public

52 Newly Profiled + 50 Public

Of the 100 analysis-layer species, 50 are newly profiled in the final atlas and 50 draw on public datasets, all processed and annotated under one uniform pipeline. Across the whole study, 52 species are newly profiled in total, of which 50 enter the final atlas.

Two Platforms

SeekOne + 10x Chromium

The 109 newly constructed libraries use two complementary single-nucleus platforms in parallel: SeekOne for 90 libraries and 10x Chromium for 19. Where both techniques cover the same species, cell composition and gene expression are highly reproducible between them.

Capacity

Balanced Sampling

Each species contributes at most 5,000 cells to the analysis layer, balancing the phylogeny: 205,000 fish, 184,583 mammalian, 49,052 bird, 38,486 reptile and 15,000 amphibian cells.

What Is Included in the Distribution

The release delivers the data needed to reproduce the atlas and to run new comparative analyses: expression matrices, per-species matrices, full cell metadata, and the unified cell-type annotations that drive the interpretation throughout the project.

Matrices

  • The full integrated matrix, 4,414,988 cells across 165 libraries.
  • The analysis-layer matrix, 492,121 cells by 11,578 orthologous genes.
  • A per-species matrix for every analysis-layer species, each capped at 5,000 cells.
  • A cross-species ortholog table linking gene models across the 100 species.
QC for the newly generated libraries retained cells with 200 to 6,000 detected genes, removed doublets with Scrublet, and applied no mitochondrial-percentage filter.

Metadata & Annotations

  • Full-cell metadata covering species, library, sample batch, and sequencing platform.
  • A unified cell-type annotation for 483,837 of the 492,121 analysis-layer cells (98.3%).
  • Cluster-to-type curation assignments for the 43 atlas-level clusters derived by expert manual annotation.
  • Lineage-resolved subcluster annotations for neurons, macroglia and microglia.
The remaining 8,284 analysis-layer cells carry an explicit "unannotated" label and are retained for transparency.

Two species profiled in-house, Misgurnus anguillicaudatus and Salarias fasciatus (two libraries each), were excluded before the analysis layer because of low genome-mapping rates (23.3 to 51.3%, versus at least 85.1% for all retained in-house libraries) and the resulting low gene complexity, attributable to incomplete reference annotation.

Supplementary Table Guide

The manuscript is accompanied by eight supplementary tables that document the resource end to end, from species sampling and library inclusion to cluster curation, phylostratigraphy and label mapping. All eight tables are provided as downloadable files alongside the release.

Table Contents
ST1 Species composition and inclusion decisions: genome sources, library batches, sequencing platform, full-matrix and analysis-layer cell counts for all 100 species, plus the two in-house species excluded before the analysis layer.
ST2 Cluster-to-type curation table: the expert manual assignment of each of the 43 atlas-level clusters to a major cell type, executed programmatically, with confidence flags where evidence is limited.
ST3 Phylostratigraphic framework: the 15-branch phylostratigraphy from Euteleostomi to Homo sapiens used for gene-age enrichment analyses.
ST4 / ST5 Analysis-layer expression data: the 492,121 cell by 11,578 gene matrix and its supporting structure used for all cross-species comparisons.
ST6 Public dataset index: the original accessions and provenance of the publicly available datasets reanalyzed in this study under the uniform pipeline.
ST7 Label mapping: the full mapping between biological names and cluster labels used for cross-species and reference-atlas comparisons.
ST8 Ligand-receptor interaction nomenclature: the stratum classification (core, tetrapod-added, amniote-added, lineage-specific) and per-pair presence across species for the interactome analyses.

How the Data Will Be Released

All data resources will be released together with the manuscript through public repositories. No accession identifiers or download links are active before publication; they will be added here and on the Publications page the moment the release goes live.

Raw & Processed Reads

Genome Sequence Archive (GSA)

Raw and processed single-nucleus RNA-seq data for the 52 newly profiled species will be deposited in the Genome Sequence Archive of the National Genomics Data Center (NGDC), Beijing Institute of Genomics, Chinese Academy of Sciences, comprising 109 snRNA-seq libraries.

Status: accession to be activated at publication.
Derived Resources

Zenodo

The integrated atlas object, full-cell metadata, per-species matrices and the cross-species ortholog table will be deposited in Zenodo with a commitment to at least 5 years of availability.

Status: DOI to be activated at publication.
Analysis Code

GitHub

All custom analysis code for data processing, cross-species integration, cell-type annotation and evolutionary analyses is maintained in the project repository, private during peer review and made public upon publication, with reviewer access available on request.

Repository: github.com/DongshengChen-TY/brainstorm-atlas

Formats for Analysis and Interoperation

The release is provided in standard single-cell formats so the data drop directly into existing analysis workflows, from interactive exploration to command-line pipelines.

.h5ad

H5AD (AnnData)

The integrated atlas object and the analysis-layer object are provided in the H5AD format, bundling expression matrix, cell metadata and cluster annotations into a single self-describing file for direct use with the scverse ecosystem.

.csv

CSV

Full-cell metadata, per-species matrices, the cross-species ortholog table and all supplementary tables are provided as comma-separated value files for easy loading into any spreadsheet or statistics environment.

For an introduction to loading and working with the matrices, see the Documentation page. The interactive Cell Browser lets you explore cell-type identities without downloading anything.

Explore Before You Download

Browse the integrated cross-species atlas interactively, compare cell types and lineages, and find the species behind each annotation.