Weekly reads 06/07/26
From cell competition to genetic edges: how intercellular interactions redefine disease and discovery
After a week off, this week’s reads focus on how the hidden dynamics of cell-cell interactions are reshaping our understanding of disease—from the competitive origins of glioma to the genetic architecture of intercellular communication. Meanwhile, Yang et al. introduce EdgeMap, a system that breaks down trait heritability into parts that are specific to individual cells and parts where cell-cell communication plays a role. The finding is a new insight that the genetic determinants of many complex traits (diseases) can be seen in both individual cells and across the interfaces of the cells connected by ligand-receptor pairs. Tyagi et al. study on Alzheimer’s disease shows how a neuronal protein called Arc is wrapping pathological tau in vesicles that can be released from cells and that this mechanism leads to spread of the disease between different brain cells. On the computational side, Zhao et al. developed the CONCISE method which is able to accurately infer cell-cell communications by taking advantage of spatial transcriptomics, while Palla et al. demonstrated that the use of general-purpose foundation models on tabular data can produce better predictions than the use of specialized biological architectures when it comes to a variety of cell responses to perturbations.
Preprints/articles that I managed to read this week
Critical role of cell competition in gliomagenesis
Jiang et al. bioRxiv (2026). 10.64898/2026.01.15.699806
The paper in one sentence
This study identifies cell competition among oligodendrocyte precursor cells (OPCs) as a critical driver of gliomagenesis, demonstrating that pre-malignant Trp53,Nf1-null OPCs outcompete wildtype counterparts through mTORC1-dependent mechanisms, and that blocking this competition prevents both pre-malignant expansion and malignant progression.
Summary
Using a mouse genetic mosaic system (MADM), the authors show that sporadic Trp53,Nf1-null OPCs outcompete wildtype OPCs during pre-malignant expansion, leading to glioma formation. Through “in-tissue” phosphoproteomic profiling, they identify mTORC1 signaling as a key, competition-specific node—required for OPC competition but not for general OPC biology. Blocking competition by strengthening wildtype OPCs (via Nf1 loss) or inhibiting mTORC1 in mutant OPCs prevents gliomagenesis. Patient biopsy analysis reveals OPC depletion in tumor regions, supporting the relevance of OPC competition in human gliomas. The study also shows that Nf1 loss and EGFRvIII drive competition, while Trp53 loss, G1 checkpoint disruption, or Pten loss do not, highlighting pathway-specific mechanisms.
Personal highlights
OPC competition as a driver of gliomagenesis: Pre-malignant Trp53,Nf1-null OPCs outcompete wildtype OPCs, leading to their expansion and eventual malignant transformation, as demonstrated by a 300-fold increase in the G/R ratio (GFP⁺ mutant/RFP⁺ wildtype OPCs) with only marginal changes in total OPC density.
mTORC1 as a competition-specific signaling node: Phosphoproteomic profiling and genetic validation reveal that mTORC1 activity is required for OPC competition but is dispensable for normal OPC proliferation and survival, as shown by the failure of Rptor-null OPCs to compete.
Competition blockade prevents tumor progression: Strengthening wildtype OPCs (via Nf1 knockout) or inhibiting mTORC1 in mutant OPCs prevents both pre-malignant expansion and malignant glioma formation in mouse models.
Human relevance: Single-nucleus RNA-seq and immunohistochemistry of patient biopsies show a specific paucity of OPCs in tumor cores and infiltrating regions, consistent with competition-driven elimination of normal OPCs.
Mutation-specific effects: Nf1 loss and EGFRvIII drive OPC competition, while Trp53 loss, Rb inactivation, or Pten loss do not, underscoring the context-dependent nature of cell competition in gliomagenesis.
Why should we care?
This work shifts the paradigm of gliomagenesis from a purely cell-intrinsic process to one driven by competitive interactions between mutant and wildtype cells. The identification of mTORC1 as a competition-specific vulnerability suggests a new therapeutic strategy: rather than directly killing tumor cells, therapies could target the competitive advantage of pre-malignant cells or strengthen their normal counterparts to prevent tumor initiation and progression. However, the study has limitations: it focuses primarily on OPC-like glioma states, and the reliance on mouse models may not fully capture the complexity of human gliomas. Additionally, while the patient data are compelling, the absence of OPCs in tumor regions could also be due to technical challenges in detection. Nevertheless, the findings provide a conceptual framework for understanding how early competitive interactions shape tumor evolution, with potential implications for unconventional therapies that disrupt cell competition rather than targeting tumor cells directly.
Intercellular communication is a heritable dimension of human tissue architecture
Yang et al. bioRxiv (2026). 10.64898/2026.03.29.715138
The paper in one sentence
The study introduces EdgeMap, a computational framework that integrates spatial transcriptomics with GWAS to reveal that genetic risk for complex traits is organized not only within cells but also across the ligand-receptor interfaces that connect neighboring cells.
Summary
This work presents EdgeMap, a method that decomposes trait heritability into two orthogonal components: cell-intrinsic (node) and cell-cell communication (edge). By combining spatial transcriptomics with GWAS summary statistics, the authors construct a spatial neighbor graph and quantify per-cell communication intensity for ligand-receptor (LR) pairs. Gene-level scores for intrinsic expression and spatially concentrated signaling are then mapped to SNP-level LD scores and jointly regressed, enabling the separation of heritability into node and edge contributions. Applied across 17 complex traits (cardiovascular, psychiatric, metabolic, immune) and five human tissues (heart, brain, hippocampus, liver, gut), the study finds that edge heritability is significantly enriched in biologically coherent trait-tissue pairings (3.8-fold; P = 4.4 × 10⁻⁶) and replicates across independent tissue sections, GWAS cohorts, and platforms (e.g., Visium HD). Per-pair decomposition identifies 67 trait-specific communication channels (FDR < 0.10), organized into pathway families such as neurexin-neuroligin synaptic signaling in bipolar disorder, vascular adhesion in cardiovascular traits, and lipoprotein-clearance pathways in liver. Critically, 64% of edge genes are absent from standard gene-level prioritization methods (e.g., MAGMA, Open Targets L2G), suggesting that intercellular communication represents a complementary dimension of genetic architecture.
Personal highlights
Novel heritability decomposition: EdgeMap introduces a framework to separate cell-intrinsic from intercellular genetic signals, revealing that genetic risk can concentrate at ligand-receptor interfaces.
Biological coherence: Edge heritability shows 3.8-fold enrichment in a priori expected trait-tissue pairings (P = 4.4 × 10⁻⁶), supporting the biological relevance of intercellular communication in genetic architecture.
Pathway-level resolution: Identifies 67 trait-specific LR channels (FDR < 0.10), including known disease-relevant pathways (e.g., neurexin-neuroligin in bipolar disorder, PCSK9-SORT1 in liver LDL).
Complementary gene prioritization: 64% of edge genes are missed by standard methods, highlighting intercellular interfaces as an underexplored source of genetic targets.
Robust validation: Results replicate across independent tissue sections, GWAS cohorts, and spatial platforms, with controls ruling out major confounders.
Why should we care?
This study challenges the cell-centric paradigm in genetic interpretation by demonstrating that genetic risk is not only organized within cells but also across the molecular interfaces that connect them. The identification of trait-specific communication channels, many of which are missed by traditional gene-prioritization methods, suggests that intercellular signaling hubs could represent a rich, underexplored space for therapeutic targets. This is particularly relevant for diseases where cell-cell interactions are central to pathology (e.g., psychiatric disorders, cardiovascular disease). The alignment with known biology (e.g., PCSK9-SORT1 in lipid metabolism) and the overlap with approved drug targets (53% of edge genes) further underscores the potential translational value. However, individual trait-tissue signals are moderate (z ≈ 2–3), and the true causal nature of these edges remains to be established. In addition, the approach depends on the quality of ligand-receptor databases and spatial resolution; while the authors address potential confounders, further validation is needed to confirm that these edges represent true biological mechanisms rather than statistical artifacts. Lastly, the study focuses on common variants and may not capture rare or context-specific interactions.
Arc mediates intercellular tau transmission via extracellular vesicles
Tyagi et al., Cell (2026). https://doi.org/10.1016/j.cell.2026.06.008
The paper in one sentence
The neuronal protein Arc packages pathological tau into extracellular vesicles (EVs), enabling its intercellular transmission and driving the spread of tau pathology in Alzheimer’s disease and tauopathies.
Summary
This study combines mouse models, human brain tissue analysis, and molecular modeling to demonstrate that Arc, a protein known for its role in synaptic plasticity, directly binds tau, particularly its phosphorylated forms, and packages it into EVs. These Arc-containing EVs are released from neurons and can seed tau aggregation in recipient cells. In transgenic mice expressing human mutant tau (rTg4510), knocking out Arc reduces EV-associated tau, nearly abolishes intercellular tau transmission, and leads to intracellular tau accumulation and early neuronal toxicity. Importantly, EVs isolated from postmortem human Alzheimer’s disease (AD) brains contain both Arc and phosphorylated tau, with Arc levels strongly correlating with phosphorylated tau in these vesicles. The authors propose a model where Arc-dependent EV release initially helps neurons eliminate toxic tau but ultimately facilitates its spread, contributing to disease progression.
Personal highlights
Arc as a tau packaging factor: Arc directly binds tau, with a higher affinity for phosphorylated tau, and mediates its encapsulation into small EVs (~130 nm) released from neurons.
Arc is essential for tau transmission: Genetic knockout of Arc in mice and primary neurons dramatically reduces intercellular tau spread, demonstrating its central role in this process.
Human relevance: Brain-derived EVs from AD patients co-package Arc and phosphorylated tau, with Arc levels correlating with phosphorylated tau in human EVs, supporting the clinical significance of this mechanism.
Molecular mechanism: Structural modeling reveals dynamic, multivalent interactions between Arc and tau, particularly through tau’s aggregation-prone motifs (e.g., PHF6), suggesting how tau is recruited into Arc-containing capsids.
Dual role of Arc-EVs: While Arc-dependent EV release may initially protect neurons by removing intracellular tau, it also promotes the spread of seed-competent tau to neighboring cells, potentially accelerating pathology.
Why should we care?
This work provides a mechanistic explanation for how tau pathology spreads between neurons in Alzheimer’s disease. By identifying Arc as a key regulator of tau packaging into EVs, the authors offer a molecular target for therapies aimed at halting disease progression. However, there are important caveats. The mouse models used express supraphysiological levels of mutant human tau, which may not fully recapitulate the subtler, age-dependent tau pathology seen in human AD. Additionally, while Arc-dependent EV transmission appears critical, other mechanisms (e.g., free tau release, tunneling nanotubes, or microglial uptake) likely contribute to tau spread. Finally, the therapeutic potential of targeting Arc remains speculative: disrupting Arc function could have unintended consequences, given its well-established role in synaptic plasticity and memory.
Spatial co-expression and cell-cell communication inference from spatially resolved transcriptomics with CONCISE
Zhao et al., bioRxiv (2026). 10.64898/2026.06.22.733860
The paper in one sentence
The authors introduce CONCISE, a statistical method for inferring cell-cell communication from spatially resolved transcriptomics that explicitly models spatial autocorrelation, count-based data properties, and measurement errors to reduce false positives and improve reliability.
Summary
This preprint presents CONCISE, a unified statistical framework designed to address major confounding factors in spatial transcriptomics (ST) data—spatial autocorrelation, variation in total molecular counts, and measurement errors—which can lead to spurious co-expression signals and inflated false-positive rates in cell-cell communication (CCC) inference. The method uses a measurement-expression model to separate true expression from observed counts, incorporates spatial kernels to account for autocorrelation, and employs moment-based estimation with analytical null distributions for efficient hypothesis testing. Through simulations, real-data permutation experiments, and negative-control analyses across multiple ST platforms (10x Visium, Stereo-seq, MERFISH, CosMx), the authors demonstrate that CONCISE achieves well-calibrated p-values, robust false-positive control, and improved detection power compared to existing tools (e.g., MERINGUE, SpatialDM, Copulacci, LIANA+). Applications to breast cancer, mouse embryo, intestinal inflammation, and non-small cell lung cancer (NSCLC) datasets reveal biologically meaningful interactions, including inflammation-associated fibroblast signaling in IBD and tumor-immune/stromal communication networks in cancer.
Personal highlights
Unified modeling of confounds: Jointly accounts for spatial autocorrelation, count-based data properties, and measurement errors—key sources of spurious signals in ST data that most existing methods overlook.
Better false-positive control: Demonstrates well-calibrated inference and type I error control in simulations and real datasets, even when >80% of genes exhibit spatial autocorrelation (e.g., in human breast cancer Visium data).
Computationally efficient: Uses analytical null distributions and moment-based estimation to analyze 1,000 ligand-receptor pairs in under 6 minutes, outperforming permutation-based methods (e.g., Copulacci, which requires >18 hours).
Platform-agnostic: Validated across diverse ST technologies (Visium, Stereo-seq, MERFISH, CosMx), highlighting broad applicability.
Raw-count embeddings improve single-cell foundation models
ref
The paper in one sentence
This study demonstrates that single-cell transformer foundation models perform best when using non-normalised, log-transformed raw counts and random gene ordering, challenging the need for elaborate preprocessing steps like library-size normalisation and rank-based tokenisation.
Summary
The authors systematically benchmarked seven preprocessing strategies for single-cell RNA-seq foundation models, including rank-based ordering, fold-change, and value-projection methods. They found that raw, log1p-transformed counts with random gene ordering, outperformed all other approaches across five diverse benchmarks (cell type classification, doublet detection, and gene-level tasks). Their method, Gene Intelligence, also matched or surpassed the performance of much larger models (e.g., Geneformer, Transcriptformer) on cell-level tasks while using 10- to 200-fold fewer parameters. Notably, the model achieved state-of-the-art results in doublet detection, outperforming eight established methods on benchmark datasets. The study suggests that simpler, count-based embeddings retain biologically meaningful signals (e.g., total RNA content) that are lost during normalisation, enabling the model to learn task-specific corrections.
Personal highlights
Preprocessing simplicity wins: Non-normalised, log1p-transformed raw counts with random gene ordering outperformed sophisticated rank-based or fold-change strategies across all benchmarks.
Gene order doesn’t matter: Random gene ordering performed as well as, or better than, principled sorting schemes, including those based on global rank or fold-change.
Parameter efficiency: The smallest Gene Intelligence model (4.7M parameters) outperformed models 200× larger on gene-level tasks, while matching large models (e.g., Transcriptformer) on cell-level tasks.
State-of-the-art doublet detection: Fine-tuned Gene Intelligence achieved the highest average AUROC across 11 benchmark datasets, surpassing dedicated tools like Scrublet and DoubletFinder.
Why should we care?
This work challenges a persistent dogma in single-cell transcriptomics: that normalisation and elaborate preprocessing are essential to account for technical noise. By showing that raw counts, long considered too “noisy” for direct use, can yield better, more efficient models, the study suggests that current pipelines may be overcorrecting data, thereby removing biologically relevant signals. The main takeaway is that simplicity can outperform complexity: the field’s reliance on normalisation and ranking may have obscured the inherent informativeness of raw expression data. Critically, the results also highlight a trade-off: while Gene Intelligence excels in efficiency and certain tasks, its reliance on raw counts assumes the model can learn to handle technical artifacts, an assumption that may not hold universally across all datasets or experimental conditions.
Tabular Foundation Models Are Competitive Cellular Perturbation Predictors Across Biological Scales
Palla et al. bioRxiv (2026). 10.64898/2026.06.28.735106
The paper in one sentence
General-purpose tabular foundation models (e.g., TabICL, TabPFN) match or outperform specialized biological architectures for predicting cellular responses to genetic and chemical perturbations across single-cell, pseudobulk, and organismal scales.
Summary
This study evaluates Tabular Foundation Models (TFMs)—pretrained, domain-agnostic regression models—against state-of-the-art biology-specific models (e.g., PRESAGE, scGPT, scLAMBDA, STACK, Prophet) in four distinct settings: (1) cell-level cross-cell-type perturbation prediction, (2) pseudobulk perturbation prediction across five Perturb-seq datasets, (3) a genome-wide CRISPR screen in primary human CD4⁺ T cells, and (4) embryo-level cell-type composition prediction in a zebrafish developmental atlas. The authors demonstrate that TFMs, despite lacking biological pretraining, consistently achieve competitive or superior performance by leveraging PCA-based dimensionality reduction and in-context learning. The work suggests that strong posterior-predictive regression and effective featurization may outweigh the need for bespoke biological inductive biases in perturbation prediction tasks.
Personal highlights
TFMs outperform specialized models in pseudobulk prediction: TabICL and TabPFN achieve the lowest relative MSE and highest cosine similarity (top-20 DE genes) across all five Perturb-seq datasets, surpassing domain-specific models like PRESAGE, scGPT, and scLAMBDA. Notably, TabPFN is the only model with positively correlated predictions in the genome-wide CD4⁺ T cell screen, where other models perform near the mean baseline.
Cell-level prediction via optimal transport matching: The authors introduce a novel pipeline to adapt TFMs for single-cell prediction: they use optimal transport (OT) to match control cells between source and target cell types in PCA space, then train TFMs on matched pairs. This approach outperforms nearest-neighbor matching and even rivals STACK, a single-cell foundation model pretrained on large-scale scRNA-seq data.
Scalability to organismal-level predictions: In the zebrafish zscape atlas, TFMs predict cell-type composition changes under genetic perturbations with higher Spearman correlation and R² than Prophet, a model explicitly designed for this task. This demonstrates that TFMs generalize beyond transcriptomics to functional phenotypes.
Data efficiency and robustness: TFMs maintain strong performance with as little as 25% of training data, while specialized models like PRESAGE degrade more sharply. Diversity-aware support-set selection (e.g., farthest-point sampling) further improves TFM accuracy, suggesting that in-context example curation is key to their success.
Mechanistic insights into TFM success: Ablations reveal that TFM performance depends on projecting targets onto high-variance subspaces (e.g., top PCA components) rather than the specific decomposition method (PCA, NMF, or ICA). This highlights that capturing the dominant signal in the data is more critical than biological inductive biases.
Why should we care?
This preprint challenges a dominant paradigm in computational biology: that predicting cellular responses to perturbations (e.g., drug treatments or gene knockouts) requires highly specialized models with built-in biological knowledge. Instead, the authors show that general-purpose AI models, pretrained on synthetic tabular data, can perform just as well or even better than bespoke architectures, without any biology-specific fine-tuning. The results are compelling but not universal. TFMs struggle in noisy, high-variability settings (e.g., primary CD4⁺ T cells), where the signal-to-noise ratio is low. Moreover, their performance gains are task-dependent: they shine in pseudobulk prediction but are only competitive (not superior) in zebrafish composition prediction. Finally, while TFMs avoid biological pretraining, they still depend on careful decomposition of high-dimensional data (e.g., PCA), which may implicitly encode biological structure.
Scalable multi-group nonnegative spatial factorization for spatial genomics data with cell-type heterogeneity
Chumpitaz-Diaz et al. bioRxiv (2026). 10.64898/2026.06.29.735224
The paper in one sentence
This paper introduces smNSF, a scalable probabilistic framework that integrates spatial coordinates and cell-type labels via multi-group Gaussian processes to disentangle gene-driven spatial patterns from cell-type composition effects in spatial transcriptomics data.
Summary
The authors present scalable multi-group nonnegative spatial factorization (smNSF), which extends nonnegative spatial factorization (NSF) by incorporating cell-type labels through multi-group Gaussian processes (MGGPs). This allows the model to capture complex spatial variation in a cell-type-specific manner while enforcing nonnegativity for interpretability. To address scalability, they develop a locally conditioned Gaussian process (LCGP) approximation, reducing the computational complexity from cubic to linear in the number of inducing points, enabling training on datasets with hundreds of thousands of cells. Across seven diverse ST datasets (Slide-seqV2, MERFISH, osmFISH, 10x Visium), smNSF recovers sparse, interpretable spatial factors and organizes them into three specificity classes: cell-type enriched, cell-type specific, and universal, through cell-type conditional posteriors. The method reveals spatial programs (e.g., oligodendrocyte lineage, vascular programs, immune compartments) that are masked in standard analyses. Benchmarking on the DLPFC dataset shows competitive performance with state-of-the-art spatial domain identification methods, achieving the highest spatial autocorrelation (Moran’s I) among 20 methods
Personal highlights
Unified spatial and cell-type modeling: Integrates spatial coordinates and cell-type labels into a single probabilistic framework using MGGPs, enabling cell-type-aware spatial decomposition.
Scalable inference with LCGP: The LCGP approximation reduces computational complexity, allowing training on large datasets (tested up to 71,939 cells in 3D).
Diagnostic for group label interpretation: Reveals that scattered cell-type labels produce enriched conditional posteriors (61% of cell-type groups), while spatially confined region labels do not (0/28 groups), providing a criterion for when conditional analysis is informative.
Interpretable spatial programs: Identifies three classes of spatial factors (cell-type enriched, cell-type specific, universal) across diverse tissues, recovering known biological programs.
Competitive benchmark performance: Achieves the highest spatial autocorrelation among 20 methods on DLPFC and competitive clustering accuracy.
Other papers that peeked my interest and were added to the purgatory of my “to read” pile
Leveraging WNT Hyperactivation to Kill Colorectal Cancer While Rejuvenating Healthy Intestine
Multiomic screening platform uncovers the impact of histone mutations on chromatin and cell fate
Regenerative cell turnover coupled to a tumor-hostile niche protects naked mole-rats from cancer
Tumor Genotype Dictates Mitochondrial and Immune Vulnerabilities in Liver Cancer
TREND: A generalizable synthetic enhancer discovery platform for targeted immunotherapy
Senescence-directed nanotherapy ameliorates fibrosis and overcomes immune exclusion in cancer
Spatio-DARLIN enables robust and efficient in situ lineage tracing in mice at single-cell resolution
Macrophage-mediated brain-bone marrow crosstalk promotes chronic stress-induced glioma growth
Lactate binds and inhibits the innate immune sensor STING to promote tumor immune evasion
TROP2 targeting reveals therapy-driven cell state dynamics in colorectal cancer
Dual tumour–myeloid targeting of glioblastoma with GPNMB CAR-T cells
LIF-Induced Tumor Plasticity Establishes an Immunosuppressive Myeloid Niche in LKB1-Mutant Lung Cancer
A Chemically Defined Synthetic Cell Capable Of Growth And Replication
A dietary switch promotes sensory neuron–dependent cancer-associated cachexia
TranscriptFormer: A generative cell atlas across 1.5 billion years of evolution
Direct comparison of CRISPR knockout and interference with Perturb-seq
Score Distributions, Not Cells: Evaluating Single-Cell Perturbations Under Class Overlap
Cell-JEPA: Latent Representation Learning for Single-Cell Transcriptomics
Sex-Dimorphic Neural Memory Shapes Pancreatic Tissue Resilience
Aneuploidy selects for the acquisition of driver genes in breast cancer
Universal cell embedding provides a foundation model for cell biology
Epigenetic Reactivation of Lineage Differentiation to Target Leukemia
Diet–microbiome synergy underlies obesity-associated immunotherapy efficacy
Tissue tension fosters macrophage-driven lipid peroxidation-induced DNA damage
Cellular architecture and neighborhood-informed virtual spatial tumor profiling from histopathology
Evolving patterns of co-mutations from tumor initiation to metastatic progression
Thanks for reading.
Cheers,
Seb.



