Weekly reads 27/7/2026
Designing life, decoding mutations, and a cancer that spreads like a parasite
This week’s reads encompasses a study, which introduces GO-CRE, a reinforcement learning framework that reconstructs sequence-grammar trajectories to design interpretable and tunable cis-regulatory elements (CREs), revealing how regulatory grammars emerge during optimization and enabling real-time corrections to avoid unproductive paths. Meanwhile, a critical assessment of SComatic exposes the challenges of detecting somatic mutations from single-cell RNA-seq in mosaic tissues, offering methodological improvements to reduce false positives and bias toward highly expressed genes. In synthetic biology, researchers achieved a milestone by engineering a chemically defined synthetic cell capable of growth, replication, and Darwinian selection, paving the way for controlled studies of life’s minimal requirements. Finally, a surprising discovery in ecology: brown bullhead catfish melanoma is identified as the first transmissible cancer in fish, demonstrating how cancer can evolve into a parasitic lineage.
Preprints/articles that I managed to read this week
Reconstructing sequence-grammar trajectories enables interpretable and tunable cis-regulatory element design
Ma et al. bioRxiv (2026). 10.64898/2026.07.10.737719
The paper in one sentence
The authors introduce GO-CRE, a reinforcement learning framework that makes synthetic cis-regulatory element (CRE) design interpretable and tunable by reconstructing and steering sequence-grammar trajectories during optimization.
Summary
This study presents GO-CRE, a closed-loop framework that combines HybriDNA—a computationally efficient hybrid Transformer-Mamba2 DNA language model—with reinforcement learning (RL) and trajectory-level interpretation to design synthetic CREs with cell-type-specific activity. By deconvolving iterative sequence changes into k-mer and motif-level features, GO-CRE reconstructs a sequence-grammar landscape that reveals distinct phases of optimization: search (broad exploration), commitment (directional movement toward target-associated features), and optimization (stabilization with diverse feature acquisition). In HepG2 and K562 cells, trajectories progressively acquire lineage-aligned regulatory grammars (e.g., HNF/FOXA motifs in HepG2, GATA/RUNX in K562), while SK-N-SH trajectories remain confined to local, unproductive basins. Critically, trajectory analysis in HepG2 identified a low-complexity polyG trap; introducing a polyG penalty redirected optimization toward hepatocyte-associated features, demonstrating the framework’s tunability. Lentiviral MPRA validation confirmed that generated CREs exhibit cell-type-specific activity, with HepG2 designs showing higher average reporter activity than endogenous CREs despite distinct, more compact motif organizations (e.g., a recurrent HNF1B–FOXA1–HNF4A arrangement). The work establishes sequence-grammar trajectory reconstruction as a foundation for both interpretable synthetic CRE design and systematic analysis of regulatory grammar dynamics.
Personal highlights
Interpretable RL for CRE design: GO-CRE transforms black-box optimization into an observable process by reconstructing sequence-grammar trajectories, enabling diagnosis of failure modes (e.g., polyG traps) and active steering of design.
Phased optimization dynamics: Identifies three stages—search, commitment, and optimization—with coordinated transitions between sequence-update programs, revealing how regulatory grammar emerges during RL.
Trap diagnosis and correction: In HepG2, trajectory analysis uncovered a polyG-associated local basin; penalizing polyG features redirected optimization toward productive HNF/FOXA motif programs, improving MinGap scores.
Experimental validation: LentiMPRA confirmed target-selective activity of generated CREs in K562 and HepG2, with HepG2 designs outperforming endogenous CREs in average activity while retaining sequence diversity.
Lineage-specific motif convergence: Generated CREs converge on compact, lineage-aligned motif organizations (e.g., HNF1B–FOXA1–HNF4A in HepG2) distinct from the heterogeneous architectures of endogenous CREs.
Why should we care?
This work addresses how to rationally design regulatory DNA, the genetic switches controlling gene expression, rather than relying on trial-and-error or opaque algorithms. By making optimization trajectories visible, GO-CRE reveals why some designs succeed (e.g., escaping low-complexity traps to adopt lineage-specific motifs) and why others fail (e.g., SK-N-SH’s confinement to unproductive grammar basins). The ability to diagnose and correct unproductive paths in real time (e.g., via polyG penalties) is a major advance over endpoint-only optimization. However, the uneven success across cell lines (strong in HepG2/K562, weak in SK-N-SH) underscores that context matters: the framework’s effectiveness depends on the quality of training data and the cellular environment.
Somatic mutation inference from single-cell transcriptomics: A survey in the esophagus
Méndez-Alejandre et al., bioRxiv (2026). 10.64898/2026.07.07.737010
The paper in one sentence
This study evaluates the SComatic algorithm for detecting somatic mutations from single-cell RNA sequencing in the genetically mosaic esophageal epithelium, revealing substantial technical limitations but also providing a customized filtering approach that improves mutation detection specificity.
Summary
The authors apply the SComatic algorithm to public scRNA-seq datasets from human and mouse esophageal samples to assess its ability to detect somatic mutations in a polyclonal, normal tissue. They find that initial variant calls are overwhelmed by technical artifacts (61% shared across libraries) and germline variants. Through a customized filtering pipeline—including removal of repetitive region calls and variants affecting too many cells—they reduce candidate mutations to 11% (mouse) and 2% (human) of the initial calls. While these filtered mutations allow rough clonal mapping to differentiation trajectories, they are biased toward highly expressed housekeeping genes, missing many known cancer drivers like NOTCH1 and TP53. The study demonstrates that low read depth and sparse cellular sampling favor detection of passenger mutations and hinder driver mutation phenotypic inference, showcasing current limitations while offering methodological guidance for future research.
Personal highlights
High false positive rate: Most initial scRNA-seq variant calls are technical artifacts or germline variants, not true somatic mutations, highlighting the need for rigorous filtering.
Customized filtering pipeline: A tailored approach dramatically enriches for candidate somatic mutations, reducing noise from repetitive regions and widespread variants.
Bias toward highly expressed genes: scRNA-seq mutation detection favors transcripts with high expression, missing many cancer-associated driver genes that are lowly expressed.
Limited clonal resolution: Sparse sampling and low read depth lead to underestimation of clone sizes and preclude fine-grained subclonal analysis.
Methodological benchmark: The study provides actionable recommendations for improving somatic mutation inference from scRNA-seq in mosaic tissues, including multi-region sampling and repeated sequencing.
Why should we care?
This work offers a critical, balanced assessment of the current state of somatic mutation detection from single-cell transcriptomics. While the approach holds promise for studying clonal evolution in normal tissues, the study reveals significant limitations: high false positive rates from technical noise, a strong bias toward highly expressed genes, and an inability to detect many known driver mutations. The customized filters improve specificity, but sensitivity remains low—detecting roughly 1 mutation per 3 cells when genomic studies show hundreds to thousands per cell in the esophagus.
A Chemically Defined Synthetic Cell Capable Of Growth And Replication
Gaut et al. bioRxiv (2026). 10.64898/2026.07.01.735724
The paper in one sentence
Researchers engineered a synthetic cell with a fully defined 90kb genome that completes a multi-generation cell cycle, including genome replication, growth via genetically encoded feeding, and division, while demonstrating selection and competition.
Summary
This study presents a synthetic minimal cell constructed from chemically defined, non-living components that exhibits key hallmarks of life. The cell encapsulates a 90kb genome distributed across seven plasmids, encoding proteins for translation, transcription, genome replication (via Phi29 polymerase), and membrane growth. Growth is achieved through genetically encoded liposome fusion, mediated by α-hemolysin (αHL) protein expressed inside the cell interacting with Ni-NTA lipid tags on feeder liposomes. The authors demonstrate five generations of a complete cell cycle with mechanical division, and later achieve genetically encoded division through protein crowding on the membrane surface. Critically, they show that synthetic cells with advantageous mutations (stronger promoters for αHL) grow faster, produce more offspring, and outcompete slower-growing cells, demonstrating selection and competition for resources in resource-limited environments.
Personal highlights
First complete synthetic cell cycle: Demonstrated five generations of growth, genome replication, and division in a synthetic cell with a fully defined chemical composition.
Genetically encoded growth: Synthetic cells express αHL protein that enables fusion with feeder liposomes, coupling gene expression to membrane growth and nutrient uptake.
Genetically encoded division: Achieved division without a cytoskeleton by inducing membrane curvature through protein crowding via streptavidin binding to membrane-displayed tags.
Darwinian selection in synthetic cells: Cells with a stronger promoter for αHL (T7Max) outcompete those with a weaker promoter, showing that beneficial mutations can spread through a population.
Resource competition: In limited feeder liposome conditions, faster-growing cells gain a progressive advantage, illustrating how competition can drive selection in synthetic populations.
Why should we care?
This work represents a significant milestone in synthetic biology by demonstrating that life-like behaviors (growth, replication, selection, and competition) can emerge from a chemically defined system built from non-living components. Unlike previous synthetic cell attempts that relied on ill-defined extracts or external manipulation, this cell’s composition and behavior are fully specified, enabling precise engineering and modeling. The ability to couple genotype to phenotype through genetically encoded growth and division allows for evolutionary dynamics to be studied in a controlled, minimal system. While these synthetic cells are not yet autonomous (requiring external feeding and division triggers) and lack the complexity of natural cells, they provide a powerful platform for understanding the minimal requirements for life and could ultimately enable the design of custom cellular systems for biotechnological applications.
Brown bullhead catfish melanoma represents a novel transmissible cancer
Curd et al. Nature (2026). 10.1038/s41586-026-10828-6
The paper in one sentence
Researchers discovered that melanomas in brown bullhead catfish from Lake Memphremagog represent the first documented transmissible cancer in fish, with cancer cells spreading between individuals like a parasitic lineage.
Summary
This study investigated the unusually high prevalence (23–37%) of melanomas in brown bullhead catfish (Ameiurus nebulosus) in Lake Memphremagog, a freshwater lake spanning Vermont (USA) and Quebec (Canada). Using whole-genome sequencing of tumor and matched normal tissues, the authors found that tumor mitochondrial and nuclear genomes were more closely related to each other across different fish than to their respective hosts. Hundreds of thousands of genetic variants were shared among tumor samples but absent from host fish, vastly exceeding levels seen in conventional cancers. Phylogenetic analyses of both mitochondrial and nuclear genomes revealed that tumor samples formed a monophyletic clade distinct from host tissues and unaffected fish, indicating a single clonal origin. Extensive screening ruled out viral or microbial agents as the cause, supporting the conclusion that the cancer cells themselves act as the infectious entity. This makes brown bullhead melanoma the fourth documented naturally occurring transmissible cancer (after dogs, Tasmanian devils, and bivalves) and the first in fish or a freshwater ecosystem. The study also explores potential transmission mechanisms, including physical contact during spawning or environmental transfer via water or sediment.
Personal highlights
First transmissible cancer in fish: Genomic evidence confirms that brown bullhead melanoma is a clonal cancer lineage that transmits between individuals, marking the first such case in fish and the first in a freshwater ecosystem.
Genomic divergence of tumors from hosts: Tumor mitochondrial and nuclear genomes are more closely related to each other across different fish than to their hosts’ genomes, with 245,189 tumor-specific SNVs shared among samples, far exceeding expectations for conventional cancers.
No microbial cause identified: Comprehensive screening for viruses and microorganisms found no unique agents associated with tumors, strongly supporting the hypothesis that the cancer cells themselves are the transmissible agent.
Phylogenetic evidence of clonality: Both mitochondrial and nuclear phylogenetic trees show tumor samples forming a distinct monophyletic clade, separate from host tissues and unaffected fish, indicating a single founder event.
Ecological and evolutionary implications: The rapid spread of the cancer (30% prevalence by 2015) suggests efficient transmission, potentially linked to reproductive behaviors, aquatic environments, or weakened immune systems, raising concerns about its impact on fish populations.
Why should we care?
This study expands our understanding of transmissible cancers beyond mammals and bivalves to include fish in freshwater ecosystems. The key takeaway is that cancer, typically viewed as a non-communicable disease, can in exceptional cases evolve to behave like a parasite, spreading between individuals. Critically, while the genomic evidence is compelling, the study leaves important questions unanswered: the exact transmission mechanism remains hypothetical (e.g., physical contact during spawning or environmental transfer), and the long-term ecological impact on brown bullhead populations is uncertain. Additionally, it is unclear whether this cancer lineage exists in other populations or water bodies.
Other papers that peeked my interest and were added to the purgatory of my “to read” pile
LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration
Tertiary lymphoid structures harbour stem-like tumour-specific T cells
Subnuclear genome compartmentalization controls bivalent chromatin activity
Muon Reduces the Training Cost of Regulatory DNA Transformers
Using Deep Learning to predict replication timing reveals baseline control of genomic DNA sequence
Genetic background sets the trajectory of experimental cancer evolution
Promoter strength and position govern promoter competition through transcript-dependent insulation
scLEMBAS: Context-Aware Modeling of Signaling Pathway Activity at Single-Cell Resolution
Division of Synthetic Cells Using a Genomically Encoded One-Protein Divisome
Uncovering spatially resolved functional genomics with CRISPR screen sequencing
Genome instability triggers intercellular DNA transfer between human cells
Spatialproteomics: an interoperable toolbox for analyzing highly multiplexed fluorescence image data
Analog intrinsic recoding measures RNA dynamics without chemical conversion
Comprehensive benchmarking of RNA velocity methods across single-cell datasets
Toward generalizable and interpretable AI in regulatory genomics
Thanks for reading.
Cheers,
Seb.


