← Back to Journals

PLoS Computational Biology

Publisher:
PLOS
ISSN:
1553-734X
Category:
MATHEMATICAL & COMPUTATIONAL BIOLOGY
Impact factor:
3.8

Feed status

10 parsed articles

Last update: Not fetched

Latest articles

Predictive modeling of gene expression and localization of DNA binding site using deep convolutional neural networks

2026-04-01

Arman Karshenas, Tom Röschinger, Hernan G. Garcia

by Arman Karshenas, Tom Röschinger, Hernan G. Garcia Despite the sequencing revolution, large swaths of the genomes sequenced to date lack any information about the arrangement of transcription factor binding sites on regulatory DNA. Massively Parallel Reporter Assays (MPRAs) have the potential to dramatically accelerate our genomic annotations by making it possible to measure the gene expression levels driven by thousands of mutational variants of a regulatory region. However, the interpretation of such data often assumes that each base pair in a regulatory sequence contributes independently to the overall gene expression. To enable the analysis of this data in a manner that accounts for possible correlations between distant bases along a regulatory sequence, we developed the Deep learning Adaptable Regulatory Sequence Identifier (DARSI). This convolutional neural network leverages MPRA data for training specific models for each operon to predict gene expression levels directly from raw regulatory DNA sequences. By harnessing this predictive capacity, DARSI systematically identifies transcription factor binding sites within regulatory regions at single-base pair resolution. To validate its predictions, we benchmarked DARSI against curated databases, confirming its accuracy in predicting known transcription factor binding sites. Additionally, DARSI predicted novel unmapped binding sites, paving the way for future experimental efforts to confirm the existence of these binding sites and to identify the transcription factors that target those sites. Thus, DARSI provides a new framework for MPRA experimental data analysis, it generates experimentally actionable predictions that can feed iterations of the theory-experiment cycle aimed at reaching a predictive understanding of transcriptional control. Here, we developed a deep learning approach—called DARSI—that leverages these massively parallel reporter assays to predict levels of gene expression from DNA sequences and help locate these important binding sites. By training our model to recognize DNA sequence patterns that affect gene expression, our method not only finds known binding sites with high accuracy, but also predicts new binding sites that call for future experimental scrutiny.

DOI: 10.1371/journal.pcbi.1014092

Network-based exploration of 4-(phenylsulfonyl)morpholine molecules for metastatic triple-negative breast cancer suppression

2026-03-31

Jung-Chen Su, Chen-Ling Lee, Fan-Wei Yang, Yan-Chih Chen, Te-Lun Mai

by Jung-Chen Su, Chen-Ling Lee, Fan-Wei Yang, Yan-Chih Chen, Te-Lun Mai Triple-negative breast cancer (TNBC) is an aggressive and heterogeneous subtype of breast cancer, with limited treatment options due to the absence of estrogen receptors, progesterone receptors, and human epidermal growth factor receptor 2 (HER2) expression. This characteristic renders TNBC resistant to hormone-based and HER2-targeted therapies, leaving cytotoxic chemotherapy as the predominant strategy and highlighting the urgency for novel interventions. In this study, we investigated the mechanism of action of GL24, a potent 4-(phenylsulfonyl)morpholine-based small molecule with selective tumor suppression effects on metastatic TNBC cells, while being ineffective against TNBC cells derived from the primary tumor site, using gene co-expression analysis. By considering the distinct phenotypic responses induced by GL24, we tailored our co-expression analysis approach, selecting gene pairs that exhibited differential co-expression in effective cells while excluding gene pairs that also showed differential patterns in non-effective cells. Constructing a co-expression network from these differential pairs, followed by enrichment analysis and functional annotation, revealed specific gene interactions and molecular pathways associated with GL24-mediated TNBC inhibition. These insights supported the previously established findings that showed convergence on apoptosis based on differentially expressed genes, while also providing complementary information by highlighting pathways involved in metabolic alterations, proliferation, and migration or invasion. This expanded understanding advances the knowledge of the mechanisms of GL24 in combating TNBC.

DOI: 10.1371/journal.pcbi.1014132

BIOPOINT: A particle-based model for probing nuclear mechanics and cell-ECM interactions via experimentally derived parameters

2026-03-31

Sandipan Chattaraj, Julius Zimmermann, Francesco Silvio Pasqualini

by Sandipan Chattaraj, Julius Zimmermann, Francesco Silvio Pasqualini Morphogenesis arises from biochemical and biomechanical interactions across multiple spatial and temporal scales. Experimental studies alone cannot fully resolve these dynamics, motivating computational models. Subcellular element modeling (SEM) is well suited for simulating emergent cellular and tissue morphologies, but traditional SEM frameworks do not explicitly include nuclear deformation or direct cell–extracellular matrix (ECM) interactions: capabilities typically associated with continuum approaches based on the finite-element method (FEM) approaches, FEM excels at modeling cell and tissue mechanics, but struggle to accommodate the large, non-linear deformations driven by local, geometry-changing events that define morphogenesis. Here, we introduce BIOPOINT, a particle-based framework that augments SEM with FEM-like mechanical capabilities by incorporating (1) a deformable, multi-particle nucleus capable of capturing nuclear stress and strain distributions and (2) an explicit ECM layer represented by structured static particles with tunable adhesive potentials. To ensure biological relevance, we calibrate BIOPOINT against single-cell indentation experiments (SKOV3). We then apply this calibrated parameter set, without additional refitting, to two independent scenarios: (i) cell (EC and hMSC) spreading on ECM micropatterns, capturing qualitative coupling between cell and nuclear shape; and (ii) confined migration (MDA-MB-231) through rigid constrictions, qualitatively reproducing the characteristic sequence of nuclear elongation and partial recovery. Differently from previous work, we present a full SEM model that uses heterogeneous particles to model nuclei, cells, and ECM via phase separation. By combining SEM’s strength in modeling emergent cell and tissue geometry with a mechanically sound handling of nuclear and ECM interactions, BIOPOINT provides a versatile platform for studying cell behaviors, like shape acquisition and migration through confinement that are relevant to morphogenesis. Implemented within the widely used, open-source LAMMPS ecosystem, BIOPOINT offers an accessible and extensible tool for the community.

DOI: 10.1371/journal.pcbi.1014113

CAPYBARA: A generalizable framework for predicting serological measurements across human cohorts

2026-03-30

Sierra Orsinelli-Rivers, Daniel Beaglehole, Tal Einav

by Sierra Orsinelli-Rivers, Daniel Beaglehole, Tal Einav The rapid growth of biological datasets presents an opportunity to leverage past studies to inform and predict outcomes in new experiments. A central challenge is to distinguish which serological patterns are universally conserved and which are specific to individual datasets. In the context of human serology studies, where antibody-virus interactions assess the strength and breadth of the antibody response and inform vaccine strain selection, differences in cohort demographics or experimental design can markedly impact responses, yet few methods can translate these differences into the value±uncertainty of future measurements. Here, we introduce CAPYBARA, a data-driven framework that quantifies how serological relations map across datasets. As a case study, we applied CAPYBARA to 25 influenza datasets from 1997-2023 that measured vaccine or infection responses against multiple influenza variants using hemagglutination inhibition (HAI). To demonstrate how a subset of measurements in each study can infer the remaining data, we withheld all HAI measurements for each variant and accurately predicted them with a 2.0-fold mean absolute error—on par with experimental assay variability. Although studies with similar designs showed the best predictive power ( e.g ., children data are better predicted by children than adult data), predictions across age groups, between vaccination and infection studies, and across studies conducted <10 years apart showed comparable 2‒3-fold accuracy. By analyzing feature importance in this interpretable model, we identified global cross-reactivity trends that can be directly applied in future longitudinal or vaccine studies to infer broad serological responses from a small subset of measurements.

DOI: 10.1371/journal.pcbi.1014129

An improved dataset for predicting mammal infecting viruses from genetic sequence information

2026-03-27

Tyler Reddy, Austin Schneider, Aaron R. Hall, Adam Witmer, Nick Hengartner

by Tyler Reddy, Austin Schneider, Aaron R. Hall, Adam Witmer, Nick Hengartner There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al. , increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

DOI: 10.1371/journal.pcbi.1014125

Deconvolving cell-type-specific gene expression profiles from bulk RNA-seq samples

2026-03-26

Sichen Zhu, Zhengqi Wang, Kevin D. Bunting, Peng Qiu

by Sichen Zhu, Zhengqi Wang, Kevin D. Bunting, Peng Qiu Bulk RNA sequencing (bulk RNA-seq) and single-cell RNA sequencing (scRNA-seq) are two important high-throughput sequencing platforms that have wide applications in biomedical research. Bulk RNA-seq reflects the average gene expression of all cells in the sample at a low experimental cost, whereas scRNA-seq enables transcriptomics profiling at a single-cell level, although with higher experimental costs. To integrate the strengths of both sequencing approaches and capitalize on the wealth of existing bulk RNA-seq datasets, we developed a U-Net-based deep learning algorithm, BLUE, to deconvolve bulk RNA-seq samples into cell-type proportions and cell-type-specific gene expression profiles. Built upon a U-Net backbone, BLUE leverages its powerful feature extraction and representation learning capabilities to achieve accurate predictions for cell-type-specific gene expression profiles, which significantly outperform existing deconvolution algorithms. Given the accurate prediction from BLUE, we developed an integrative framework for subtyping cancer patients and identifying cell-type-specific gene signatures that can function as prognostic biomarkers for cancer.

DOI: 10.1371/journal.pcbi.1014101

Bayesian network models to assess antimicrobial resistance patterns of <i>Streptococcus suis</i> isolated from swine production systems in the United States between 2014–2021

2026-03-26

Ruwini Rupasinghe, Brittany L. Morgan Bustamante, Rebecca C. Robbins, Maria J. Clavijo, Beatriz Martínez-López

by Ruwini Rupasinghe, Brittany L. Morgan Bustamante, Rebecca C. Robbins, Maria J. Clavijo, Beatriz Martínez-López Multidrug resistance (MDR) is frequently evident in Streptococcus suis, generating distinct antimicrobial resistance (AMR) profiles, which limits the effective antimicrobial drug (AMD) options against S. suis in pigs and humans. Despite its significance, there is a lack of studies and pertinent methodologies that uncover complex interactions among AMDs and associated resistance patterns. This study aimed to identify associations between phenotypic resistance patterns of S. suis isolates from swine production systems in the United States against common AMDs using Bayesian network analysis (BNA). Data from 259 unique S. suis isolates collected from 91 farms were included. Phenotypic susceptibility interpretations (resistance vs susceptible) of minimum inhibitory concentrations (MICs) were evaluated for 13 commonly used AMDs: ceftiofur (CEF), penicillin (PEN), enrofloxacin (ENR), gentamicin (GEN), neomycin (NEO), spectinomycin (SPC), sulfadimethoxine (SUL), tiamulin (TIA), tilmicosin (TIL), clindamycin (CLN), chlortetracycline (CHL), oxytetracycline (OXY), and tetracycline (TET). BNA was conducted using the R package bnlearn to identify joint resistance patterns and estimate conditional dependencies among resistance outcomes. Results revealed a high prevalence of MDR: 248 isolates (95.6%) were resistant to more than one AMD, and 209 isolates (80.7%) were resistant to at least one AMD in three or more classes. The Bayesian network comprised of 11 edges connecting 13 AMD nodes, highlighting statistical dependencies between AMDs resistances. PEN, TIA, and TIL were the most central nodes, with PEN connected to SUL, TIA, GEN, and CEF; TIA to PEN, SPC, TIL, and CLN; and TIL to SUL, TIA, CLN, and OXY. Other associations included CEF–SPC, TET–CLN, CEF–ENR, and OXY–CHL. These relationships implicate systematic dependencies between AMDs and may have resulted from mechanisms like cross-resistance and co-resistance. While these relationships are statistically derived and hypothesis-generating, they underscore the importance of understanding AMR patterns in guiding more effective AMD use. This approach can help prevent overuse, reduce treatment failures, and support AMR mitigation efforts for improved animal and public health outcomes.

DOI: 10.1371/journal.pcbi.1014117

Accounting for sensitivity of latent learning to behavioral statistics with successor representations

2026-03-24

Matheus Menezes, Xiangshuai Zeng, Sen Cheng

by Matheus Menezes, Xiangshuai Zeng, Sen Cheng Latent learning experiments were critical in shaping Tolman’s cognitive map theory. In a spatial navigation task, latent learning means that animals acquire knowledge of their environment through exploration, such that pre-exposed animals learn faster on a subsequent learning task than naive ones. This enhancement has been shown to depend on the design of the pre-exposure phase. Here, we hypothesize that the deep successor representation (DSR), a recent computational model for cognitive map formation, can account for the modulation of latent learning because it is sensitive to the statistics of behavior during exploration. In our model, exploration aligned with the future reward location significantly improves reward learning compared to random, misdirected, or no exploration, as reported by experiments. This effect generalizes across different action selection strategies. We show that these performance differences follow from the spatial information encoded in the structure of the DSR acquired in the pre-exposure phase. In summary, this study sheds light on the mechanisms underlying latent learning and how such learning shapes cognitive maps, impacting their effectiveness in goal-directed spatial tasks.

DOI: 10.1371/journal.pcbi.1014131