2026-04-02
Gabriel Dewa, C. Mee Ling Munier, Sara Ballouz, Raymond Louie
The application of single-cell RNA sequencing (scRNA-seq) for biomarker discovery promises unprecedented resolution in identifying potential biomarkers by capturing and analysing cellular heterogeneity. Traditionally, biomarker discovery efforts within single-cell transcriptomics have primarily relied on conventional statistical approaches, particularly through the application of differential gene expression analysis, to identify candidate biomarkers. However, in recent years, with the rapid advancement and growing popularity of artificial intelligence and machine learning, their application in scRNA-seq biomarker discovery has become increasingly prominent. Currently, machine learning-based approaches for scRNA-seq biomarker discovery exhibit considerable methodological diversity, which can be distinguished by factors such as the level of discovery, choice of supervised learning algorithm, feature selection methods, classification metrics, and downstream biological analyses. This review provides a comprehensive overview of the current landscape of machine learning methods for scRNA-seq biomarker discovery, offering researchers a complete and detailed understanding of the field.
DOI: 10.3389/fbinf.2026.17673622026-04-01
Liping Zhou, Chenyang Fei, Quanxia Liu
Parkinson’s disease (PD) is a complex neurodegenerative disorder for which current treatments are often symptomatic and lack disease-modifying effects. The traditional Chinese medicine herb pair Tianma-Gouteng, composed of Gastrodia elata Bl (Tianma) and Uncaria rhynchophylla(Miq.) Miq. ex Havil. (Gouteng), has demonstrated clinical efficacy in treating PD motor symptoms, yet its multi-target mechanisms remain unclear. This study employs an integrated approach combining bioinformatics and computational chemistry to elucidate these mechanisms and identify key active components. Methods involved network pharmacology to identify active compounds and PD-related targets, followed by protein-protein interaction network analysis and functional enrichment. Molecular docking and 100-ns molecular dynamics (MD) simulations were utilized to evaluate the binding stability and dynamics of core component-target complexes. Additionally, Density Functional Theory (DFT) was conducted to analyze the electronic properties and reactivity of key compounds. Network pharmacology analysis identified 42 active components and 261 PD-related targets. Core targets identified were AKT1, TP53, and STAT3, which are involved in the regulation of PI3K-AKT signaling, mitochondrial apoptosis, and neuroinflammation. MD simulations demonstrated that quercetin (QU) and kaempferol (KA) formed highly stable complexes with AKT1 and TP53, exhibiting low average root-mean-square deviation (RMSD <0.2 nm), stable radius of gyration (Rg fluctuation <0.05 nm), and sustained protein-ligand hydrogen bonds. In contrast, complexes with 4–4′-hydroxybenzyloxy and 20-hexadecanoylingenol showed conformational instability, consistent with higher entropy penalties. DFT calculations revealed that QU and KA possess low HOMO-LUMO gaps, indicating high chemical reactivity, along with strong nucleophilic regions and intramolecular hydrogen bonds that facilitate target binding. The Tianma-Gouteng pair exerts anti-PD effects through the synergistic modulation of AKT1-mediated PI3K-AKT signaling, STAT3-driven neuroinflammation, and TP53-regulated apoptosis. Quercetin and kaempferol are identified as pivotal components due to their stable target binding and favorable electronic properties, providing a promising foundation for the development of novel PD therapeutics.
DOI: 10.3389/fbinf.2026.17962162026-03-30
Amal Alnouri, Andreas Hinterreiter, Christian Steinparz, Labinot Bajraktari, Sebastian Burgstaller-Muehlbacher, Markus Bauer, Gregorio Alanis-Lobato, Marc Streit
IntroductionFinding new uses for existing drugs, known as drug repurposing, is a widely adopted drug development strategy in the pharmaceutical industry. Computational drug repurposing leverages vast biomedical data to prioritize repurposing candidates. Once these candidates are prioritized, domain experts face the burden of evaluating their true potential.MethodsIn this work, we propose a visualization-based approach to address this challenge for a multimodal class of computational drug repurposing, where heterogeneous evidence modalities are integrated. We conducted a design study in close collaboration with domain experts, from which we derived a domain abstraction of the expert assessment process. Grounded in this abstraction, we developed an interactive visualization approach that explicitly models the expert reasoning process. We applied the proposed approach to create a prototype implementation, molIEreVIS, in the context of an operational drug repurposing pipeline. We used this prototype to collect qualitative feedback from domain experts actively engaged in assessing computational drug repurposing candidates.ResultsThe results demonstrate the potential of our approach to support insights and reasoning in this process and reveal directions for enhancements and future work.
DOI: 10.3389/fbinf.2026.17564592026-03-20
Hua Li, Yingchun Jiang, Jijia Li
ObjectiveThis study aims to explore the potential molecular mechanisms by which di (2-ethylhexyl) phthalate (DEHP) exposure induces pulmonary arterial hypertension (PAH).MethodsWe conducted differential expression analysis on multiple genomics datasets to pinpoint PAH-associated genes. Subsequently, an integrative approach combining machine learning algorithms and network toxicology was employed to examine the binding interactions between DEHP and the identified target proteins.ResultsOur analysis identified 60 genes as potential targets of DEHP in PAH. Further refinement using machine learning prioritized twelve core regulatory genes: ALKBH2, AOC2,BCL2L10,CTBP2,DNM2,ERLIN2,HPS6,RABGGTA,PON2,SLC4A7,SORT1, and PDE4D. Among these, HPS6, CTBP2,RABGGTA, SORT1,ALKBH2,BCL2L10, AOC2,and PON2 were significantly downregulated, whereas SLC4A7,PDE4D, ERLIN2,and DNM2 were markedly upregulated (P < 0.05).ConclusionThese findings demonstrate that DEHP promotes PAH pathogenesis by modulating specific genes and associated pathways. The twelve core genes identified through machine learning are proposed as key regulators in this process, providing crucial insights for future mechanistic investigation into DEHP-induced PAH.
DOI: 10.3389/fbinf.2026.17116372026-03-19
Reda Chahir, Salaheddine Redouane, Jacob Galan, Hicham Hboub, Lahoussaine Aserrar, Salma Chakir, Ahmed Salim Lahlou, Hinde Aassila, Rachid El Fatimy, Naoual Oukkache
DOI: 10.3389/fbinf.2026.18264092026-03-18
Yimin Shen, Xiaotian Xu, Xiaoxi Hao, Cuimin Sun, Wei Lan
IntroductionRapid diagnosis of bacterial pneumonia is crucial for clinical diagnosis and treatment, but traditional methods are time-consuming. The wide application of machine learning techniques in medical diagnosis provides an effective way to solve this problem. However, the complexity of medical datasets and the problem of class imbalance poses serious challenges to classical machine learning algorithms.MethodsAiming at the multiclass imbalanced problem in complete blood count (CBC) datasets, this study proposes a novel ensemble learning algorithm, Forest of Evolutionary Multi-Classifiers Based on Bagging with Error-Correcting Output Coding (Forest-EMCBE). The algorithm integrates Multi-Objective Genetic Algorithm, Error-Correcting Output Codes (ECOC), and balanced sampling strategy, which enhances the generalization ability of the classifiers through a three-layer integrated structure.ResultsTo validate the effectiveness of the proposed method, we trained the diagnostic model on a CBC dataset, which contains 1,457 samples and 4 different classes of bacterial pneumonia results, and compared it with 11 state-of-the-art algorithms. The experimental results demonstrate the superior performance of the Forest-EMCBE algorithm on the CBC dataset, outperforming all other compared algorithms.DiscussionBased on the Shapley value-based feature importance analysis method, this study dissects the contributions of key features to the prediction outcomes and further elucidates the differential impacts of features such as age, gender, and neutrophil percentage on predicting infections by different bacterial species.
DOI: 10.3389/fbinf.2026.17926432026-03-17
Tope Abraham Ibisanmi, Xiaotao Jiang, Mark Willcox, Naresh Kumar
The accelerating antimicrobial resistance (AMR) crisis continues to render more and more conventional antibiotics ineffective. Antimicrobial peptides (AMPs) are promising alternatives to traditional antibiotics due to their broad-spectrum activity, diverse mechanisms of action, and lower propensity for resistance. Traditional discovery approaches face limitations arising from the vast sequence space and the challenge of balancing efficacy with low toxicity. Addressing these challenges is critical for developing next-generation antimicrobial agents, and computational methods are increasingly driving progress. Public repositories, and techniques such as molecular docking enable in silico evaluation of peptide target interactions, identifying candidates with strong binding potential. Molecular dynamics (MD) simulations offer deeper insights into how AMPs disrupt membranes, form pores, or act synergistically, while Steered MD extends this to probing membrane penetration. Artificial intelligence (AI) methods, including machine learning and deep learning, capture complex sequence activity relationships, predict novel AMPs from genomic and metagenomic data, and design new peptides de novo using generative models. Despite rapid advances, most existing reviews treat these approaches in isolation, leaving a fragmented understanding of their interplay. This paper addresses that gap by unifying computational strategies, highlighting synergies, and critiquing limitations. Ultimately, integrating these methodologies offers a path toward more efficient AMP discovery to fight AMR.
DOI: 10.3389/fbinf.2026.17494042026-03-16
Zixuan Li, Zhiguo Yu, Peng Li
Identification of essential proteins is fundamental for understanding cellular processes and disease mechanisms. However, many existing computational methods do not adequately model dynamic expression activity and often underutilize global network context, which limits prediction accuracy. To address these issues, we propose a Correlation-guided Subgraph Graph Neural Network (CSGNN) for essential protein identification by integrating correlation-guided graph construction with attention-based representation learning. First, we derive an activity-aware expression matrix from periodic gene expression patterns, and we construct a weighted protein network by computing Pearson correlation coefficients between gene pairs. This correlation-guided network further defines first-order and second-order neighborhoods, which provide multi-scale subgraph contexts for each protein. Next, we employ a two-layer attention-based graph convolution to learn node embeddings by aggregating information within these correlation-defined neighborhoods. Finally, we form an interaction-aware node representation by integrating each protein embedding with its neighborhood context, and we use a lightweight multilayer perceptron to output an essentiality probability for each protein. Proteins are then ranked by the predicted scores to identify essential candidates. Experiments on yeast and E. coli datasets demonstrate that CSGNN consistently outperforms traditional baselines, indicating improved accuracy and robustness for essential protein identification.
DOI: 10.3389/fbinf.2026.1731178