2026-07-25
Abstract Wild silkmoths (Saturniidae) are one of the most emblematic and well-studied families of moths. Yet, the absence of a robust phylogenetic framework based on a comprehensive taxonomic sampling impedes our understanding of their evolutionary history. We sequenced and analyzed 1024 ultraconserved elements (UCEs) and their flanking regions to infer the relationships among 338 species of Saturniidae representing all described subfamilies, tribes, and genera. We investigated systematic biases in genomic data and performed dating and historical biogeographic analyses using extinction free and state-dependent speciation and extinction models to document the evolutionary history of wild silkmoths in space and time. Using Gene Genealogy Interrogation, we showed that saturation of nucleotide sequence data blurs our understanding of early divergences and first biogeographic events. Our results support a Neotropical origin of saturniids, but remain undecisive with respect to the extent of the ancestral range for the family. The “traditional” hypothesis of an origin restricted to the Neotropics is supported by models that account for founder-effect dispersal, but a new alternative scenario recovered from other models is proposed and considered a better fit to the hypothetical modes of diversification in these moths. It estimates a broader ancestral range covering the Neotropical, West Nearctic and East Palearctic bioregions and emphasizes the critical role of Beringia as a route between the New World and the Old World during the Eocene. Interestingly, all the early branching lineages of Saturniidae that diversified into today’s recognized subfamilies are characterized by a very strong geographic conservatism, except for one noticeable exception, the Saturniinae subfamily, that is now present on all continents but Antarctica. Overall, our results provide a framework for in-depth investigations into the spatial and temporal dynamics of all saturniid lineages and for the integration of their evolutionary history into further global studies of biodiversity and conservation. Rather unexpectedly for a taxonomically well-known family such as Saturniidae, the proper alignment of taxonomic divisions and ranks with our phylogenetic results leads us to propose substantial rearrangements of the family classification, including the description of one new subfamily and two new tribes.
DOI: 10.1093/sysbio/syag0562026-07-21
Abstract Recent advances in systematics and evolutionary biology have revealed extraordinary complexity in phylogenetic diversification, challenged our traditional views of the Tree of Life, and changed our understanding of the nature of phylogenetic entities themselves. These empirical advances call for an updated conceptual basis for phylogenetics. Phylogeny is best conceptualized as a complex multidimensional system – encompassing spatial, temporal, and hierarchical dimensions. The hierarchical dimension confers important but poorly studied influences – such as emergence, constraint, non-linearity, and non-extrapolationism – on phylogenetic possibilities. Spatially and temporally, phylogenetic diversification patterns should be viewed as resulting from multimodal phenomena – i.e., as a consequence of both vertical and horizontal modes of evolution. Further, phylogeny should not be viewed as a single absolute history – it comprises interacting, multi-layered lineage histories (i.e., genomes, cells, organisms, species). Causal interactions and feedback dynamics amongst these dimensions represent a classic biocomplexity system. Therefore, we suggest that this is an opportune time for development of a phylogenetic systems theory.Because our view of phylogeny comprises a complex interacting system of multiple lineage histories, we also argue that a process ontology provides the logical foundation for contemporary phylogenetics. A process ontology framework also reconciles recent observations that the boundaries of biological entities (including the boundaries of genomes, cells, organisms, and species lineages) are fuzzy and open to varying degrees of exchange at various points in their histories. While these fuzzy boundaries accommodate the biocomplexity and multimodal evolutionary processes described above, they also pose challenges for our notions of phylogenetic entities themselves and their biological individuality – leading to exciting new areas of study in systematics, theoretical biology and philosophy of biology.
DOI: 10.1093/sysbio/syag0532026-07-10
Abstract How stereotypical, and hence predictable, are evolutionary and accumulation dynamics? Here we consider processes – from genome evolution to cancer progression – involving the irreversible accumulation of binary features (characters). We seek models of how these characters evolve in the form of transition networks, describing transitions between sets of characters, that reflect the simplest possible sets of character dynamics that can explain all the observations. A transition network supporting a single, deterministic dynamic pathway is maximally simple and lowest cost, and branches (corresponding to different possible “next steps” for evolution) increase cost, particularly if these branches are “deep”, occurring at early stages in the dynamics. In this sense, the optimal description measures how stereotypical the evolutionary or accumulation process is – how predictable are its dynamics in independent samples or lineages. The problem is solvable in polynomial time for cross-sectional observations by building on an existing method, and we provide a polynomial-time estimate in the more general case of pairs of observed states. We use this approach to define a “stereotypy index” reflecting the extent of evolutionary predictability, and to efficiently estimate likely orderings of evolutionary events, common precursor steps, relationships between characters, and pathways of evolution. We demonstrate use cases in the evolution of antimicrobial resistance, organelle genomes, squamate and human morphology, and cancer progression, demonstrating that evolution in many cases evolution is significantly more stereotypical than expected from random character evolution. We provide a software implementation at https://github.com/StochasticBiology/hyperDAGs .
DOI: 10.1093/sysbio/syag0522026-07-10
Abstract Macroevolutionary studies have shown that the shape of phylogenetic trees differs in space, time, and between taxa. It is commonly assumed that these differences in tree shape reflect variability in the underlying ecological and evolutionary processes that produced them, and mechanistic eco-evolutionary models are increasingly used to explore this link. A concern in this context is whether conclusions drawn from such mechanistic models are robust to idiosyncrasies in how eco-evolutionary processes are formalized in the models. Here, we use eight mechanistic macroevolutionary models to study how 52 metrics of phylogenetic tree shape respond to variation in the strength of five fundamental processes: competition, dispersal, environmental filtering, niche conservatism, and speciation. We find that models agree on how some tree metrics respond to changes in these processes, in particular dispersal and speciation. However, no tree metric uniquely correlated with a single process, suggesting that single tree metrics have limited utility as shortcuts for inferring the underlying eco-evolutionary processes. Moreover, while it was possible to infer the underlying processes if the data-generating model was known, inference was not consistent across the different models. We conclude that the relationship between phylogenetic patterns and eco-evolutionary processes in macroevolutionary analysis is likely sensitive to the structural and mechanistic details of how a given process is implemented within models.
DOI: 10.1093/sysbio/syag0502026-07-09
Abstract Resolving deep phylogeny is complicated by systematic errors and biological sources of gene tree discordance, a challenge that is clearly presented by Collembola, one of the earliest diverging hexapod lineages. To address the long-standing conflict among the four collembolan orders, we expanded taxon sampling (113 species) and marker representation (4,070 single-copy orthologs), conducting a systematic evaluation of potential error sources. Using both concatenation and multispecies coalescent approaches under site-homogeneous (LG) and site-heterogeneous (CAT-PMSF) models, we recovered two primary competing topologies: a Neelipleona-first and an Entomobryomorpha-first hypothesis. Pronounced branch-length heterogeneity across orders and two exceptionally short internal branches suggested susceptibility to systematic error. Analyses of empirical data, combined with extensive simulations, showed that the Neelipleona-first topology arises under conditions known to induce long-branch attraction, including branch-length imbalance, compositional heterogeneity, and model misspecification, with fast-evolving and compositionally constrained sites further amplifying these artifacts. Coalescent simulations demonstrated that incomplete lineage sorting and gene tree estimation error jointly account for much of the deep gene tree-species tree discordance. In contrast, analyses using site-heterogeneous models and multispecies coalescent approaches, both intended to reduce systematic errors, consistently supported the Entomobryomorpha-first topology, recovering Entomobryomorpha + (Symphypleona + (Poduromorpha + Neelipleona)). Our findings clarify the mechanistic origins of phylogenomic conflict in Collembola and highlight the need to jointly consider systematic error and ILS when resolving ancient radiations. We propose the name 'Brachyantennamorpha' for the clade Poduromorpha + Neelipleona.
DOI: 10.1093/sysbio/syag0492026-07-08
Abstract High-throughput sequencing data, such as target capture, RNA-Seq, genome skimming, and high-depth whole genome sequencing, are used for phylogenomic analyses. Integrating these mixed data types into a single phylogenomic dataset requires several bioinformatic tools and significant computational resources. Here, we present Captus , a novel pipeline to analyze mixed data efficiently. Captus assembles these data types, searches for loci of interest, and produces paralog-filtered alignments. If reference target loci are not available for the studied taxon, Captus can also be used to discover new putative homologs via sequence clustering. Compared to other software, Captus allows the recovery of a greater number of more complete loci across more species. We apply Captus to assemble a comprehensive dataset, comprising the four types of sequencing data for the angiosperm order Cucurbitales, a clade of about 3,100 species in eight mainly tropical plant families, including begonias (Begoniaceae) and gourds (Cucurbitaceae). Our phylogenomic results support the currently accepted circumscription of Cucurbitales except for the position of the holoparasitic Apodanthaceae, which group with Rafflesiaceae in Malpighiales. A subset of mitochondrial gene regions supports the earlier divergence of Apodanthaceae in Cucurbitales. However, the nuclear regions and majority of mitochondrial regions place Apodanthaceae in Malpighiales. Within Cucurbitaceae, we confirm the monophyly of all currently accepted tribes but also reveal hybridization and incomplete lineage sorting both in Cucurbitales and within Cucurbitaceae. We show that contradicting results among earlier phylogenetic studies in Cucurbitales can be reconciled when accounting for gene tree conflict and demonstrate the efficiency of Captus for complex datasets.
DOI: 10.1093/sysbio/syag0462026-06-10
DOI: 10.1093/sysbio/syag0432026-06-08
<span class="paragraphSection"><div class="boxTitle">Abstract</div>Phylogenomic discordance is pervasive and cannot always be resolved by increasing the amount of sequencing data alone. Biological processes such as polyploidy, hybridization, and incomplete lineage sorting are major contributors to discordance and must be accounted for to avoid misleading evolutionary interpretations. To better understand how these processes influence phylogenetic reconstruction, we conducted a comprehensive phylogenomic study in the complex genus <span style="font-style:italic;">Packera</span>. With over 90 species and varieties, 40% of which exhibit polyploidy, aneuploidy, or other cytological complexities, <span style="font-style:italic;">Packera</span> presents significant challenges for phylogenetic reconstruction. Given these complexities, we assessed different published paralog processing methods on the resulting evolutionary relationships and phylogenetic support of this group. We then applied three of these methods to evaluate their impact on tree topology and our understanding of <span style="font-style:italic;">Packera</span>’s evolutionary history by constructing a time-calibrated phylogeny, reconstructing historical biogeography, and testing for ancient reticulation. Phylogenetic outcomes varied based on the paralog processing method used, with no method outperforming others. Our findings highlight the large impact of orthology inference and paralog processing on phylogenomic analyses, particularly in polyploid-rich groups such as <span style="font-style:italic;">Packera</span>, and we offer guidance on methodological impacts along with practical recommendations. We note that gaining a robust understanding of <span style="font-style:italic;">Packera</span>'s evolutionary history requires more than computational approaches alone. While technological advancements have greatly expanded our ability to analyze genomic data, effective phylogenomic research still relies on strong taxon sampling and detailed species knowledge. Without careful attention to biological context, phylogenomic studies risk misinterpreting evolutionary history and processes. By integrating genomic results with knowledge of the study system, we can begin to improve the accuracy of evolutionary reconstructions and gain deeper insights into the complex history of plant diversification.</span>
DOI: 10.1093/sysbio/syag0412026-05-29
<span class="paragraphSection"><div class="boxTitle">Abstract</div>Biogeography is intrinsically linked to evolution as a process and to systematics as a practice. Phylogenetic biogeography, in particular, studies the distribution of life in space over time through the lens of common ancestry. Over the past century, new biological and geological discoveries, theoretical frameworks, and methodological techniques revolutionized how we understand why species have come to live where they do. This perspective piece orients readers to major advances, changes, and conflicts from the phylogenetic biogeography literature, much of which was centered around articles published in this very journal. As part of our survey, we also highlight areas that were historically active, remain biologically significant, and deserve renewed attention.</span>
DOI: 10.1093/sysbio/syag0422026-05-27
<span class="paragraphSection"><div class="boxTitle">Abstract</div>Phylogenetic trees are often inferred from protein sequences sampled from diverse taxa across the tree of life. The compositions of these amino acid sequences may be heterogeneous across both sites and branches, particularly if deep phylogenetic divergences are the focus. Under some conditions, failure to model this compositional heterogeneity can lead to phylogenetic artefacts. However, the computational cost of phylogenetic inference with models accounting for compositional heterogeneity can be prohibitive. The originally proposed site-and-branch-heterogeneous GFmix model accounts for changing relative frequencies of G, A, R, and P (GARP) vs. F, Y, M, I, N, and K (FYMINK) amino acids resulting from extreme variation in G+C content among taxa. This GFmix model modifies a fitted site-heterogeneous profile mixture model in a branch-specific manner using parameters that reflect branch-specific amino acid compositions. This approach has been shown to improve likelihoods and reduce compositional artefacts. However, the original implementation of the model includes constraints which may sacrifice accuracy for computability and is limited to modeling variation in GARP/FYMINK composition. Here we investigate the properties of the original GFmix model in greater depth and present several improvements to the model. The improved GFmix models permit fewer constraints on branch-specific composition parameters, allow modeling of user-defined compositional heterogeneity, and provide for full maximum-likelihood optimization of parameters. We have also developed new methods for detecting compositional heterogeneity directly from sequence data. Analyses of simulated site-and-branch-heterogeneous data indicates that the improved GFmix models better estimate branch-specific compositions and branch lengths in heterogeneous trees. We applied the various versions of the GFmix model to a real dataset with known compositional heterogeneity artefacts. We find that the most complex GFmix model with full maximum likelihood parameter optimization consistently supports the correct tree over the artefactual tree with improved likelihoods. All implementations of the GFmix model and related scripts are available from <a href="https://www.mathstat.dal.ca/~tsusko/software.html">https://www.mathstat.dal.ca/~tsusko/software.html</a>.</span>
DOI: 10.1093/sysbio/syag0402026-05-22
<span class="paragraphSection"><div class="boxTitle">Abstract</div>Species trees need to be dated for many downstream applications. Some scalable molecular dating methods take a phylogenetic tree with branch lengths in substitution units, as well as a set of calibrations, as input and convert the branch lengths of the species tree to time units, while being consistent with the pre-specified calibrations. When dating species trees from multi-locus genome-scale datasets, the branch lengths and sometimes the topology of the species tree are estimated using concatenation. However, concatenation does not address gene tree heterogeneity across the genome. While Bayesian dating methods can address some forms of gene tree heterogeneity, such as incomplete lineage sorting, they are not scalable to large numbers of species. In this paper, we introduce a new scalable pipeline for dating species trees that addresses gene tree discordance for both topology and branch length estimation. The pipeline uses discordance-aware methods that account for incomplete lineage sorting for estimating the topology and branch lengths, and maximum likelihood-based methods for the dating step. Our simulation study on datasets with gene tree discordance shows that this pipeline produces more accurate and less biased date estimates than pipelines that use concatenation. Furthermore, it is substantially more scalable and can handle datasets with thousands of species and genes. Our results on two biological datasets demonstrate that this new pipeline improves the inference of node ages and branch lengths for certain nodes, particularly those closer to the tree tips, and improves downstream analyses of diversification.</span>
DOI: 10.1093/sysbio/syag038