Editorial: Multi-modal learning with large-scale models
2026-03-13
Xianmin Wang, Jing Li
DOI: 10.3389/fnbot.2026.1816779Feed status
7 parsed articles
Last update: Not fetched
2026-03-13
Xianmin Wang, Jing Li
DOI: 10.3389/fnbot.2026.18167792026-03-05
Canyang Liu, Yichen Liu, Yongqi Zhou, Buqin Su
Balancing exploration and exploitation remains a fundamental challenge in reliable mobile robot control, as conventional policies often converge on suboptimal behaviors. Inspired by the brain's division of labor for adaptive control, we propose SpikeAEC, a fully spiking, neuromodulated Actor-Explorer-Critic architecture designed to address this dilemma online within a closed-loop system. SpikeAEC comprises three specialized subnetworks operating in parallel: the Actor, inspired by the basal ganglia, proposes exploitative actions; the Explorer, modeled after the ACC-GPe-STN pathway, generates adaptive exploratory actions gated by a vigilance signal modulated by the accumulated global temporal-difference (TD) error; and the Critic, based on the ventral striatum, computes the TD error. The final action is selected by a separate, TAN-based Arbitrator, which probabilistically chooses between the Actor's and Explorer's action proposals according to recent performance and the TD error. These subnetworks are coupled through a unified three-factor learning framework that uses the TD signal and phasic neuromodulators (acetylcholine and dopamine) from the Arbitrator to drive pathway-specific synaptic plasticity. This online plasticity enhances the quality of action proposals and accelerates policy refinement. In simulation, SpikeAEC outperforms leading brain-inspired methods by converging 24% faster, reducing trajectory length by 18%, and increasing cumulative reward by over 5% against the top-performing baseline, all while maintaining consistency with established neurophysiological principles.
DOI: 10.3389/fnbot.2026.17577952026-03-03
Bangcheng Zhang, Qi Xia
As automotive manufacturing advances toward the industrial 5.0 era, traditional rigid automation production models are transitioning toward the embodied intelligence paradigm. Confronted with mass customization, diverse products, and small-batch production, the environment of automotive manufacturing exhibits high dynamism and unstructured characteristics. Different from traditional industrial intelligence based on static, hard-coded logic, robots enhance their cognitive abilities through closed-loop interaction with dynamic environments, inspired by bionic neural mechanisms, this shift enables robots to perform flexible and reliable operations in complex production scenarios. This paper analyzes the core role and key technologies of neural intelligence algorithms in reshaping perception, decision, and execution of industrial robot, while providing a systematic review of industrial robot evolution within the automotive industry, and provides a reliable path for future development.
DOI: 10.3389/fnbot.2026.17960432026-02-17
Shanqin Wang, Mengjun Miao, Miao Zhang
ProblemDeep learning technology promotes the development of single-image dehazing. However, many existing methods fail to fully consider the haze density and its spatial distribution, which limits the improvement of dehazing performance.Proposed solutionTo address this issue, we propose an attention-based multi-scale feature aggregation network (AMSA-Net) for single-image dehazing.MethodAMSA-Net is an encoding and decoding structure. Its encoder and decoder are composed of multi-scale hybrid attention feature aggregation module (MSHA-FAM). The module can perceive the haze density and spatial information in the haze image, which helps to improve the dehazing effect. MSHA-FAM is composed of two key components: the scale-aware coordinate residual module (SCRM) and multi-scale feature refinement residual module (MSFRRM). SCRM uses improved coordinate attention to effectively capture haze density and spatial characteristics, thus significantly improving dehazing effect. MSFRRM extracts semantic features through up-sampling and down-sampling, and uses improved pixel attention mechanism to enhance key features. In the overall MSHA-FAM pipeline, SCRM first learns the density and spatial distribution characteristics of haze, then refines it through MSFRRM, so as to remove haze more effectively.Key resultsThe experimental results demonstrate that our proposed AMSA-Net is superior to the comparison methods in terms of dehazing quality. Ablation studies further verify the effectiveness of the proposed modules.ImpactIn this work, we present AMSA-Net, which has achieved good dehazing performance and can provide high-quality input for subsequent computer vision tasks.
DOI: 10.3389/fnbot.2026.16981002026-02-06
Xiaoxiao Cao
IntroductionThe integration of virtual simulation with intelligent modeling is crucial for advancing the scientization and personalization of volleyball physical training. This study aims to overcome the convergence instability and feature misalignment in modeling multimodal kinematic and physiological sequences.MethodsA dynamical framework based on a Dual-Stream Long Short-Term Memory network integrated with a temporal attention mechanism is proposed. The framework decouples heterogeneous feature learning and optimizes temporal weight distribution.ResultsExperimental validation on complex motion state estimation demonstrates that the proposed model reduces load modeling error to 3.8% and achieves a motion classification accuracy of 93.1%. The velocity trajectory fitting coefficient of determination is 0.91 with a peak deviation of 0.05 m/s.DiscussionThese results confirm the effectiveness of the attention-based DS-LSTM in optimizing multimodal sequence modeling for training state estimation and feedback.
DOI: 10.3389/fnbot.2026.17604942026-02-02
Heba G. Mohamed, Muhammad Nasir Khan, Fawad Naseer, Muhammad Tahir, Mohsin Jamil
IntroductionTelepresence robots (TPRs) must co-navigate with humans in constrained hospital environments, where safety depends on anticipating rather than merely reacting to human motion. Existing approaches rarely integrate short-horizon human-motion forecasting with safety-constrained control, which reduces robustness in dense corridors and ward bays. This study addresses this gap by evaluating an anticipatory, safety-aware co-navigation framework for TPRs.MethodsWe developed a modular framework that couples a lightweight transformer-based forecaster that predicts multi-agent trajectories under occlusion with a safe reinforcement learning (RL) controller. The forecaster produces short-term distributions over pedestrian states that are embedded into the RL policy state and cost as risk-aware occupancy features. Safety is enforced via constrained policy optimization augmented by a run-time control barrier function (CBF) shield that filters unsafe actions. We benchmarked the approach against a social-force or dynamic window approach (DWA), an attention-based crowd-RL policy, and model predictive control (MPC) with CBF. Experiments were conducted across two hospital-like benchmarks (a crowded corridor and a four-bed ward), totaling 2,400 episodes. Outcomes included task success, collision count, minimum human–robot clearance, near-miss events ( ≤ 0.3 m), time-to-goal, CBF violations, and ablations removing forecasting and the CBF shield.ResultsRelative to the best-performing baseline, the proposed method improved task success by 21.6% and reduced collisions by 47.3%. Median minimum human–robot clearance increased by 0.19 m, and near-miss events decreased by 38.5%. Time-to-goal was maintained within +2.7% of MPC+CBF while incurring zero CBF violations under the shield. Ablation studies showed that removing forecasting degraded success by 14.2%, whereas removing the CBF shield increased constraint breaches from 0% to 6.1% of steps.DiscussionAnticipatory perception combined with Safe-RL yields substantially safer and more reliable telepresence co-navigation in human-dense clinical layouts without sacrificing efficiency. The framework is modular, enabling alternative forecasters and safety shields. Limitations include sensitivity to forecast drift during abrupt changes in crowd flow. Future work will explore on-device adaptation, shared-autonomy overlays to incorporate operator intent, and prospective evaluations in live hospital workflows.
DOI: 10.3389/fnbot.2025.16975182026-01-05
Qian Han, Jinfu Lao, Jinyong Zhang
To address these challenges, we propose a subdomain adaptation framework driven by transferable semantic alignment and class correlation. First, source and target domains are divided into subdomains according to class labels, and a joint subdomain distribution alignment mechanism is introduced to reduce intra-class distribution divergence while enlarging inter-class disparities. Second, a domain-adaptive semantic consistency loss is employed to cluster semantically similar samples and separate dissimilar ones in a unified representation space, enabling precise cross-domain semantic alignment. Third, pseudo-label quality in the target domain is improved via temperature-based label smoothing, complemented by a class correlation matrix and a loss function capturing inter-class relationships to exploit intrinsic intra-class coherence and inter-class distinction. Extensive experiments on multiple public datasets demonstrate that the proposed method achieves superior average classification accuracy compared to existing approaches, validating the effectiveness of semantic alignment and class correlation modeling. By explicitly modeling intra-class coherence and inter-class distinction without additional architectural complexity, the framework effectively mitigates domain shift, enhances semantic alignment, and improves recognition performance on the target domain, offering a robust solution for deep unsupervised domain adaptation.
DOI: 10.3389/fnbot.2025.1665528