2025-12-10
Badiu Badams, Usman Ullah Sheikh, Syed Abd Rahman Syed Abu Bakar, Norhaliza Abdul Wahab
Detecting small and partially hidden objects in rivers and water bodies remains a major challenge for real-time waste detection systems. These objects are often missed due to their small size, low contrast, and cluttered surroundings. Further complicating the task is the lack of dedicated datasets focused on small floating debris, limiting the development of more capable detection models. To bridge this gap, we developed D_six, a custom dataset of 495 high-resolution images capturing six classes of floating waste under real-world conditions. In this study, we improve the YOLOv5s object detection model by integrating atrous convolutions at three key backbone layers: P1/2, P3/8, and P5/32. These layers represent different scales of the feature pyramid, and the strategic placement of atrous convolution at each level plays a crucial role in helping the model recognize small and occluded objects more effectively. Using a dilation rate of 6, the model’s receptive field is expanded without increasing its size or slowing it down. When trained and evaluated on the D_six data set, the FloYO-Net (Floating Object YOLO Network) consistently outperformed the standard YOLOv5s, achieving a mean Average Precision (mAP@0.5) of 0.828 and mAP@0.5:0.95 of 0.509, compared to 0.787 and 0.498 respectively. Improvements were especially notable for hard-to-detect items like plastic bottles and plastic drink containers, with average precision gains of 6.6% and 7.1%, respectively. These results demonstrate that atrous convolution — when thoughtfully placed — can significantly improve detection accuracy, making it a powerful enhancement for real-time environmental cleanup systems.
2025-12-10
Sherin Babu, Binu Thomas
Particulate Matter induced air pollution is known to have significant negative impacts on both the environment and human health. This research evaluates the effectiveness of various decision tree ensemble models in predicting daily PM10 concentrations in Thiruvananthapuram, Kerala, from July 2017 to December 2019. Seven decision tree ensemble models, namely Random Forest, Extra Trees, Gradient Boosting, AdaBoost, LightGBM, XGBoost, and Histogram-Based Gradient Boosting are employed here. To address missing data in the dataset, kNN imputation is utilized for a cohesive dataset suitable for model training. The models utilize both meteorological and air pollutant variables, with performance assessment using metrics such as the coefficient of determination (R²), root mean square error (RMSE) and mean absolute error (MAE). The findings indicate that the Extra Trees regression model provided the best prediction performance (R² = 0.9397, RMSE = 6.664 μg/m³, MAE = 4.950 μg/m³). Histogram-Based Gradient Boosting and Random Forest also demonstrate strong predictive capabilities. The explainability of the best prediction models is conducted by the feature importance analysis process. Feature importance analysis highlighted sulfur dioxide (SO2) as the most significant pollutant influencing PM10 levels, alongside meteorological factors like wind speed and rainfall, enhancing both prediction accuracy and interpretability of results. This research represents the first comprehensive effort to predict PM10 levels in Thiruvananthapuram using machine learning techniques, addressing a gap in regional air quality studies.
2025-12-10
Fatima Aliyu Shugaba, Usman Ullah Sheikh, Mohd Afzan Othman, Nurulaqilla Khamis, Muhammad Habibullah Abdulfattah
Handwritten Arabic text recognition (HATR) presents unique challenges due to complex character shapes, contextual variations, cursive connections, and the presence of diacritical marks. This study introduces AHAD (Arabic Handwritten Alphabet with Diacritics), a novel benchmark dataset of 71,061 handwritten Arabic character images annotated with five primary vowel diacritics; Fathah, Kasrah, Dammah, Shaddah, and Sukoon, covering 492 distinct classes that combine character identity, contextual form, and diacritic. Leveraging this dataset, we propose an incremental learning framework based on Convolutional Neural Networks (CNNs) to address fine-grained recognition of handwritten Arabic characters with its corresponding diacritics. The model was initially trained on a 114-class dataset of handwritten Arabic characters (in all contextual forms) of non-diacritic characters and fine-tuned in two phases using the AHAD dataset. The two-phase strategy includes output layer expansion, learning rate adjustment, and gradual unfreezing of deeper layers to enhance knowledge retention and prevent catastrophic forgetting. The proposed method achieved a validation accuracy of 92.96% and a test accuracy of 93.26%. Our findings demonstrate the effectiveness of incremental learning for diacritic-aware Arabic handwriting recognition and establish AHAD as a strong baseline for future research in this field.