2026-02-02
Shubham Singh, Arun Kaushik
This manuscript focuses on reliability acceptance sampling plans applied to the Burr type-XII distribution within the framework of the generalized hybrid censoring scheme. This study involves determining the sample size ( n ) and acceptance constant ( k 1 ) for given values of producer's and consumer's risks, utilizing the asymptotic normality theory of maximum likelihood estimators for model parameters. Additionally, a Monte-Carlo simulation study is conducted to assess whether the reliability acceptance sampling plans adhere to specified risk levels for finite sample sizes. Optimum reliability acceptance sampling plans are obtained by employing a variance minimization criterion under cost constraints. We have provided the algorithm for the computation of sample size ( n ) and acceptance constant ( k 1 ). Furthermore, the algorithms for computing the optimum reliability acceptance sampling plans under generalized hybrid type-I censoring scheme and generalized hybrid type-II censoring scheme are also provided.
2025-10-30
A. M. Mathai, M. Kumar
Mixture distributions are widely utilized in various practical problems, such as clinical experiments and electronic component life testing. Despite this, the literature does not extensively cover acceptance sampling plans associated with these distributions. In this paper, variable acceptance sampling plans are designed for a mixture of exponential-Rayleigh distributions using partially accelerated life tests. Under progressive Type-II censoring scheme, the maximum likelihood estimators of the unknown parameters of the mixture distribution are derived for Arrhenius and linear life-stress relationships. Based on these relationships, optimal variable sampling plans are formulated. The plan parameters are determined by solving corresponding optimization problems. The study presents numerical findings, a comparative analysis, and sensitivity assessments. Finally, the practical applicability and relevance of the proposed acceptance sampling plans are demonstrated using real-world datasets. These datasets include breast cancer patients' records and failure lifetime data from communication transmitter-receivers in a commercial aircraft.
2025-05-28
Artur Lemonte
This paper proposes a new control chart based on the two-parameter standard two-sided power distribution for monitoring rates and proportions, that is, when the quality characteristic of interest belongs to the unit interval (0,1). Control charts based on the well-known beta and Kumaraswamy distributions are usually considered to deal with this kind of data. The standard two-sided power distribution has many similarities to the beta and Kumaraswamy distributions and a number of advantages in terms of tractability. We evaluate and compare the performance of the new control chart with the beta and Kumaraswamy control charts through Monte Carlo simulation experiments. The simulation results reveal that the control chart based on the standard two-sided power distribution outperforms the beta and Kumaraswamy control charts in terms of run length analysis. An empirical application to a real data set is considered to illustrate the new control chart in practice, and comparisons with the two most traditional control charts for rates and proportions (beta and Kumaraswamy) are made.
2025-05-28
Bernhard Meindl
National statistical offices (NSIs) routinely publish aggregated data in the form of statistical tables. However, ensuring data privacy is a critical aspect of this process. Anonymization techniques must be applied to these tables to safeguard the privacy of individual data contributors and prevent unauthorized inference about specific units from the published outputs as often required by law. The R package cellKey offers a possible solution to this challenge by implementing a post-tabular perturbation method. This method modifies table cell values after aggregation, ensuring that sensitive information is adequately masked. It is versatile, suitable for both frequency tables and magnitude tables. A key feature of the cellKey package is its ability to maintain consistency across multiple tables that share identical cells. This ensures that anonymized data across different tables remains coherent while still protecting privacy. This approach makes the package especially useful for scenarios involving complex datasets with interrelated tables. The cellKey package is user-friendly and can empower NSIs and other data holders to publish statistical outputs that uphold both data utility and privacy, meeting the growing demands for secure and accessible data dissemination.
2025-04-23
Gregor Laaha, Johannes Laimighofer, Nur Banu Özcelik, Svenja Fischer
Environmental models typically rely on stationarity assumptions. However, environmental systems are complex, and processes change over states or seasons, leading to often overlooked heterogeneity. This paper explores methods to incorporate process heterogeneity into statistical models to improve their performance. It considers problems from natural hazards and earth system sciences, demonstrating the effects of process heterogeneity and proposing methodological advances through model extensions. The first problem addresses flood frequency analysis, where floods are generated by different processes in catchment and atmosphere. A mixture model combining peak-over-threshold distributions of flood types can handle this heterogeneity, especially regarding tail heaviness, making it relevant for flood design. The second problem involves minimum flow frequency analysis, with heterogeneity from different summer and winter processes. A mixture distribution model for minima and a copula-based estimator can incorporate seasonal distributions and event dependence, showing significant performance gains for extreme events. The third problem examines process heterogeneity in rainfall models. Clustering event characteristics (e.g., duration, intensity) using Gower's distance and a lightning index helps distinguish between convective and stratiform events, showing potential to enhance rainfall generators. The fourth problem deals with parameter variation in temporal models of environmental variables, using daily streamflow series. A tree-based machine learning model shows that prediction performance and model parameters vary with quantile loss optimization, suggesting the need for different or combined models for full time series in the presence of process heterogeneity. The study highlights the importance of considering process heterogeneity in modeling from the outset and encourages a better understanding of statistical assumptions and the enrichment of physical knowledge in environmental statistics.
2025-04-23
Bernhard Spangl
The problem of recursive filtering in linear state-space models is considered. The solution to this problem is the classical Kalman filter which is optimal in the sense that it minimizes the variance of the estimated states, if the error processes of the state and observation equations are both Gaussian. However, the Kalman filter is well known to be sensitive to outliers, so robustness is an issue. Two approximate conditional-mean (ACM) type filters for vector-valued observations are proposed that generalize existing univariate filters of similar type to the multivariate case. These new ACM-type filters are compared by simulations in a multivariate setting with additive outliers to the classical Kalman filter and the robust least squares (rLS) filter, another approach robustifying the Kalman filter. Additionally, different settings of tuning parameters and their impact are investigated. The results of the simulation experiments show that in the presence of additive outliers the multivariate ACM-type filters not only outperform the classical Kalman filter, as expected, but they also outperform the rLS filter.
2025-04-23
Torsten Hothorn
Reproducibility of statistical simulations is crucial but proved being a challenge in its own right. Recently, lack of reproducibility of important simulation studies stimulated developments of reporting guidelines and specific protocols trying to improve on this situation. The problem, of course, is not new and issues regarding reproducibility of numerical results, for example statistical analyses or simulations, have long been known. Documented lack of progress regarding reproducibility in the past decade naturally leads to the question if problem awareness was not powerful enough to lead to improved reproducibility. As a benchmark case, I tried to reproduce a simulation study in a so far unpublished manuscript by Fritz Leisch and myself. The results show that, time and again, the devil is in the details and much self-discipline and extensive record-keeping and documentation are mission critical.