Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 01:26:24 EST

0

Posted in stat.AP · 2026-09-03 · Mayleen Cortez-Rodriguez

Natural Disasters and the Nonprofit Sector

When natural disasters strike, individuals, communities, and even entire countries can suffer. Researchers have studied the impacts of disasters on various factors of interest, from mental health, to poverty, to economic activity. However, the impact of disasters on the nonprofit sector is understudied despite the nonprofit sector's...

💬 0 commentsarXiv:2609.04136v1PDF
0

Posted in stat.CO · 2026-09-03 · Shuyang Cao, Alex Stringer

Fast Computation of Nested Cross-Validation for Penalized Regression

Cross-validation is a resampling procedure that provides a point estimate of generalization error for any predictive model. Cross-validation is widely used for model selection and evaluation. Uncertainty in the cross-validation estimate is challenging to quantify, and estimation of its variance is known to require multiple runs of the...

💬 0 commentsarXiv:2609.04126v1PDF
0

Posted in stat.ME · 2026-09-03 · María Eugenia Riaño

Model-assisted estimation with a training subsample: a two-phase sampling approach with design-based variance estimation

When a flexible prediction model is fitted on a training subsample drawn from a probability sample, the model-assisted estimator actually reported arises from one realized partition, yet existing theory quantifies uncertainty only for partition-averaged, cross-fitted, or symmetrized versions of it. We represent the training subsample...

💬 0 commentsarXiv:2609.04082v1PDF
0

Posted in stat.ME · 2026-09-03 · Emanuele Giorgi, Claudio Fronterre, Peter Diggle

Comment on: "The Two Cultures of Prevalence Mapping: Small Area Estimation and Model-Based Geostatistics"

Small Area Estimation (SAE) and Model-Based Geostatistics (MBG) provide complementary approaches to prevalence mapping, with their relative advantages depending on the inferential goals and characteristics of the available data. We argue that a fuller comparison should consider model interpretability, the role of epidemiologically...

💬 0 commentsarXiv:2609.03805v1PDF
0

Posted in stat.ME · 2026-09-03 · Benjamin Poignard, Yoann Potiron

Parametric estimation of Hawkes processes based on ordinary least squares

We develop a parametric estimation framework for self-exciting Hawkes processes whose intensity functions admit a parametric form. The estimation procedure is based on ordinary least squares. To apply the least squares estimation, we restrict to a kernel class that can be expressed as a sum of the product of a parameter and a...

💬 0 commentsarXiv:2609.03696v1PDF
0

Posted in stat.ME · 2026-09-03 · Žikica Lukić, Bojana Milošević

Change-point analysis: a new perspective for unstable financial markets

We introduce two new classes of nonparametric change-point tests for sequences of univariate non-negative random variables. The proposed procedures are based on the empirical modified Hankel transform and the Laplace transform, respectively, and provide new transform-based tools for detecting distributional changes. We derive the...

💬 0 commentsarXiv:2609.03614v1PDF
0

Posted in stat.ME · 2026-09-03 · Giulia Patanè, Sonja Greven, Alessandra Menafoglio

Random mixtures in Bayes Hilbert spaces

We present a framework for the analysis and unmixing of random density mixtures in the Bayes Hilbert space. General identifiability results for mixtures in Hilbert spaces are established and applied to the Bayes Hilbert space setting. Building on these results, we propose a penalised maximum likelihood approach for the unmixing of...

💬 0 commentsarXiv:2609.03523v1PDF
0

Posted in stat.ML · 2026-09-03 · Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang

Towards a Statistical Understanding of Mixture-of-Experts

Mixture-of-experts (MoE) architectures increase model capacity by combining a collection of expert predictors through input-dependent routing, while often activating only a small subset of experts for each input. Despite their growing importance in modern large-scale models, the statistical roles of their design choices, especially...

💬 0 commentsarXiv:2609.03501v1PDF
0

Posted in stat.ML · 2026-09-03 · Quang Hoang Trung, Quang Huu Hieu, Nguyen Van Hoang Phuc, Vo Nguyen Le Duy

ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models

Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among likely token alternatives. Existing local approaches often select candidate tokens from either the teacher or...

💬 0 commentsarXiv:2609.03355v1PDF
0

Posted in stat.ME · 2026-09-03 · Bob Wilson

Randomization Inference for Matched Pairs with Binary Outcomes

We give an exact randomization-based confidence set for the average treatment effect (ATE) in matched-pair studies with a binary outcome, requiring neither monotonicity nor any distributional assumption beyond the within-pair coin flip. At its core is an analytic solution to the worst-case allocation of attributable effects: two...

💬 0 commentsarXiv:2609.03227v1PDF
0

Posted in stat.AP · 2026-09-02 · Yulin Guo, Veera Sundararaghavan, Boris Kramer

Uncertainty quantification of fatigue initiation life for powder bed fusion metal additive manufacturing

Predicting fatigue life with quantified uncertainties is essential for the qualification of critical components produced by laser-based powder bed fusion additive manufacturing. We present a framework that propagates microstructure and defect uncertainties directly to a fatigue initiation life distribution for a specific part. In...

💬 0 commentsarXiv:2609.03163v1PDF
0

Posted in stat.ML · 2026-09-02 · Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios

A Closed-Form Formula for Consistent Lipschitz Regression on Metric Spaces with Sparse Neural Network Realizations

Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks. We address this by introducing a...

💬 0 commentsarXiv:2609.03129v1PDF
0

Posted in stat.ME · 2026-09-02 · Kun Xia, Jianrui Zhang, Qing Lu, Chenxi Li

Multimarker genetic association tests for panel count data

The existing multimarker survival tests focus on time to event outcomes. However, recurrent events are common in real world clinical and biomedical studies, especially in the research of chronic and recurrent diseases. In this paper, we develop a suite of set based genetic association tests for panel count outcomes under a unified...

💬 0 commentsarXiv:2609.03113v1PDF
0

Posted in stat.ML · 2026-09-02 · Zihao Shi, Huajun Xi, Bingyi Jing, Hongxin Wei

Occupancy-based Quantile Risk Control

Conformal risk control is an emerging framework for the safe deployment of machine learning models with finite-sample guarantees. To accommodate a broader class of risk notions, quantile risk control extends this framework to quantile-based risk measures. However, existing methods either suffer from excessive conservatism or lack...

💬 0 commentsarXiv:2609.03104v1PDF
0

Posted in stat.ME · 2026-09-03 · Xiaorui Wang, Juan-Juan Cai, Huixia Judy Wang, Jian Qing Shi, Yanlin Tang

Causal Inference for Heterogeneous Extreme Quantiles with Heavy-Tailed Outcomes

We propose a framework for estimating conditional extreme quantile treatment effects (CEQTEs) in observational studies with heavy-tailed outcomes. Our procedure first estimates intermediate conditional quantiles using inverse-probability-weighted (IPW) quantile regression and then extrapolates them to extreme levels using extreme...

💬 0 commentsarXiv:2609.03933v1PDF
0

Posted in stat.ME · 2026-09-02 · Marc Delord

Non-Invariance in Nested Prediction Models under Selective Predictor Availability

Selective measurement of predictors is common in routinely collected health data. We used nested prediction models as a framework for characterising the consequences of a selectively measured predictor, with a restricted model defined in the target population and an extended model including the selectively measured predictor defined...

💬 0 commentsarXiv:2609.02836v1PDF
0

Posted in stat.ML · 2026-09-02 · Zhaoming Li, Paul Hand

Full-Model Optimality for Tunable Linear Generative Priors in Compressed Sensing

Generative models have been studied experimentally and theoretically as priors for inverse problems such as compressed sensing. Recent work by Gunn et al. studied the use of generative priors with tunable complexity, where a family of generative priors with varying complexity is maintained and a specific complexity can be selected at...

💬 0 commentsarXiv:2609.02790v1PDF
0

Posted in stat.ME · 2026-09-02 · Jiwon Kang, Yun Am Seo

Quantum mutual information statistics for detecting dependence-structure change points in time series

Detecting when the dependence between two components of a multivariate time series changes, while the marginals drift freely, requires a dependence-specific statistic. We take the inferential object to be a density operator -- the trace-normalised second moment of unit-norm random Fourier features of ranks -- rather than a probability...

💬 0 commentsarXiv:2609.02787v1PDF
0

Posted in stat.ME · 2026-09-02 · Sahil Loomba, Dean Eckles

Off-policy causal estimation in networks

In the presence of interference, where the treatment assigned to one unit can affect the outcomes of others, many causal estimands depend on the treatment-assignment policy under which the experiment is conducted. This policy dependence creates a fundamental challenge for off-policy estimation, where the goal is to estimate causal...

💬 0 commentsarXiv:2609.02756v1PDF
0

Posted in stat.ML · 2026-09-02 · Jia-Nan Wang, Zixun Huang, Kairui Li, Lei Wu

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain $η_{\mathrm{SGD}}^{\mathrm{crit}}\eqsim 1$,...

💬 0 commentsarXiv:2609.02728v1PDF
0

Posted in stat.ME · 2026-09-02 · Johannes Brachem, Thomas Kneib

Reconciling Interpretability with Covariate-Dependent Shape Flexibility in Penalized Transformation Models for Distributional Regression

A central challenge in distributional regression is to allow the shape of the conditional distribution of the response variable to vary flexibly with covariates while retaining directly interpretable effects on its mean and standard deviation. We extend the penalized transformation model (PTM) family into a conditional-shape PTM,...

💬 0 commentsarXiv:2609.02662v1PDF
0

Posted in stat.ME · 2026-09-02 · Patrick B. Langthaler, Jun Ma, Jonas Beck

Detecting Early and Late Divergences in Survival Curves Using Nonparametric Effect Measures

Clinical trials often show treatment curves that diverge early and converge later, or vice versa patterns that are poorly captured by the proportional-hazards assumption. We develop a joint inferential framework for two nonparametric functionals of censored survival data: the Kaplan--Meier-based Mann--Whitney effect and a novel...

💬 0 commentsarXiv:2609.02596v1PDF
0

Posted in stat.CO · 2026-09-02 · Glory Mary Givi, Cédric Travelletti, Grégory Mermoud

TrunX: A massively parallel, differentiable implementation of the 3-PG forest growth model in JAX

Process-based forest models are widely used to simulate forest growth and responses to environmental change, but their calibration and application often require many computationally expensive model evaluations. We present an implementation of the Physiological Processes Predicting Growth (3-PG) model in JAX that uses just-in-time...

💬 0 commentsarXiv:2609.02557v1PDF
0

Posted in stat.AP · 2026-09-02 · Lee Suddaby, Gordon J Ross

Did Mary Shelley Write Frankenstein? A Stylometric Analysis

The novel Frankenstein was published anonymously in 1818, and was first credited to Mary Shelley in a French translation of 1821. Since its publication, several claims - both contemporaneous and recent - have been made suggesting that Frankenstein was actually written by Mary's husband, Percy Bysshe Shelley. We review the background...

💬 0 commentsarXiv:2609.02527v1PDF
0

Posted in stat.AP · 2026-09-02 · James Bailie

Big data, differential privacy, and national statistical organisations

Differential privacy (DP) has emerged in the computer science literature as a measure of the impact on an individual's privacy resulting from the publication of a statistical output such as a frequency table. This paper provides an introduction to DP for official statisticians and discuss its relevance, benefits, and challenges from a...

💬 0 commentsarXiv:2609.02495v1PDF