Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 04:19:01 EST

0

Posted in stat.OT · 2026-08-27 · Jonas Bjermo, Frank Miller

Algorithms for optimizing model-based incomplete block designs

Because of time limitations or participation burden, the treatments in an experimental design can be too large for a single subject. Instead of addressing this using combinatorial incomplete block designs, we propose a model-based approach that optimizes model parameters. This offers distinct advantages: it incorporates...

💬 0 commentsarXiv:2608.27056v1PDF
0

Posted in stat.ML · 2026-08-27 · Abdullah Karasan

Representation Measurements Under Function-Preserving Reparameterizations

Hidden coordinates are not uniquely determined by a language model's input--output function, so representation-derived measurements should be invariant to function-preserving changes of basis. This study shows that column-permutation parallel analysis violates function-preserving reparameterization invariance because its reference...

💬 0 commentsarXiv:2608.27020v1PDF
0

Posted in stat.ML · 2026-08-27 · Toni Karvonen, Chris J. Oates

Why not to use the Gaussian kernel

Kernels measure similarity or correlation in tasks such as regression and classification. The Gaussian kernel, other names of which include squared exponential and radial basis function kernel, is one of the most popular in Gaussian process regression. We argue that the Gaussian kernel is best avoided and should never be used as a...

💬 0 commentsarXiv:2608.26974v1PDF
0

Posted in stat.ME · 2026-08-27 · Xiaoxiao Ling, Andrea Gabrio, Gianluca Baio

A Bayesian Longitudinal Model for Imputing Item-Level Missing Data in Trial-Based Economic Evaluations

Trial-based economic evaluations are widely used to assess the cost-effectiveness of healthcare interventions and inform decision-making. Cost and effectiveness outcomes are typically collected using multi-item questionnaires administered at multiple time points, and are often subject to item-level missingness. In principle,...

💬 0 commentsarXiv:2608.26929v1PDF
0

Posted in stat.AP · 2026-08-27 · Gurjeet Sangra Singh, Frantzeska Lavda, Alexandros Kalousis

Climate Physics Dynamic Matching

Deep generative models such as flow matching and diffusion models have shown potential for learning complex dynamical systems, but typically act as black boxes that neglect underlying physical structure, while physics-based models governed by partial differential equations are often incomplete due to missing source terms, or uncertain...

💬 0 commentsarXiv:2608.26907v1PDF
0

Posted in stat.ME · 2026-08-27 · Jianming Wu, Xinyu Zhang, Jie Zeng

Expected Shortfall Model Averaging

Expected shortfall (ES) is widely used to measure tail risk in finance and economics, but its prediction is challenging due to non-elicitability and model uncertainty. This paper proposes a two-stage cross-validation model averaging method for ES forecasting. In the first stage, conditional value-at-risk is estimated using quantile...

💬 0 commentsarXiv:2608.26805v1PDF
0

Posted in stat.ML · 2026-08-27 · Athanasios Vlontzos, David Gustafsson, Michael O'Riordan, Ciarán M. Gilligan-Lee

Incremental Recommendation via Causal Models

Recommendation impressions are a finite resource, hence delivering a recommendation to a user who would discover the content organically yields no incremental value and displaces other recommendations that could. We address this by extending an existing production recommendation model to a causal architecture using holdback data that...

💬 0 commentsarXiv:2608.26804v1PDF
0

Posted in stat.AP · 2026-08-27 · Matthias von Davier

Integrating Network Psychometrics and LLMs: The Ising-Embeddings-Model applied to Reliability Auditing

Scoring consistency for constructed-response items in large-scale assessments is typically estimated through double-scoring, which uses small samples and assumes independence among responses. We present an integrated framework combining network psychometrics with the Linguistic-Integrated Reliability Audit (LiRA) via a modified Ising...

💬 0 commentsarXiv:2608.26790v1PDF
0

Posted in stat.ME · 2026-08-27 · Georgios Gavrilopoulos, Johanna Ziegel

Uncertainty quantification for expectation-calibrated predictions

The existing literature on model calibration focuses mainly on classification and probabilistic prediction. In this work, we address calibrated point predictions for the conditional mean. Although existing impossibility results preclude exact out-of-sample calibrated predictions, we develop calibrated confidence intervals that provide...

💬 0 commentsarXiv:2608.26703v1PDF
0

Posted in stat.ML · 2026-08-27 · Yanhang Zhang, Wei Liu, Yuhong Yang

A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs

Model selection becomes particularly challenging under strong predictor dependence and model-class uncertainty, especially when there are exponentially many models. We propose a Descriptive-Complexity Information Criterion (DCIC) that regularizes large candidate model collections through Kraft-admissible code lengths. Under...

💬 0 commentsarXiv:2608.26618v1PDF
0

Posted in stat.ME · 2026-08-27 · Yen-Chi Chen

On efficiency gains via augmenting a tiny sample with a massive auxiliary sample

In this paper, we study the problem of augmenting a tiny target sample with a massive auxiliary sample. Utilizing Tukey's factorization, there are two popular approaches: the inverse probability weight (IPW) and the full-likelihood (FL) methods. We show that the IPW approach suffers from the limited target sample problem while the FL...

💬 0 commentsarXiv:2608.26610v1PDF
0

Posted in stat.ME · 2026-08-27 · Shiyao Liu, Junni L. Zhang

Analyzing Within-Subject Experiments: Identification, Testing, and Sensitivity

Recent work encourages political scientists to move from post-only toward within-subject designs for improved precision from repeated measurements. We formalize a potential-outcomes framework for two-period within-subject designs that allows for unequal allocation and heterogeneous treatment and carryover effects. We characterize the...

💬 0 commentsarXiv:2608.26606v1PDF
0

Posted in stat.AP · 2026-08-27 · Stephen Jun Villejo, Peter Diggle, Guangquan Li, Ella White, Matthew Wade, Christopher Williams, Davey L. Jones, Alisha Davies, Marta Blangiardo

A spatio-temporal block aggregation model for latent log Gaussian outcomes: application on modelling wastewater virus concentration in Wales

Wastewater-based epidemiology has emerged as a valuable tool for monitoring community-level infectious disease dynamics, providing population-wide signals that complement clinical surveillance. However, wastewater measurements are often observed as aggregated values over irregular spatial units. This work develops an approach to link...

💬 0 commentsarXiv:2608.27207v1PDF
0

Posted in stat.ME · 2026-08-27 · Kejun Chen, Xianqi Wei, Qianqian Zhu

Personalized Federated Learning for Tensor Regression

The growing availability of tensor-valued data across multiple institutions creates opportunities for collaborative analysis, but also raises challenges related to data privacy, high dimensionality, and client heterogeneity. This paper introduces a personalized federated tensor regression framework that addresses all three...

💬 0 commentsarXiv:2608.27191v1PDF
0

Posted in stat.ME · 2026-08-26 · Robin Denz, Filippo Saatkamp, Katharina Meiszl, Nina Timmesfeld

The Symmetric Pair Matching Design: A Self-Controlled Method with Automatic Adjustment for Time Effects

Self-controlled study designs eliminate confounding by individual-level characteristics that remain constant during the observation time and are thus widely used in pharmacoepidemiology and vaccine safety research. However, existing methods remain vulnerable to time effects, including temporal trends and seasonality in the exposure or...

💬 0 commentsarXiv:2608.25979v1PDF
0

Posted in stat.ME · 2026-08-26 · Amitakshar Biswas, Adam B Kashlak

Random Invariance Testing on Quadratic Form Statistics with Application to Autocorrelation

Randomization testing with permutations is a very common nonparametric approach to hypothesis testing. However, randomization testing can be done with other group transformations including random rotations. In this work, we consider the problem of invariance in quadratic form statistics under a unified framework with closed form...

💬 0 commentsarXiv:2608.25918v1PDF
0

Posted in stat.ML · 2026-08-26 · David P. Hofmeyr

Efficient Estimation of High Information Projections using Nearest Neighbours

An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data. Following similar intuitive motivation to a number of existing techniques, the proposed method is based on enhancing the nearest neighbour relationships in the data. The proposed projection...

💬 0 commentsarXiv:2608.25887v1PDF
0

Posted in stat.ML · 2026-08-26 · Mahamat Hamdan Nassouradine, Clément Gauchy, Pierre-Emmanuel Angeli, Sébastien da Veiga

Multi-output Gaussian process prediction of physical fields under linear equality constraints

We address the simultaneous prediction of multiple high-dimensional physical fields governed by linear equality constraints, a setting that arises in many real-world applications in physics machine learning. Gaussian process (GP) regression is a widely used surrogate modeling approach due to its effectiveness in small-sample regimes...

💬 0 commentsarXiv:2608.25709v1PDF
0

Posted in stat.AP · 2026-08-26 · Elkanah Nyabuto, Philipp Otto

Learning Volatility Dependence Networks in UK Equity Markets using Penalised Spatiotemporal ARCH Models

Spatiotemporal ARCH models capture temporal volatility persistence and cross-sectional dependence but typically require a predefined spatial weight matrix. This is restrictive in financial markets, where the dependence network is rarely known. We develop a LASSO-penalised quasi-maximum likelihood estimator that jointly learns a sparse...

💬 0 commentsarXiv:2608.25588v1PDF
0

Posted in stat.ML · 2026-08-26 · Caixing Wang, Zhibo Chen, Yue Wang

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

Random feature methods provide a scalable approximation to kernel ridge regression (KRR), but the regularization parameter that yields the oracle learning rate depends on unknown smoothness and capacity parameters. In this work, we propose a neighboring early-stopping rule for adaptive regularization in KRR with random features...

💬 0 commentsarXiv:2608.25513v1PDF
0

Posted in stat.ME · 2026-08-26 · Shunxing Yan, Fang Yao

Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality

Functional data analysis is an important statistical field that treats data as random functions. In practice, the random functions are often not fully observed but instead measured at discrete times. While simpler problems, such as mean and covariance estimation, have been widely studied for discretely observed data, optimal...

💬 0 commentsarXiv:2608.25468v1PDF
0

Posted in stat.ME · 2026-08-26 · Yusaku Ohkubo, Yukito Iba

{poscosea} : A Computationally Efficient Sensitivity Analysis for Bayesian Models using the posterior covariance representation

Bayesian methods are essential in modern data analysis in ecology and evolutionary biology. They provide a flexible framework for modeling complex data-generating processes, while quantifying uncertainty based on the classical subjective interpretation of probability. However, Bayesian inference may provide misleading measures of...

💬 0 commentsarXiv:2608.25426v1PDF
0

Posted in stat.ME · 2026-08-26 · Jilin Wu, Ruike Wu, Zhijie Xiao, Mengxi Zhang

Robust Nonparametric Testing for Structural Changes in Multivariate Volatility via Multiple Quantiles

We propose an omnibus nonparametric test for structural changes in the multivariate volatility matrix. The test aggregates bounded generalized quantile scores over a range of quantile levels and has a weighted leave-$q$-out $U$-statistic representation. Deleting nearby index pairs renders the centering effect induced by serial...

💬 0 commentsarXiv:2608.25310v1PDF
0

Posted in stat.ME · 2026-08-26 · Sokbae Lee, Yuan Liao, Myung Hwan Seo, Youngki Shin

SAUSS: Stochastic Approximation with Unbiased Simulated Scores for Limited Dependent Variable Models

Multinomial choice models allow flexible substitution patterns but become computationally demanding with many alternatives or observations. With a fixed per-observation simulation budget, simulated maximum likelihood introduces simulation bias, while each optimization step requires a full-sample likelihood evaluation. We propose...

💬 0 commentsarXiv:2608.25304v1PDF