Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 01:26:24 EST

0

Posted in stat.ME · 2026-08-31 · Jack Storror Carter

Parameterising Gaussian Graphical Models

Gaussian graphical models (GGMs) describe the dependence structure among jointly Gaussian random variables. However, the most common parameterisation of GGMs, the precision matrix, describes both the dependence and scale of the variables. This has been shown to lead to model selection methods that depend on the scale of the variables,...

💬 0 commentsarXiv:2609.00288v1PDF
0

Posted in stat.ML · 2026-08-31 · Mitch Hill

Exact Global MCMC with Denoising Diffusion

This work shows that diffusion models learned with standard denoising loss can provide effective global MCMC proposals for complex high-dimensional target densities. The method is motivated by the observation that sequentially applying a forward and reverse diffusion process defines a Markov chain with a target stationary distribution...

💬 0 commentsarXiv:2609.00279v1PDF
0

Posted in stat.ML · 2026-08-31 · Zihang Liang, Haochen Zhang, Lingzhou Xue

Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost

We study federated online reinforcement learning with linear function approximation. While recent multi-agent reinforcement learning algorithms achieve strong regret guarantees, they typically require sharing raw trajectories. This reliance incurs a communication cost that scales linearly with the number of episodes and violates the...

💬 0 commentsarXiv:2609.00193v1PDF
0

Posted in stat.AP · 2026-08-31 · Ben O'Brien, Lewis J. Lehe

A Tool for Reconstructing Transit Vehicle Trajectories: A Case Study at IndyGo

Automatic vehicle location (AVL) data produced by transit vehicles is invaluable in performance studies, but turning raw AVL points into a detailed view of vehicle stop-and-gos is burdensome: the datasets are sparse, noisy, and prone to blunders. While recent research has explored methods of reconstructing trajectories describing the...

💬 0 commentsarXiv:2608.31078v1PDF
0

Posted in stat.ME · 2026-08-31 · Anton van Beek, Adam. M Boyce, Will J. Dawson, James B. Robinson

Posterior Geometry and Identifiability in Multi-Response Bayesian Calibration

Calibration under model misspecification is inherently ill-posed because calibration parameters and structural discrepancy are statistically confounded without additional assumptions. Bayesian formulations address this ambiguity through prior and covariance modeling choices, including multi-response observations and cross-source...

💬 0 commentsarXiv:2608.31047v1PDF
0

Posted in stat.ML · 2026-08-31 · James Crowley, Faez Ahmed, Anton van Beek

Learning the Geometry of Admissible Hypotheses through Inductive Bias in Training Distributions

Scientific discovery often requires reasoning over competing hypotheses that are consistent with experimental observations. For mixed-variable and combinatorial hypothesis spaces, however, constructing probabilistic representations remains challenging because both the active model components and their associated parameters are...

💬 0 commentsarXiv:2608.31028v1PDF
0

Posted in stat.CO · 2026-08-31 · Rahul Singh, Abhinek Shukla

Scalable Statistical Inference in Stochastic Gradient Descent

Constructing confidence regions for stochastic gradient descent (SGD) ideally requires estimating the asymptotic covariance matrix, a severe computational bottleneck in high dimensions. Traditional cancellation-based batch means methods bypass this estimation but require inverting a sample batch covariance matrix. This introduces...

💬 0 commentsarXiv:2608.30845v1PDF
0

Posted in stat.ME · 2026-08-31 · José María Lago, Albert Castellana, Edgars Nemše

Aggregate Disambiguation Systems

Natural-language tasks can elicit different verdicts from protocol-following evaluators that receive the same declared information. We study aggregate disambiguation systems (ADSs). Given a task and a candidate solution, each evaluator casts a binary vote on whether the solution should be accepted, and the system aggregates the votes...

💬 0 commentsarXiv:2608.30805v1PDF
0

Posted in stat.ME · 2026-08-31 · Yongqi Zhong, Anne-Renee Hartman, Jing Zhang

From Test Performance to Risk-Based Effect Sizes: A Unified Wald-Type Framework to Design Clinical Validation Studies for Binary and Survival Outcomes

Clinical validation studies of predictive tests are usually designed to focus on sensitivity ($Se$) and specificity ($Sp$), while statistical power is often calculated on regression-effect scales (e.g., risk ratio, hazard ratio). However, these quantities are statistically connected. Here, we provide closed-form links from...

💬 0 commentsarXiv:2608.30801v1PDF
0

Posted in stat.ML · 2026-08-31 · Nan Zheng, Hoi Yiu Cheung, Vibhu Sharma, James T. Thorson, Noel G. Cadigan

Implementing neural network mixed-effects models in Template Model Builder (TMB)

Neural network mixed-effects models (NMMs) have gained traction by combining the strong representation and predictive power of artificial neural networks with the capacity of mixed-effects modeling to capture complex correlation structures. However, existing estimation approaches rely heavily on manual derivations of objective...

💬 0 commentsarXiv:2608.31133v1PDF
0

Posted in stat.AP · 2026-08-31 · Manganaw N'Daam, Edoh Katchekpele, Tchilabalo Abozou Kpanzou

Long-Memory Estimation and Fractionally Integrated Modeling of White Maize Prices in Togo

Agricultural commodity prices often exhibit strong temporal persistence, which may limit the performance of conventional time series models. This study investigates long memory in logarithmic monthly white maize prices from six major markets in Togo between January 2001 and June 2022. Long memory is examined using the...

💬 0 commentsarXiv:2608.30569v1PDF
0

Posted in stat.ML · 2026-08-31 · Fariborz Setoudehtazang, Geoffrey J. McLachlan

Informative Label Missingness in Multiclass Classification Information Geometry and Excess Risk

Informative label missingness can change the usual efficiency ordering between completely and partially labelled classifiers because the pattern of missing labels may itself carry information about the classification model. We develop a general likelihood-based theory for this phenomenon in parametric multiclass classification. An...

💬 0 commentsarXiv:2608.30561v1PDF
0

Posted in stat.ME · 2026-08-31 · Bosen Cui, Yuhong Yang, Fan Yang

Power and sample size calculations for causal mediation analysis with a binary mediator in randomized trials

Mediation analyses are increasingly conducted in randomized trials, but a sample size adequate for the total treatment effect may leave the natural indirect effect (NIE) or natural direct effect (NDE) substantially underpowered. Randomization does not extend to the mediator, so precision depends on the conditional mediator...

💬 0 commentsarXiv:2608.30412v1PDF
0

Posted in stat.ML · 2026-08-31 · Mingzhi Song

Estimating Population-Risk Curves Along Nonconvex Gradient Flows from the Training Sample

We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample. Flow approximate leave-one-out (Flow-ALO) propagates a deletion response and evaluates omitted observations at approximate deleted paths. The risk-curve error decomposes into response approximation, exact-LOO...

💬 0 commentsarXiv:2608.30261v1PDF
0

Posted in stat.CO · 2026-08-31 · Takato Ueno, Shuji Kijima

GPU-Parallelization of Markov Chain Pool Decoding with Unbiased MCMC

Markov chain pool decoding (MCPD) devised by Knill et al. (1996) identifies likely positive clones from noisy pooled-test results. The standard MCPD estimates clone-wise posterior probabilities using Gibbs sampling, but it may allocate excessive computational effort to low-scoring clones. This paper focuses on parallelizing MCPD on...

💬 0 commentsarXiv:2608.30239v1PDF
0

Posted in stat.ML · 2026-08-31 · Darinka Dentcheva, Xiangyu Tian

Fairness in multi-class multi-group classification problems via contextial coherent risk measures

We propose a new design of fair classifiers for multi-class classification problems in the presence of vector-valued sensitive attributes. In that scenario each sensitive attribute has multiple values and forms several groups relevant to the fairness consideration. Naturally those groups are overlapping and one should also analyze the...

💬 0 commentsarXiv:2608.30223v1PDF
0

Posted in stat.ME · 2026-08-31 · Lorenzo Gasparollo, Mats J. Stensrud

Causal inference with staggered entries and effects that change over calendar time

Studies with staggered entry, in which individuals enroll at different calendar times, are ubiquitous in medicine and related disciplines. Because these studies usually have a fixed administrative end of follow-up, identification of the estimand of interest relies on assumptions about the right-censoring mechanism. The assumptions are...

💬 0 commentsarXiv:2608.30099v1PDF
0

Posted in stat.CO · 2026-08-31 · Jongmin Mun

Multifidelity Computer Model Emulation Via Diffusion Model Steering and Targeted Maximum Likelihood

We develop a multifidelity method for fusing low-resolution simulations with computationally expensive high-resolution simulations, which are run infrequently and are therefore prone to bias. We formulate this fusion as a constrained optimization under missing-not-at-random (MNAR) selection bias. This formulation searches for the...

💬 0 commentsarXiv:2608.30096v1PDF
0

Posted in stat.ME · 2026-08-30 · Anirban Mondal, Paromita Banerjee, Abhijit Mandal

Robust K-means Clustering using the Density Power Divergence Measure

We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm. Since Mahalanobis...

💬 0 commentsarXiv:2608.30093v1PDF
0

Posted in stat.ML · 2026-08-30 · Shulei Wang

Learning Representations through Token Prediction: Geometry, Approximation, and Downstream Guarantees

Token prediction is a central pre-training objective for modern language models. Despite its empirical success, why token prediction learns broadly useful representations remains incompletely understood. We develop a statistical framework connecting token prediction with representation geometry, encoder approximation, and downstream...

💬 0 commentsarXiv:2608.30072v1PDF
0

Posted in stat.ML · 2026-08-30 · Yasin Khadem Charvadeh, Grace Y. Yi, Mithat Gönen, Pouya Faroughi

A Deep Latent Variable Framework for Jointly Modeling Missingness, Measurement Error, and Heterogeneity

Missing data, measurement error, and population heterogeneity are pervasive challenges in analyzing data arising from modern observational studies and machine learning applications. Although these problems frequently coexist and interact, they are often treated separately in existing works. We propose a unified probabilistic framework...

💬 0 commentsarXiv:2608.30040v1PDF
0

Posted in stat.AP · 2026-08-30 · Zongyue Teng, Ningkun Zhou, Xinyu Zhang, Robert Wallis, Qingyan Xiang

Nonlinear trajectories of lung function recovery in patients with pulmonary disease: empirical evaluation of longitudinal modeling approaches

Introduction: Longitudinal lung function recovery after pulmonary disease commonly follows nonlinear trajectories, and failure to adequately model these trajectories can lead to biased or misleading estimates of treatment effects. However, an important methodological gap remains as there is limited assessment of statistical methods...

💬 0 commentsarXiv:2608.30015v1PDF
0

Posted in stat.ME · 2026-08-30 · Omar Alzeley, Michail Tsagris

Modelling compositional data with structural zero values

Compositional data are positive multivariate data whose sum equals 1. A popular method to analyze such data is via log--ratio transformations, which are however not applicable when zero values are present. In this paper we present a conditional logistic normal distribution suitable for compositional data with structural zero values....

💬 0 commentsarXiv:2608.29954v1PDF
0

Posted in stat.ME · 2026-08-30 · Eric Goldman, Fushing Hsieh

Design of Experiment in Complex Systems based on Computational Taxonomy

Via Computational Taxonomy (CT), we develop Design of Experiment(DoE) based on rigorously redefined constituting ingredients of complex system dynamics: randomness, nonlinearity and even class, through a data-driven constructed Taxonomic Hierarchy. As an opposite quest of Classification without man-made assumptions and structures, we...

💬 0 commentsarXiv:2608.29883v1PDF
0

Posted in stat.ME · 2026-08-30 · Dongxu Yang, Wanfeng Liang, Le Zhou, Long Feng

Spatial-sign-based multilinear principal component analysis for tensor data

Multilinear principal component analysis (MPCA) reduces the dimension of tensor-valued data while preserving their mode-specific structure, but its quadratic scatter criterion can be unstable under heavy-tailed distributions and contamination. We propose spatial-sign-based multilinear principal component analysis (SMPCA), a robust...

💬 0 commentsarXiv:2608.29862v1PDF