Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 12:09:11 EST

0

Posted in stat.ML · 2026-07-28 · Daniel Kua, Yan Song

Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?

Deep generative models (DGMs) are widely used for complex high-dimensional data and increasingly applied to spatial and spatio-temporal modeling. Their generated samples implicitly represent the learned data distribution and associated uncertainty. However, for real-world data, assessing whether DGMs have learned the underlying...

💬 0 commentsarXiv:2607.25929v1PDF
0

Posted in stat.ME · 2026-07-28 · Arjun Sondhi

Bias-corrected Cox regression with AI-extracted covariates via calibration summary statistics

Large-scale observational studies increasingly rely on AI pipelines to extract structured variables from unstructured clinical records. A common workflow separates the data vendor, who validates extraction accuracy with a gold-standard sample, from the downstream researcher, who receives only the extracted dataset and summary accuracy...

💬 0 commentsarXiv:2607.25868v1PDF
0

Posted in stat.ME · 2026-07-28 · Marie-Félicia Beclin, Apolline Courrèges-Vartanian, Geneviève Lefebvre, Tat-Thang Vo

Causally Interpretable Meta-Mediation Analysis With Missing At Random Mediator and Outcome Data

Meta-analyzing natural indirect effect estimates from multiple studies is increas- ingly used to synthesize evidence on causal pathways of interest. However, stan- dard mediation meta-analysis approaches are typically based on structural equation modeling, which fails to account for mediator-outcome confounding, is not read- ily...

💬 0 commentsarXiv:2607.25822v1PDF
0

Posted in stat.AP · 2026-07-28 · Samuel Pawel, Saverio Fontana, Jinyu Chen, Leonie Stoltefuß, Frank Weber, Guido Skipka, Sibylle Sturtz, Ralf Bender, Leonhard Held

Edgington's Combination Method for Two-Study Meta-Analysis: An Empirical Evaluation in 1226 Meta-Analyses

Two-study meta-analyses are common in evidence synthesis but pose major statistical challenges. With only two studies, the between-study variance cannot be reliably estimated, rendering standard random-effects methods unstable. Here, we investigate meta-analyses based on Edgington's p-value combination method as an alternative...

💬 0 commentsarXiv:2607.25819v1PDF
0

Posted in stat.ME · 2026-07-28 · Luca Benetti, Gianluca Baio, Anna Heath

Calculating the Expected Value of Sample Information accounting for missing data

The Expected Value of Sample Information (EVSI) is a powerful instrument to determine the value of additional evidence to inform an economic model. However, EVSI has been applied only to idealized data collection mechanisms, thereby reducing its potential applications in realistic studies. In this paper, we define a methodology to...

💬 0 commentsarXiv:2607.25775v1PDF
0

Posted in stat.ME · 2026-07-28 · Pier Giovanni Bissiri, Riccardo Corradin, Andrea Ongaro

Nonparametric Bayesian inference for the Gini-Simpson index

Many statistical problems concern the analysis of species distributions or, more generally, of discrete labeled quantities. Assessing species diversity constitutes a key step toward understanding population structure, and the Gini-Simpson index is among the most widely adopted diversity measures. In this manuscript, we examine several...

💬 0 commentsarXiv:2607.25737v1PDF
0

Posted in stat.ME · 2026-07-28 · Tran Trong Khoi Le, Pham Hien Trang Tu, Nhat Long Ngo, Tat-Thang Vo

On the magnitude, sign and ranking of recanting-twin path-specific effects

The framework of recanting twin path-specific effects has recently been propose to address the issue of intermediate confounding in causal mediation analysis, enabling the decomposition of the average treatment effect into identifiable fine-grained path-specific effects (PSEs). An open question, however, is the extent to which...

💬 0 commentsarXiv:2607.25709v1PDF
0

Posted in stat.ME · 2026-07-28 · Masahiro Fujisawa, Masaki Adachi, Takuo Matsubara

Generalised Robust Bayes for Joint Inference of Model and Contamination

Generalised Bayesian inference (GBI) has emerged as a compelling robust alternative to standard Bayesian inference, mitigating sensitivity to data contamination by replacing the log-likelihood with a robust loss or divergence. However, existing robust GBI frameworks typically provide only qualitative robustness: while they can make...

💬 0 commentsarXiv:2607.25665v1PDF
0

Posted in stat.AP · 2026-07-28 · Jihyun Park, Jieun Kim, Taehan Bae, Jae Youn Ahn

Can a small additional claim lower the premium? Credibility orders for collective risk models

The collective risk model is a fundamental framework in insurance ratemaking for modeling aggregate losses by combining claim frequency and claim severity components. A key structural requirement for a reliable experience rating system is a monotone ordering property: policyholders with worse past experience should receive a...

💬 0 commentsarXiv:2607.25623v1PDF
0

Posted in stat.AP · 2026-07-28 · Léa Gondian, Thimothée Thiery

Validation of methods to estimate the uncertainty of buildings energy savings in a controlled numerical setting and Bayesian energy signature with autocorrelated errors

In the field of building energy efficiency, the measurement and verification (M&V) of energy savings following energy efficiency measures often relies on the use of a calibrated statistical model. In order to obtain reliable estimates, the estimation of uncertainties associated with this procedure is recognized as a crucial aspect of...

💬 0 commentsarXiv:2607.25382v1PDF
0

Posted in stat.ME · 2026-07-28 · Kotaro Sasaki, Hisashi Noma

Penalized likelihood inference for beta-binomial meta-analysis of proportions of rare events

In meta-analyses of proportions, the event of interest is often rare, resulting in sparse event counts and frequent zero-event studies. The beta-binomial model has been used as a flexible random-effects model for pooling overdispersed and rare-event proportions. However, the commonly used maximum likelihood estimator (MLE) may be...

💬 0 commentsarXiv:2607.25320v1PDF
0

Posted in stat.AP · 2026-07-28 · Juan Francisco, Mandujano Reyes

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks

Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items. Existing neural IRT approaches, including PSN-IRT, estimate these quantities using point estimates, limiting uncertainty...

💬 0 commentsarXiv:2607.25257v1PDF
0

Posted in stat.ME · 2026-07-28 · Deepani Hemachandra, Jagath Senarathne, Mahasen Dehideniya

A Copula-Based Regression Framework for Enhanced Prediction under Heteroscedasticity

Classical regression approaches, including ordinary least squares, rely on strong assumptions such as constant variance and normality of residuals, which are often violated in real-world data. Although log-transformation is commonly used to stabilise variance, it may introduce re-transformation bias and fail to address...

💬 0 commentsarXiv:2607.25250v1PDF
0

Posted in stat.ME · 2026-07-28 · Chunlei Ge, W. John Braun

Differential Equation-Constrained Exponential-Type Local Polynomial Regression Under Model Misspecification

The issue of model misspecification is critical, yet it is often regarded as unavoidable in applied statistical modeling. Model misspecification can be mitigated by incorporating informative features and strengthening model formulations, such as through the integration of domain knowledge or structural constraints. In this paper, we...

💬 0 commentsarXiv:2607.25248v1PDF
0

Posted in stat.ML · 2026-07-28 · Zeyu Bian, Ying Zhou, Yifan Cui

Learning from the Unseen: Offline Reinforcement Learning with Hidden Actions

Standard offline reinforcement learning (RL) algorithms typically assume that the actions in the dataset are observed without error. However, in many real-world applications, the true actions are unobserved and only noisy proxies are available, causing existing RL methods to yield biased and potentially misleading conclusions. We...

💬 0 commentsarXiv:2607.25241v1PDF
0

Posted in stat.ML · 2026-07-28 · Michael Pokojovy, J. Marcus Jobe, Simon Lacoste-Julien

Lloyd's $K$-Means Clustering Algorithm Is Frank-Wolfe in Disguise

Lloyd's $K$-means algorithm, also known as naïve $K$-means, is a widely used ad hoc optimization heuristic, designed to minimize the sum of squared errors (SSE) across all $K$-partitions of a dataset via iterative cluster refinement. In this work, we establish a novel connection between Lloyd's algorithm and the Frank-Wolfe (FW)...

💬 0 commentsarXiv:2607.25190v1PDF
0

Posted in stat.ME · 2026-07-27 · Garrett Frady, Dipak K. Dey, Shariq Mohammed

Bayesian Feature Extraction using Gaussian and Diffused-gamma Priors for High Dimensional Spatio-Temporal Data

High-dimensional data with sparse structure and spatio-temporal dependence arise in many scientific domains. We develop a Bayesian feature-extraction framework for spatio-temporal settings that employs Gaussian and Diffused-gamma priors to induce structured sparsity. The modeling framework specifies a general likelihood via Bregman...

💬 0 commentsarXiv:2607.24378v1PDF
0

Posted in stat.ML · 2026-07-24 · Yichen Gu, Yuxuan Song, Weizhou Qian, Yixin Wang, Joshua Welch

Amortized Bayesian Causal Discovery of Extended Factor Graphs

Learning causal graphs from interventional data is a challenging problem with broad applications. In molecular biology, for example, a central goal is to uncover gene regulatory networks from large-scale perturbation data. An ideal algorithm for this task should scale to thousands of nodes, incorporate interventions even when their...

💬 0 commentsarXiv:2607.22934v1PDF
0

Posted in stat.ML · 2026-07-21 · Nived Rajaraman

The Price of Hidden Curvature: An $\widetildeΩ (d^{5/4} \sqrt{T})$ Lower Bound for Bandit Convex Optimization

We establish a $\widetildeΩ(d^{5/4}\sqrt T)$ lower bound on the minimax expected regret of stochastic bandit convex optimization of $1$-Lipschitz functions on the Euclidean ball. This presents the first nontrivial regret lower bound that grows faster than $d\sqrt{T}$ for this problem, establishing that stochastic bandit convex...

💬 0 commentsarXiv:2607.18652v2PDF
0

Posted in stat.ME · 2026-07-20 · Muhammad Qasim, Kai Wang, Ishan S Bhatt

Adaptive Penalization and Bootstrap-Smoothed Inference for Two-Sample Mendelian Randomization with Summary Data

Two-sample Mendelian randomization (MR) uses genetic variants as instrumental variables to estimate causal effects from observational data using summary association statistics. However, horizontal pleiotropy can invalidate standard MR estimators and lead to biased causal inference. Pleiotropy-robust methods have been proposed to...

💬 0 commentsarXiv:2607.18503v1PDF
0

Posted in stat.ME · 2026-07-20 · Kihyun Han, Yanyuan Ma, Karen Marder, Tanya P. Garcia

SPYCE: A Doubly Robust Estimator for Trials Targeting Early Huntington Disease under Outcome-Dependent Censoring

Clinical trials for neurodegenerative diseases must identify sensitive endpoints -- outcomes that change rapidly enough to detect treatment effects. In Huntington disease, this requires measuring how outcomes change as participants approach Stage 1. Yet many participants exit studies before reaching this stage, making their time to...

💬 0 commentsarXiv:2607.18501v1PDF
0

Posted in stat.ME · 2026-07-20 · Masahiro Kojima, Hisato Sunami, Masaaki Kuriki

A Globally Calibrated Bayesian Optimal Phase II Design for Adaptive Enrichment Trials

Adaptive enrichment can allow the development of an experimental treatment to continue when its activity is insufficient in an all-comer population but remains promising in a prespecified biomarker-positive subgroup. However, a straightforward sequential application of separately calibrated phase II designs to the two populations can...

💬 0 commentsarXiv:2607.17692v2PDF
0

Posted in stat.ML · 2026-07-20 · Vignesh Tirukkonda, Gautam Dasarathy

Mixing-Free and Signal-Optimal Learning of Gaussian Graphical Models from Glauber Dynamics

Gaussian graphical model selection is usually studied under independent sampling, but in many applications the data arise as a single trajectory of a dependent stochastic process. We study exact recovery of the graph from one trajectory of random-scan Gaussian Glauber dynamics. Existing techniques for this problem either inherit the...

💬 0 commentsarXiv:2607.18559v1PDF
0

Posted in stat.ME · 2026-07-20 · Soham Bakshi, Lingjun Gao, Zijun Gao, Snigdha Panigrahi

Flexible Inference for Winners with Conditional Validity

Researchers often select top-performing options or winners, based on a data-driven criterion, such as treatments, models, or model features and then report effect estimates for the selected winners. Naive post-selection estimates, however, are known to suffer from the winner's curse, producing systematically overoptimistic results. We...

💬 0 commentsarXiv:2607.18545v1PDF
0

Posted in stat.ME · 2026-07-20 · Shuhe Wang, Matthew T. Slaughter, Jennifer C. Nelson, Brian D. Williamson

Using binary silver labels in electronic health records-based computable phenotyping algorithms

Gold-standard phenotype labels are often unavailable at scale in electronic health record (EHR) studies because they require manual chart review. Weakly supervised phenotyping methods instead use silver-standard labels, such as diagnosis-code counts, natural language processing (NLP) mentions, medication indicators, or laboratory...

💬 0 commentsarXiv:2607.18431v1PDF