Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 05:21:11 EST

0

Posted in stat.ME · 2026-08-26 · Dan Han, Vicki Modisette, Ting Li, Akidul Haque

Empirical-Bayes Elastic-Net Computation for Exponential Random Graph Models

Exponential random graph models (ERGMs) describe dependence among network ties, but inference becomes difficult when the likelihood is intractable and candidate network statistics are strongly correlated. We introduce BERGM Elastic Net, an adaptive empirical-Bayes approach that combines lasso shrinkage with ridge stabilization in a...

💬 0 commentsarXiv:2608.25280v1PDF
0

Posted in stat.ME · 2026-08-26 · Guannan Zhai, Feifang Hu

Valid test for multi-arm trials with generalized linear models under covariate-adaptive randomization

Modern medical research, such as dose-finding studies, seamless trials, and shared control designs, often involves comparing multiple treatments simultaneously. Despite its wide applications, most research focuses on continuous endpoints, leaving the inference for general outcome types in high demand. In this article, we propose a new...

💬 0 commentsarXiv:2608.25272v1PDF
0

Posted in stat.AP · 2026-08-25 · Mohammed Adjieteh, Vytaras Brazauskas

Quantile and Log-Quantile Least Squares for Robust-Efficient Fitting and Validation of Log-Location-Scale Loss Models

\begin{quote} {\bf\em Abstract\/}. ~A variety of models for insurance and other types of losses are special cases of the {\em log-location-scale\/} family, with the lognormal and Pareto-$I$ distributions being the most prominent examples. The latter also serves as a primary example of infinite-mean models that often present challenges...

💬 0 commentsarXiv:2608.25234v1PDF
0

Posted in stat.ME · 2026-08-26 · Manuel Pfeuffer, Roshan Prakash Rane, Kerstin Ritter, Sonja Greven

Controlling for Omitted Variable Bias in Deep Neural Networks

Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictions when these...

💬 0 commentsarXiv:2608.25930v1PDF
0

Posted in stat.ME · 2026-08-25 · Razieh Nabi, Anna Guo, Lin Liu

Toward a Semiparametric Efficiency Theory under Equality Constraints in Nested Markov Models

Probabilistic models of Directed Acyclic Graphs (DAGs) with latent variables impose equality constraints on the observed data distribution beyond ordinary conditional independencies. These so-called Verma constraints arise in nested Markov models associated with Acyclic Directed Mixed Graphs, the latent projection of latent-variable...

💬 0 commentsarXiv:2608.24602v1PDF
0

Posted in stat.ML · 2026-08-25 · Janis Aiad, Aghiles Drali, Aymen El Ouadrhiri, Anass Ettahiri, Yasser Oufqir, Simon Patry, David Cortes, Marianne Clausel, Emilie Devijver

Scalable and Versatile Identification for Hierarchical Structural Causal Models: A New Look at Project STAR

The STAR (Student-Teacher Achievement Ratio) experiment (1985, Tennessee, USA) is a landmark hierarchical dataset designed to assess the impact of class size on student outcomes, with observations nested within classes. To encode class-level interventions in such hierarchical settings, we develop a complete, scalable, open-source...

💬 0 commentsarXiv:2608.24500v1PDF
0

Posted in stat.ME · 2026-08-25 · Riccardo Rastelli, Shizhe Chen

A latent space network model for dynamic neural latent embedding

We introduce a novel latent space network model for analyzing multivariate time series of neural spike-train data. The methodology is motivated by an experimental study in mice, where neuronal responses were collected under a sequence of visual discrimination tasks. We adopt a latent variable framework to model the firing rates of...

💬 0 commentsarXiv:2608.24452v1PDF
0

Posted in stat.ML · 2026-08-25 · Rafael Oliveira

Sequential operator learning under dependent data

Learning operators from sequentially collected data arises in adaptive experimental design, Bayesian optimization, and dynamical-system modelling, where observations may be dependent, and future inputs or sensing operators may depend on preceding data. We derive time-uniform self-normalized concentration bounds for stochastic...

💬 0 commentsarXiv:2608.24426v1PDF
0

Posted in stat.ME · 2026-08-25 · Sota Osumi, Akira Okazaki, Shuichi Kawano

Groupwise Predictor Envelope Models for Multivariate Linear Regression

Envelope methods improve estimation efficiency in multivariate analysis by isolating low-dimensional structures that contain all the information material to the parameter of interest. In multivariate linear regression with random predictors, predictor envelope models achieve this goal by removing variation in the predictors that is...

💬 0 commentsarXiv:2608.24371v1PDF
0

Posted in stat.ME · 2026-08-25 · Lena Schemet, Sarah Friedrich-Welz

Wild Bootstrap and Efron's Bootstrap for Debiased Cox Regression

Cox regression with Lasso penalization is widely used for variable selection in time-to-event data, but reliable coefficient inference after selection remains difficult. We investigate bootstrap inference for the debiased Cox estimator after Cox Lasso selection. Two score-based procedures are considered: a wild bootstrap using...

💬 0 commentsarXiv:2608.24230v1PDF
0

Posted in stat.ML · 2026-08-25 · Soham Chatterjee, Rwitobroto Dey, Smarajit Bose

A Heterogeneous Mixture of Experts Framework for Interpretable Machine Learning

Mixture-of-Experts (MoE) models provide a flexible framework for partitioning complex prediction problems into simpler local learning tasks through an input-dependent gating mechanism. Existing interpretable MoE approaches, such as Mixture of Decision Trees (MoDT), achieve transparency by employing homogeneous decision-tree experts,...

💬 0 commentsarXiv:2608.24195v1PDF
0

Posted in stat.ML · 2026-08-25 · Zhongli Jiang, Min Zhang, Dabao Zhang

qshap: Fast Shapley Decomposition of $R^2$ for Gradient-Boosted Trees

Numerous methods have been developed to quantify feature attributions in individual predictions for tree ensembles. However, many applications require global measures of feature contributions to overall model performance. Although local attribution scores can be aggregated to characterize feature importance, such summaries do not...

💬 0 commentsarXiv:2608.24104v1PDF
0

Posted in stat.ME · 2026-08-25 · Jiaqi Tong, Fan Li

Orthogonal double residual learning for optimal individualized treatment rules

Individualized treatment rules (ITRs) map baseline characteristics to treatment recommendations, with the optimal ITR maximizing expected reward or policy welfare. Indirect methods may require restrictive modeling assumptions, whereas direct methods can be sensitive to nuisance estimation error and limited overlap. We propose...

💬 0 commentsarXiv:2608.24085v1PDF
0

Posted in stat.ME · 2026-08-25 · Sergei Pankratev, Palash Arora

CUPED on Steroids: Multivariate Covariate Adjustment for Switchback Experiments

Controlled-experiment Using Pre-Experiment Data (CUPED) reduces the variance of the treatment effect estimator in online experiments by adjusting the in-experiment outcome metric using its lagged pre-experiment value. This method can be strengthened by enriching its covariate set while keeping it automatable and guarding against...

💬 0 commentsarXiv:2608.24038v1PDF
0

Posted in stat.ME · 2026-08-25 · Taehyeon Koo, Elizabeth A. Stuart, Kara E. Rudolph, Caleb H. Miles

Causal Effects of Modified Treatment Policies under Positivity Violations: A Partial Identification Approach

Modified treatment policies (MTPs) are interventions based on each individual's natural treatment value. We study mean outcomes under MTPs for continuous treatments, including exposure mixtures. Positivity is the standard sufficient condition for identifying these mean outcomes without extrapolation: policy-generated values remain...

💬 0 commentsarXiv:2608.23971v1PDF
0

Posted in stat.ML · 2026-08-25 · Victor Medina-Olivares, Stefan Lessmann, Jonathan Crook

$\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions

Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regression remains attractive because its coefficients are easy to interpret, but it can miss nonlinear structure. Flexible models can improve prediction, but their explanations are often...

💬 0 commentsarXiv:2608.24582v1PDF
0

Posted in stat.ML · 2026-08-25 · Hao Chen

What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation

Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet,...

💬 0 commentsarXiv:2608.24881v1PDF
0

Posted in stat.ME · 2026-08-24 · Stanislav Škorňa, Jitka Machalová

Penalized likelihood estimation of probability density functions using compositional splines

Probability density functions are commonly estimated through preliminary smoothing or aggregation procedures, e.g., histograms or kernel density estimation, before subsequent functional representation and functional data analyses. Such a two-stage approach can lead to additional approximation bias and weaken the direct connection...

💬 0 commentsarXiv:2608.23512v1PDF
0

Posted in stat.ML · 2026-08-24 · Jiaming Qiu, Yingye Zheng, Ying-Qi Zhao

Primal--Dual Alternating Neural Learning for Timely Classification with Performance Guarantees

Timely risk classification is essential in many clinical monitoring settings, where decisions must balance the benefit of classifying patients early for subsequent intervention against the value of observing additional data. Yet most existing statistical and machine-learning methods are designed for fully observed trajectories and...

💬 0 commentsarXiv:2608.23480v1PDF
0

Posted in stat.ME · 2026-08-24 · Sarika Aggarwal, Brent A. Coull, Nima Hejazi, Rachel C. Nethery

Evaluating the effects of policy interventions subject to early adoption: A case study of prescription drug monitoring programs and opioid dispensing

Policies that require organizations to use new systems, such as prescription drug monitoring programs (PDMPs), are often implemented in phases, with an initial period of voluntary access followed by mandated compliance. This allows the policy intervention to be adopted before compliance is required (early adoption), causing outcomes...

💬 0 commentsarXiv:2608.23472v1PDF
0

Posted in stat.AP · 2026-08-24 · Steeven B. Affognon, Babacar M. Ndiaye, Pierre Mendy, Cheikh M. F. Kebe

From Daily Fluctuations to Annual Hydrological Cycles: A Wavelet-Based Analysis of Nonstationary Seasonality in Senegal River Hydropower Inflows

This study presents a reproducible framework combining Fourier and wavelet analysis to examine the seasonality of daily inflows at three sites on the Senegal River (Bafing Makana, Felou, Gouina), based on 65,631 daily observations spanning nearly 60 years (1961-2020). Using harmonic regression, Welch spectral analysis, stationary...

💬 0 commentsarXiv:2608.23470v1PDF
0

Posted in stat.ME · 2026-08-24 · Anik Burman, Margaret Gamalo, Promit Ghosal, Prosenjit Kundu

Transporting Randomized Trial Effects to Real-World Populations via Riesz-Calibrated Optimal Transport

Randomized trials support causal inference, but differences between trial and target populations can limit the transportability of treatment effects to real-world settings. Many existing approaches model the propensity of trial participation and can therefore be sensitive to model misspecification and weak overlap of the covariate...

💬 0 commentsarXiv:2608.23453v1PDF
0

Posted in stat.OT · 2026-08-24 · Kaitlyn G Fitzgerald

Becoming Good Stewards of Information: A framework for integrating ethical, civic, and professional formation throughout the statistics and data science curriculum

Recent recommendations in statistics and data science education emphasize goals that extend beyond content mastery, including statistical literacy, evidence-based decision-making, communication, ethics, responsible use of data, and civic responsibility. We argue these goals can all be understood through a common lens: helping students...

💬 0 commentsarXiv:2608.23352v1PDF
0

Posted in stat.ML · 2026-08-24 · Gordei Verbii

One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion

One implicit DDIM inversion step is the cheapest probe of whether a pretrained diffusion model encodes local manifold geometry. It is the stationarity condition of an explicit potential, $x-G(x)=\nablaΨ_t(x)$, strongly convex at the Bayes limit with modulus exactly $e^{-h_t}$ for the step's log-SNR gap $h_t$ $-$ for every data law,...

💬 0 commentsarXiv:2608.23094v1PDF
0

Posted in stat.CO · 2026-08-24 · V. Masarotto

fdWasserstein: Optimal Transport Methods for Covariance Operators of Functional Data

Data increasingly arrive as collections of curves - a voice recording, a growth trajectory, a day of sensor readings - where each observation is a whole function rather than a single number. The usual question asked of such data is how the average curve differs from one group to the next. But the average is only half the picture: two...

💬 0 commentsarXiv:2608.22921v1PDF