Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 00:31:53 EST

0

Posted in stat.ML · 2026-09-02 · Roser Homs, Olga Kuznetsova, Bernadette J. Stolz

A computational approach to maximum likelihood thresholds for colored Gaussian graphical models

Gaussian graphical models (GGMs) are essential tools for interpretable structure learning. However, in high-dimensional, small-sample regimes, the available data is often insufficient for the maximum likelihood estimator to exist. Colored Gaussian graphical models (CGGMs) mitigate this limitation by imposing symmetry constraints...

💬 0 commentsarXiv:2609.02382v1PDF
0

Posted in stat.ML · 2026-09-02 · Shizhe Zhang, Mingyang Zhao, Lei Ma

Schrödinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation

Generative modeling directly on geometric manifolds can avoid errors introduced by flattening non-Euclidean data, repeated ambient projection, and coordinate inconsistency in Euclidean representations. Schrodinger bridges provide a probabilistic generative framework for entropy-regularized transport between prescribed endpoint...

💬 0 commentsarXiv:2609.02196v1PDF
0

Posted in stat.ML · 2026-09-02 · Ming Tan, Xiyun Jiao

HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC

Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their performance depends strongly on hyperparameters such as the step size, mini-batch size, and number of leapfrog steps. Since most SGMCMC algorithms lack a Metropolis-Hastings acceptance rate, standard acceptance-based tuning...

💬 0 commentsarXiv:2609.02138v1PDF
0

Posted in stat.ME · 2026-09-02 · Difan Song, V. Roshan Joseph

Efficient Screening Designs for Expensive Black-box Models with Qualitative and Quantitative Factors

Computationally expensive black-box models often involve a large number of input factors with complex interactions and varying importance. Experimental design techniques can be used for quickly identifying the important factors, which can make the optimization of a complex computer model or the training of an expensive machine...

💬 0 commentsarXiv:2609.02087v1PDF
0

Posted in stat.ME · 2026-09-02 · Tingxuan Han, Ke Deng

Data-Adaptive Rerandomization for 2K Factorial Designs

Factorial designs allow simultaneous estimation of multiple main effects and interactions, but covariate imbalance can substantially reduce estimation precision. Existing rerandomization methods improve covariate balance yet do not fully exploit heterogeneous priorities across factorial effects or effect-specific covariate importance....

💬 0 commentsarXiv:2609.02078v1PDF
0

Posted in stat.ML · 2026-09-02 · Xiaowen Dong, Hoi-To Wai, Siheng Chen, Laura Toni, Dorina Thanou

From topology learning to graph generation: A unifying perspective

Learning graph structures from data is a fundamental problem that spans a wide range of signal processing and machine learning tasks. While significant effort has been made to tackle the problem, existing research has largely evolved along two parallel directions. The first seeks to infer the topology of an individual graph from...

💬 0 commentsarXiv:2609.02286v1PDF
0

Posted in stat.ML · 2026-09-02 · Troy Butler, Tianyi Jiang, João Silva, Harri Hakula, Timothy Wildey

Copula Transformations for Data-Consistent Inversion

Data-consistent inversion (DCI) constructs probability measures whose push-forward distributions agree with observed data, while iterative data-consistent inversion (iDCI) extends this framework to generalized stochastic inverse problems by enforcing multiple push-forward constraints sequentially. Although iDCI avoids the direct...

💬 0 commentsarXiv:2609.02832v1PDF
0

Posted in stat.AP · 2026-09-02 · Guang Yang, Wei Shi, Yuan Cao, Long Feng

Learning CNN Filters via Generalized Stein's Method

Convolutional Neural Networks (CNNs) have undoubtedly revolutionized image data analysis and the field of computer vision. As the cornerstone of CNNs, the convolution operation enables the networks to extract abstract features and uncover hidden relationships in the image data. This paper considers the problem of estimating...

💬 0 commentsarXiv:2609.02875v1PDF
0

Posted in stat.ME · 2026-09-02 · Manuela-Simona Cojocea

Statistical Inference for Probability Barycenters and Kolmogorov Moments

A probability coordinate chart is a continuous strictly increasing bijection that transports observations to the open unit interval, where averaging is always well defined. For barycentric inference, the first coordinate moment is returned through the inverse chart. Higher initial coordinate moments similarly generate initial...

💬 0 commentsarXiv:2609.02869v1PDF
0

Posted in stat.ML · 2026-09-01 · Zhaoliang Yuan, Jie Wang

Variable Selection for Feature-Based Newsvendor

Feature-based newsvendor models use observable covariates to tailor inventory decisions, aiming to balance holding and shortage costs under demand uncertainty. However, high-dimensional feature sets often hinder interpretability and inflate data collection and implementation costs. This paper studies variable selection for the...

💬 0 commentsarXiv:2609.01544v1PDF
0

Posted in stat.ME · 2026-09-01 · David Bolin, Alexandre de Bustamante Simas, Erik Karlsson Strandh, Jonas Wallin

Gaussian Processes on Directed Metric Graphs

We introduce a statistical framework for Gaussian fields indexed at arbitrary edge locations on general compact directed metric graphs. The construction is based on a stochastic differential equation with a first-order operator and conditions at the vertices. We characterise well-posedness and identify the covariance reproducing...

💬 0 commentsarXiv:2609.01435v1PDF
0

Posted in stat.ML · 2026-09-01 · Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein

On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study

Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. We formalize augmentation as a...

💬 0 commentsarXiv:2609.01410v1PDF
0

Posted in stat.ML · 2026-09-01 · Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate

Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity

The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictions for the same inputs (predictive multiplicity). Existing work primarily focuses on multiplicity within individual models, but in more complex decision systems, the impact of the Rashomon effect is less well understood. In...

💬 0 commentsarXiv:2609.01397v1PDF
0

Posted in stat.ML · 2026-09-01 · Ziqi Zhao, Qingjian Ni

Matched Queries for Curvature and Density at Branching Junctions

At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center. Recovering this missing information is necessary for describing local continuation beyond a single point, but finite observations must...

💬 0 commentsarXiv:2609.01319v1PDF
0

Posted in stat.ME · 2026-09-01 · Peter Cotton

Scalable Inversion of Contests with Correlated Performances, Including Softmax and Multinomial Probit

Multinomial probit choice probabilities over n alternatives are Gaussian orthant integrals, computed by simulation for thirty years, one expensive integral per alternative. Inversion, which is to say determining item attractiveness consistent with a prescribed choice probability vector, is even more difficult and has been considered...

💬 0 commentsarXiv:2609.01133v1PDF
0

Posted in stat.ML · 2026-09-01 · Marco Simnacher, Georg Keilbar, Benjamin König, Christoph Lippert, Sonja Greven

Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for large language model (LLM) outputs,...

💬 0 commentsarXiv:2609.00946v1PDF
0

Posted in stat.ML · 2026-09-01 · Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan

Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models

We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. The feature is observed for all data, whereas some class labels are missing. The probability of a missing label is modelled as a function of classification uncertainty, giving a feature-dependent...

💬 0 commentsarXiv:2609.00774v1PDF
0

Posted in stat.ME · 2026-09-01 · Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan

Deep Skew-t Mixture Models

High-dimensional clustering is challenging when component distributions are both heavy-tailed and directionally asymmetric. We propose a deep skew-$t$ mixture model (DStMM), a hierarchical factor-analytic mixture based on the generalised-hyperbolic skew-$t$ normal mean--variance representation. A shared inverse-gamma mixing variable...

💬 0 commentsarXiv:2609.00773v1PDF
0

Posted in stat.ME · 2026-09-01 · Mohammad Alhyari, Haziq Jamil, Hans Montcho, Håvard Rue

Deterministic Leave-One-Cluster-Out Cross-Validation for Multilevel Bayesian Structural Equation Models

We introduce a closed-form, refit-free procedure for leave-one-cluster-out (LOCO) cross-validation in multilevel Gaussian Bayesian structural equation models (SEMs), together with predictive scoring of every nested submodel. Conditional independence of clusters given the parameters expresses the cluster-deleted posterior as a...

💬 0 commentsarXiv:2609.00670v1PDF
0

Posted in stat.ME · 2026-09-01 · Hanzhang Lu, Jeffrey L. Andrews, Ryan P. Browne

An efficient EM algorithm for both element-wise and structural missingness in matrix-variate normal mixture models

Matrix-variate data with missing entries arise frequently in applications where observations are naturally organized as two-dimensional arrays. Although the matrix normal distribution provides a parsimonious model through its Kronecker covariance structure, standard EM estimation can be computationally expensive because arbitrary...

💬 0 commentsarXiv:2609.00616v1PDF
0

Posted in stat.ME · 2026-09-01 · Qi Kuang, Yin Xia

Anytime-Valid Distribution Shift Detection via Predictive Rank Martingales

Many sequential distribution shift detectors update a growing reference set with incoming observations. After a change, this update contaminates the reference set with post-change observations and can weaken subsequent evidence. Keeping the calibration sample fixed mitigates this contamination, but repeated reuse induces dependence...

💬 0 commentsarXiv:2609.00536v1PDF
0

Posted in stat.ME · 2026-08-31 · Juhee Lee, Kun Xia, Jianrui Zhang, Gongjun Xu, Qing Lu, Chenxi Li

Genetic association testing with multivariate survival phenotypes under interval censoring

Set-based genetic association tests provide a powerful framework for detecting genetic effects on complex traits by jointly analyzing multiple genetic variants. Although set-based methods have been developed for interval-censored survival outcomes, existing approaches primarily focus on a single survival phenotype and therefore do not...

💬 0 commentsarXiv:2609.00456v1PDF
0

Posted in stat.CO · 2026-08-31 · Sam Power

Non-Uniform Random Scans in Gibbs Sampling and CAVI

Gibbs sampling and coordinate ascent variational inference (CAVI) are two basic coordinate-wise methods for statistical computation. Recent analyses under strong log-concavity establish convergence rates for versions of these algorithms that update one uniformly selected block at each step. We extend both results to arbitrary fixed,...

💬 0 commentsarXiv:2609.00408v1PDF
0

Posted in stat.ML · 2026-08-31 · Caixia Xu, Piotr Fryzlewicz

A convolutional framework for detecting event-driven dynamics in energy price series

This paper develops a general convolutional neural network (CNN) framework for detecting heterogeneous event-driven dynamics in univariate time series windows. We show that the induced CNN class exactly represents classifiers based on range, maximum drawup, maximum drawdown and slope change, and uniformly approximates realised...

💬 0 commentsarXiv:2609.00402v1PDF
0

Posted in stat.ME · 2026-08-31 · Khai Nguyen, Elizabeth Juarez-Colunga, Peter Mueller

Generalized Bayesian Clustering with Regression for Unaligned Longitudinal Binary Data

We propose a generalized Bayesian clustering with regression model for unaligned longitudinal binary outcomes, motivated by seizure diary data from the Human Epilepsy Project. Seizure diaries are sparse, irregularly observed, and vary enormously across patients. A single fully-specified generative model tends to be either misspecified...

💬 0 commentsarXiv:2609.00307v1PDF