Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 08:17:18 EST

0

Posted in stat.ME · 2026-08-19 · Felix Boakye Oppong, Dimitris Rizopoulos, Thierry Gorlia, Nicole Erler

Functional forms in joint models for longitudinal and time-to-event data: A practical guide with application and interpretation

Background: Joint models for longitudinal and time-to-event data are widely used in clinical research. However, the choice of functional form linking the biomarker trajectory to event risk is often treated as a technical detail, despite its importance for model assumptions and interpretation. Default specifications may fail to capture...

💬 0 commentsarXiv:2608.18858v1PDF
0

Posted in stat.ME · 2026-08-19 · Shivshankar Nila, Ishapathik Das, N. Balakrishna

Robust Modeling of Extremes in the Presence of Inliers with Enhanced Tail Estimation

Extreme value theory provides a fundamental framework for modeling rare and extreme events; however, threshold selection remains a persistent challenge, particularly in the presence of inliers such as instantaneous or early failures. Such observations commonly arise in applications including reliability studies and environmental data,...

💬 0 commentsarXiv:2608.18735v1PDF
0

Posted in stat.ME · 2026-08-19 · Per August Jarval Moen, Sebastian Grau Nielsen, Espen Bjørge Urheim, Martin Tveten, Ingrid Kristine Glad

gridcp: Fast Online Changepoint Detection in Python

Online changepoint detection is the problem of detecting distributional changes in a data stream in real-time. A large body of methodology exists for the offline (fixed-size) setting, but applying these methods online quickly becomes infeasible since the per-observation computational cost and memory consumption typically grow at least...

💬 0 commentsarXiv:2608.18695v1PDF
0

Posted in stat.CO · 2026-08-19 · Hongru Zhao, Huiqian Feng

Convex Reparameterization and Self-Concordant Algorithms for Multivariate Regression with Covariance Estimation

Building on a reparameterization for multivariate linear regression that yields a jointly convex penalized likelihood in the reparameterized regression coefficient matrix and the precision matrix, we show that the resulting scaled Gaussian loss is standard self-concordant. This places the joint estimation problem within composite...

💬 0 commentsarXiv:2608.18441v1PDF
0

Posted in stat.ME · 2026-08-19 · Kentaro Takeda, Masahiro Kojima

A seamless dose-optimization design for monotherapy and combination therapy

The emergence of molecular-targeted agents and immune-oncology therapies has fundamentally transformed oncology drug development, necessitating evolution beyond traditional dose-finding approaches designed for cytotoxic agents. While conventional agents exhibit predictable monotonic dose-response relationships, novel anticancer agents...

💬 0 commentsarXiv:2608.18435v1PDF
0

Posted in stat.ME · 2026-08-19 · Keming Hu, Yingpei He

Centroid-Referenced Mahalanobis Matching (CRM): A Scalable, Representation-Based Framework for Causal Inference in Large Observational Studies

Matching for causal inference can be computationally expensive at scale and can silently change the target population when overlap is limited. We propose Centroid-Referenced Mahalanobis Matching (CRM), which replaces global pairwise search with stratified sampling in two reference coordinates: each unit's Mahalanobis distance from the...

💬 0 commentsarXiv:2608.18417v1PDF
0

Posted in stat.ML · 2026-08-18 · Haoshu Xu, Hongzhe Li

Inference and Uncertainty Quantification for Streaming $r$-PCA

We address two open questions in streaming PCA via Oja's algorithm: sharp operator-norm convergence for general rank under sub-Gaussian data, and distributional inference for the resulting subspace estimator. Existing convergence analyses, even in the rank-one case, either assume bounded data or leave non-vanishing remainder terms...

💬 0 commentsarXiv:2608.18374v1PDF
0

Posted in stat.ME · 2026-08-19 · Mark Cary, Charles Bokor

Regularised Iterative Generalised Least Squares with Optimal Selection of the Hyper-Parameter for Identifying Nonlinear Phenomenological Models

In some fields currently dominated by empirical approaches, such as state of health (SoH) prediction for lithium-ion batteries, phenomenological models motivated by quasi-physical thinking contain parameters to be estimated from experimental data. Often the structure of such models yields fully or partially confounded parameters,...

💬 0 commentsarXiv:2608.18742v1PDF
0

Posted in stat.AP · 2026-08-18 · Rhitankar Bandyopadhyay

Runs Above Expected (RAE) and Wicket Effect (WE): A Context-adjusted and Unified Impact Metric for Twenty20 Cricket

We develop a reproducible framework for evaluating individual batting and bowling performances in Twenty20 (T20) cricket on one interpretable scale of runs above expectation, built from two ball-level primitives. The first, Runs Above Expected (RAE), is the residual between the runs scored on a delivery and a contextual expectation of...

💬 0 commentsarXiv:2608.18020v1PDF
0

Posted in stat.ME · 2026-08-18 · K. Potter, K. R. Moran, R. Ulrich, D. C. Stenning, D. Bingham, L. Castro, G. Wilson, C. A. Maldonado

Scalable Heteroskedastic Gaussian Process Models for Large Inhomogeneous Datasets

We introduce Heteroskedastic Normalized Vecchia Gaussian Processes (HetNV), a scalable framework for Gaussian process regression with input-dependent observation noise. HetNV combines Vecchia likelihood approximations on normalized inputs with residual-based nonparametric variance estimation. The latent mean is estimated via a Vecchia...

💬 0 commentsarXiv:2608.18018v1PDF
0

Posted in stat.ME · 2026-08-18 · Xilin Mao, Bosen Cui, Yuhong Yang

Transporting Trial Evidence Under Posterior Drift and Possible Hidden Confounding

Randomized trials provide internally valid treatment-effect evidence, but trial participants may not represent the target population. In contrast, observational studies are often closer to the target population, but their treatment assignment may be affected by possible hidden confounding. We develop a robust posterior-drift framework...

💬 0 commentsarXiv:2608.17999v1PDF
0

Posted in stat.AP · 2026-08-18 · Ying Yao, Nan Zhang, Daniel J. Graham

Quantifying the Causal Operational Determinants of Service Reliability in Urban Rail Transit: Evidence from Panel Double/Debiased Machine Learning

Urban rail transit reliability is a critical measure of system performance, yet its causal determinants remain poorly quantified due to high-dimensional and interdependent influencing factors. This study investigates reliability patterns across 46 international metro operators between 1994 and 2024 using the CoMET benchmarking...

💬 0 commentsarXiv:2608.17901v1PDF
0

Posted in stat.ME · 2026-08-18 · Satabdi Saha, Christine B. Peterson

Graph-Adaptive Horseshoe for Compositional Regression

Compositional predictors, such as microbiome abundances, pose unique challenges in variable selection due to their unit-sum constraint and inherent dependencies. Existing approaches often rely on fixed association graphs derived from phylogenetic or ecological distances, which may not reflect outcome-relevant relationships. We propose...

💬 0 commentsarXiv:2608.17858v1PDF
0

Posted in stat.ML · 2026-08-18 · Kaifei Wang, Yinyu Ye, Han Zhong

Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits

Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-case regret $\mathcal{R}_{K,T}$ and instability $\mathcal S_{K,T}$, defined as the largest standard deviation of a terminal pull count, for $K$ arms and $T$...

💬 0 commentsarXiv:2608.17841v1PDF
0

Posted in stat.ME · 2026-08-18 · Markus Schepers, Werner Brannath, Esther Hoffmann, Julia Stingl, Irene Schmidtmann

Blinded sample size review for McNemar's test based on primary and surrogate endpoints

We develop blinded sample size re-estimation strategies for McNemar's test based on paired binary primary and secondary short-term surrogate endpoints. The development is motivated by a prospective randomized clinical trial on childhood glaucoma. A conditional power expression for McNemar's test given the primary endpoint at an...

💬 0 commentsarXiv:2608.17784v1PDF
0

Posted in stat.ME · 2026-08-18 · Tom Colemont, Brecht Evens, Tjonnie G. F. Li, Frederik De Ceuster

Modified Bryson-Frazier Smoothing and Hyperparameter Learning for Temporal Gaussian Process Regression

One-dimensional Gaussian processes with stationary, integrable kernel functions admit exact or arbitrarily accurate state-space representations, enabling linear-time inference through Kalman filtering and Rauch-Tung-Striebel (RTS) smoothing. However, the RTS smoother requires inversion of predicted state covariance matrices, which can...

💬 0 commentsarXiv:2608.17595v1PDF
0

Posted in stat.ML · 2026-08-18 · Huibo Xu, Shi Fu, Qixin Zhang, Dacheng Tao

Feature Priming in Online Linear Regression: Sparse-Regret Lower Bounds and a Tight Univariate Rate

In high-dimensional online prediction, the best predictor may depend on only a few features, so regret should scale with sparsity rather than the ambient dimension. Feature priming pursues this goal by estimating feature weights from past data and refitting a minimum-norm predictor on the rescaled design. Warmuth and Amid asked at...

💬 0 commentsarXiv:2608.17573v1PDF
0

Posted in stat.CO · 2026-08-18 · Zipei Nie, Guanyang Wang, Peng Zhang

The Snake Algorithm: A Rejection-Free Sampler for Binary Matrices with Fixed Margins

We study uniform sampling of binary matrices with fixed row and column sums, a recurring problem in ecological null models, Rasch-model testing, network analysis, and combinatorics. We propose the Snake algorithm, a rejection-free Markov chain Monte Carlo sampler that grows an alternating path until its first self-intersection and...

💬 0 commentsarXiv:2608.17531v1PDF
0

Posted in stat.ML · 2026-08-18 · Shuoguang Yang, Qiang Sun

Online Generalized Sparse Regression: How Does Overparametrization Help?

Regularized sparse regression has been extensively studied in the offline setting, but online formulation remains relatively under-explored. This gap stems from four key challenges: (i) the infeasibility of dynamically updating the regularization parameter in every online round, (ii) managing storage and memory complexity, (iii)...

💬 0 commentsarXiv:2608.17466v1PDF
0

Posted in stat.ML · 2026-08-18 · Kaiji Sekimoto, Muneki Yasuda

Nonlocal Transition Kernel for Efficient Learning of Restricted Boltzmann Machines

Learning restricted Boltzmann machines (RBMs) is computationally challenging because it requires expectations whose exact evaluation is generally intractable. The expectations are typically evaluated using a sampling approximation based on blocked Gibbs sampling (BGS), which is a local Markov chain Monte Carlo transition kernel....

💬 0 commentsarXiv:2608.17450v1PDF
0

Posted in stat.ME · 2026-08-18 · Eric Slud, Tim Trudell

SDR Variance Estimates in Small Domains

Successive Difference Replication (SDR) is a replication based method of variance estimation introduced by Fay and Train (1995) for estimators based on complex multistage surveys, especially those including a final systematic sampling stage. The method has been used for many years as the primary variance-estimation methodology in...

💬 0 commentsarXiv:2608.17353v1PDF
0

Posted in stat.ML · 2026-08-18 · Baishi Li, Kelvin J. L. Koa, Ke-Wei Huang

SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting

Modern probabilistic time-series forecasters often express uncertainty through forecast samples. While typically converted into nominal prediction regions using empirical quantiles, these model-implied sets lack formal coverage guarantees and frequently deviate from nominal targets under distribution shift. Existing multivariate...

💬 0 commentsarXiv:2608.17333v1PDF
0

Posted in stat.AP · 2026-08-18 · Luo Xiao, Wenyi Wang, Yumeng Zhang, Mike Lamonte, Andrea LaCroix, Chongzhi Di

A functional joint model with baseline functional covariates: linking sitting accumulation patterns to physical function and mortality among older women

In large-scale epidemiological studies, it is often of interest to investigate joint relationships between longitudinal and time-to-event outcomes with exposures that are trajectories or functions. Our motivation study is the Objective Physical Activity and Cardiovascular Health (OPACH) Study, which collected accelerometry-measured...

💬 0 commentsarXiv:2608.17278v1PDF
0

Posted in stat.ML · 2026-08-18 · Emma Ceccherini, Daniel Lawson, Anjulika Salhan

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

Categorising invoices into the correct General Ledger (GL) code underpins financial reporting and tax compliance. This is a skilled accounting judgement rather than a routine task: the correct category depends subtly on the nature of the purchasing business, the vendor and the invoice text. Whilst AI is increasingly being adopted...

💬 0 commentsarXiv:2608.18033v1PDF
0

Posted in stat.ML · 2026-08-17 · Shuai Huang, Zhe Qu, Zhaowei Hua, Guohao Shen, Rui Tang, Hongtu Zhu

Non-Crossing Deep Quantile Regression for Distributional Survival Prediction

In survival analysis the way covariates act on the risk of an event often differs between early and late failure times, yet hazard- and mean-based summaries collapse this variation into a single number. Quantile-based modeling instead describes the full conditional distribution on the original time scale, but existing censored-data...

💬 0 commentsarXiv:2608.16864v1PDF