Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 06:20:23 EST

0

Posted in stat.ML · 2026-08-24 · Kaj Nyström

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

We develop a finite-width geometric framework describing how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. Incompatibility among weight-generated covariance, gates, and backward sensitivities is quantified through three families of commutators: between gates and covariance,...

💬 0 commentsarXiv:2608.22910v1PDF
0

Posted in stat.ME · 2026-08-23 · David F. Anderson, Jingyi Ma

A general-purpose sensitivity method for multiple simultaneous parameter perturbations in stochastic reaction networks

Stochastic reaction networks are continuous-time Markov chain models for interacting populations, with applications in biochemistry, epidemiology, ecology, and related areas. We study finite-difference sensitivity estimation when a single estimator requires several nearby parameterized paths. Existing variance-reducing couplings are...

💬 0 commentsarXiv:2608.22627v1PDF
0

Posted in stat.ME · 2026-08-24 · Erin Craig, Yiling Huang, Snigdha Panigrahi

Interpretable AI with Local Distillation

Modern AI models such as tabular foundation models and gradient-boosted ensembles can outpredict classical methods, but provide little basis for reasoning about their predictions. High-stakes decisions call for models that are both accurate and interpretable as built. Local linear modeling offers a path forward: a smooth regression...

💬 0 commentsarXiv:2608.23538v1PDF
0

Posted in stat.ME · 2026-08-24 · Andrew C. Eggers, Zikai Li

Classification testing: A new framework for drawing qualitative conclusions from quantitative estimates

Social scientists rely on hypothesis testing to support their research conclusions, but the standard tests are designed for testing one hypothesis rather than adjudicating between rival possibilities. We develop a new framework, "classification testing", as an alternative. Instead of selecting one hypothesis to test, a researcher...

💬 0 commentsarXiv:2608.23315v1PDF
0

Posted in stat.ME · 2026-08-24 · Amadeo Grob, Maurizio Daniele, Johanna Ziegel

Sequentially valid inference for probabilistic inflation forecasts

Traditional statistical tests are poorly suited for the sequential evaluation of probabilistic forecast calibration. We address this limitation in macroeconomic forecasting by applying a new sequential testing method based on e-values. The e-value-based methodology enables anytime-valid inference. It allows practitioners to test...

💬 0 commentsarXiv:2608.23064v1PDF
0

Posted in stat.ME · 2026-08-24 · Lisa Leimenstoll, Melanie Schienle

Identification and Inference for Causal Effects in Extremes under General Conditions

Understanding the propagation of extreme events is important in many economic and environmental applications, yet most econometric methods for causal inference focus on average effects rather than tail behavior. This paper studies the identification of causal relations in extremes and derives resulting estimators and their asymptotic...

💬 0 commentsarXiv:2608.22957v1PDF
0

Posted in stat.ME · 2026-08-23 · Yuhao Deng, Haoyu Wei, Donglin Zeng, Rui Song, Xiao-Hua Zhou

Estimating Pathway Treatment Effects in the Presence of Intermediate Events with Multi-State Data

During clinical trials evaluating a drug's effect on a survival endpoint, intermediate events often occur in addition to the primary event. The treatment can exert its effect on the primary endpoint along multiple pathways through intermediate events. Assumptions for identifying mediation effects, such as sequential ignorability in...

💬 0 commentsarXiv:2608.22608v1PDF
0

Posted in stat.CO · 2026-08-21 · Francisco F. Queiroz, Rodrigo M. R. de Medeiros

Comprehensive Regression and Diagnostics for Non-Negative Data Using the BCSreg Package

Continuous positive data characterized by high skewness and heavy tails frequently arise in applied statistics. In other applications, these characteristics are accompanied by a point mass at zero, resulting in a non-negative response with a mixed discrete-continuous distribution. Standard regression models often fail to capture these...

💬 0 commentsarXiv:2608.21287v1PDF
0

Posted in stat.ML · 2026-08-21 · Adam Noonan

The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

Many machine-learning systems set a threshold at a quantile of a calibration set: conformal predictors that promise 90% coverage by drawing their cutoff at the calibration set's 90th percentile, abstention gates that decline to answer when a model's score falls below the calibration set's tenth percentile, safety filters that block...

💬 0 commentsarXiv:2608.21262v1PDF
0

Posted in stat.AP · 2026-08-21 · Chen Cheng, Vinh Ngoc Tran, Jiayuan Dong, Sarah Whitaker, Shannon Bergt, John Ziker, Valeriy Y. Ivanov, Xun Huan

Matching Urban Flood Sensor Placement to Monitoring Objectives Using Bayesian Optimal Experimental Design

Flood-monitoring sensors are often placed according to coverage, access, or expected inundation. However, the value of a measurement depends on the prediction or decision it is intended to inform. Using tRIBS-Urban simulations and a neural-network surrogate of the August 2014 metropolitan Detroit flood, we examine how this learning...

💬 0 commentsarXiv:2608.21182v1PDF
0

Posted in stat.AP · 2026-08-21 · Yili Hong, Xiaohong Gu

Statistical and Deep Learning Approaches for Predicting Degradation of Polymeric Materials in Photovoltaics

Polymeric materials are widely used in photovoltaic (PV) systems, making it essential to understand their service life to ensure reliable PV performance. The primary failure mechanism of polymeric materials in PV systems is photodegradation caused by ultraviolet (UV) radiation. Degradation modeling provides a framework for predicting...

💬 0 commentsarXiv:2608.21148v1PDF
0

Posted in stat.ME · 2026-08-21 · Piotr Fryzlewicz

LABS: Extending the scope of binary segmentation via a look-ahead device

Binary segmentation is widely used for multiple change-point detection because it is fast, simple to describe, and simple to implement. Its validity rests on the requirement that, at each recursive stage, the procedure identifies one of the true change-points when several are present in the current interval. This holds for detecting...

💬 0 commentsarXiv:2608.21122v1PDF
0

Posted in stat.ME · 2026-08-21 · Rok Spruk

Public Signals, Concealed Choices: Dynamic Measurement without Behavioral Identification

Members of collective institutions may leave public traces while their individual choices remain concealed. This paper separates a corpus-conditional public position from the behavioral rule linking that position to participation and secret choice. I measure the first with a dynamic ordinal state-space model and establish a...

💬 0 commentsarXiv:2608.21077v1PDF
0

Posted in stat.AP · 2026-08-21 · Yipeng Wei, Zahra Hoodbhoy, Emily R. Smith, Fang Jin, Muhammad Imran Nisar, Muhammad Farrukh Qazi, Christopher Mores, Victor Akelo, Caleb Sagam, Florence Aweyo, Charlotte Tawiah, Veronica Agyemang, Kwaku Poku Asante, Sam Newton, Santosh Joseph Benjamin, Anne George Cherian, Devakumar Devadhas, James A, Margaret P. Kasaro, Augustine Tunga, Sarmila Mazumder, Neeraj Sharma, Wilbroad Mutale, Mae Bridget Spelke, Qing Pan

Knowledge-guided Transfer Prediction In Underrepresented Populations: A GRU-D-Static Framework For Maternal And Neonatal Outcomes

Integrating summary-level scientific knowledge into neural network models provides a practical strategy for transferring prediction models trained on adequately sampled source cohorts to underrepresented target populations, where individual-level data in the target domain are often limited or unavailable. In this study, we propose...

💬 0 commentsarXiv:2608.21073v1PDF
0

Posted in stat.AP · 2026-08-21 · Charu Gupta, Gabriel Innocenzi, Christina Yap, Daniel Jackson, Fabio Rigat

Calibration of clinical trial sample size based on design utility

Clinical trial design relies on both statistical and clinical considerations for pre-specification of potentially practice-changing target treatment effects. As larger trials tend to be associated with high power and modest minimal detectable benefit, trial sample size is typically calibrated with reference to relevant precedents to...

💬 0 commentsarXiv:2608.20997v1PDF
0

Posted in stat.ME · 2026-08-21 · Zern Ke, Mingshi Cui, Feng Dai, Birol Emir, Javier Cabrera, Demissie Alemayehu

From Cumulative Weights to Marginal Density Ratios: Per-Protocol Estimation in Sequential Target Trial Emulation

Sequential target trial emulation evaluates eligibility at multiple baseline times to emulate a sequence of randomized trials using observational data. Estimating per-protocol effects in this setting is challenging because treatment deviations and loss to follow-up induce selection among individuals who remain observed and adherent...

💬 0 commentsarXiv:2608.20976v1PDF
0

Posted in stat.ME · 2026-08-21 · Eylul Fidan, Ufuk Beyaztas, Soutir Bandyopadhyay

Spatial function-on-function quantile regression

This paper introduces a novel penalized spatial function-on-function quantile regression framework for analyzing spatially indexed functional data, bridging a critical gap between spatial functional models and quantile regression. Our work makes three key contributions. First, we propose the first spatial function-on-function quantile...

💬 0 commentsarXiv:2608.20919v1PDF
0

Posted in stat.ME · 2026-08-21 · Samhita Pal, Jared D Huling

Heterogeneous Effects of Continuous Treatments via Conditional Modified Treatment Policies

For continuous treatments such as drug dose or ventilator intensity, a key clinically actionable question is whether a modest, patient-specific adjustment to the current dose would help or harm, rather than whether to treat at all. Standard estimands such as average or conditional dose-response functions require positivity across a...

💬 0 commentsarXiv:2608.20744v1PDF
0

Posted in stat.ME · 2026-08-21 · Stephany Lima de Oliveira, Frederico Machado Almeida

A modified score function for monotone likelihood in promotion time cure rate models

Survival models that incorporate a cure fraction provide a flexible framework for jointly modeling the cure and the survival distributions. However, when the data comprise a high proportion of censored observations or highly unbalanced binary covariates, maximum likelihood estimation may become unstable, leading to parameter estimates...

💬 0 commentsarXiv:2608.20641v1PDF
0

Posted in stat.ML · 2026-08-21 · Cholyeon Cho, Yuchen Wu

Minimax Optimality of Score-Entropy Discrete Diffusion

Discrete diffusion models have demonstrated strong performance across a range of datasets, including natural language data and graph-structured data. Among many variants, score-entropy discrete diffusion (SEDD) has achieved particularly strong empirical results. In SEDD, new samples are generated by iteratively evaluating a sequence...

💬 0 commentsarXiv:2608.20635v1PDF
0

Posted in stat.ME · 2026-08-20 · Soonhong Cho

Let Time Tell: Identification and Gaussian Process Estimation for Interrupted Time Series

We study causal inference in interrupted time series designs where a treatment affects every unit simultaneously, so that the contemporaneous controls used by difference-in-differences and synthetic control are unavailable and the counterfactual must be extrapolated from a unit's own pre-treatment history. We establish identification...

💬 0 commentsarXiv:2608.20610v1PDF
0

Posted in stat.ME · 2026-08-20 · Hyungjoon Kim, Andee Kaplan, Matthew D. Koslovsky

A Comprehensive Bayesian Approach to Entity Resolution for Data with Multiple Truths

In many applications, from government to ecology, integrating data from diverse and noisy sources is critical for downstream inference. However, a unique identifier to link records cleanly from the same entity may not exist. Entity resolution (also referred to as de-duplication or record linkage) merges such databases to identify...

💬 0 commentsarXiv:2608.20601v1PDF
0

Posted in stat.ME · 2026-08-20 · David McCoy, Yi Li

Targeted Deep Survival Contrasts: Valid Inference for Treatment-Specific Survival Benefit with Neural Networks

Neural survival models are increasingly asked to support counterfactual claims---how much a treatment would change survival in a population---rather than only prognostic risk scores. Answering such questions from observational data requires valid inference for treatment-specific survival contrasts under confounding and...

💬 0 commentsarXiv:2608.20598v1PDF
0

Posted in stat.ME · 2026-08-21 · Simon Rudkin, Wanling Rudkin

A Multiscale Ball Test for Conditional Mean Independence

Tests of conditional mean independence can lose power when departures are confined to a bounded part of a multivariate predictor space and the relevant spatial scale is unknown. We propose a Multiscale Ball Conditional Mean Independence (MBCMI) test that aggregates support-weighted local mean contrasts in an outcome variable across...

💬 0 commentsarXiv:2608.20727v1PDF
0

Posted in stat.ME · 2026-08-21 · Deepra Ghosh, Sanat K. Sarkar

Controlling the False Discovery Rate Control in Two-Sided Gaussian Mean Testing Under Arbitrary Dependence

The recent work of Sarkar and Zhang (2025) introduced Positive Tail Dependence Under the Null (PTDN) and developed Generalized Shifted Benjamini-Hochberg (BH) procedures for two-sided Gaussian $z$- and $t$-testing under known covariance structures. This paper develops further consequences of that framework. First, we derive explicit...

💬 0 commentsarXiv:2608.21267v1PDF