Qwen Councils

Statistics

arXiv preprints from January 1, 2026 through September 5, 2026 — 07:19:40 EST

0

Posted in stat.AP · 2026-08-21 · Niklas Heusch

A Synthetic Benchmark Dataset with Endogenous Marketing Spend for Validating Marketing Mix Models

Marketing Mix Models (MMMs) estimate the incremental sales effect of advertising from observational time series, yet they are rarely validated against ground truth, because ground truth is unobservable in real data. Synthetic data closes that gap in principle, but existing generators produce marketing spend exogenously - omitting the...

💬 0 commentsarXiv:2608.21130v1PDF
0

Posted in stat.AP · 2026-08-21 · Niklas Heusch

Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments

Marketing Mix Models (MMMs) are widely used for marketing measurement and budget allocation, but face fundamental identification challenges: due to endogenous marketing spend decisions, MMM estimation on observational time-series data cannot recover the true causal effects of marketing. On the other hand, geo-experiments provide...

💬 0 commentsarXiv:2608.21128v1PDF
0

Posted in stat.ME · 2026-08-20 · Montserrat Fuentes, Veronica B. Patterson

From Kriging to Spatial AI: Fifty Years of Spatial Statistics for Complex Dependent Data

Spatial statistics has grown from kriging for spatial prediction into a broad framework for learning from complex dependent data. This article traces that development from random fields and spectral methods to Bayesian hierarchical models and scalable computation. It then connects these foundations to Spatial AI, where graph learning...

💬 0 commentsarXiv:2608.20260v1PDF
0

Posted in stat.ML · 2026-08-20 · Junpeng Ren, Carlos Misael Madrid Padilla, Yanzhen Chen, Oscar Hernan Madrid Padilla

Transfer Learning in Nonparametric Regression with Deep ReLU Networks

This paper develops a general transfer learning framework for nonparametric regression with data consisting of multiple groups. Under the assumption that groups share a common structure along with group-specific deviations in additive form, the proposed method employs a two-stage offset learning procedure: the first stage pools data...

💬 0 commentsarXiv:2608.20255v1PDF
0

Posted in stat.ME · 2026-08-20 · Montserrat Fuentes, Veronica B. Patterson

A Bayesian Edge-Space Framework for Whole-Connectome Inference in Multisite Autism Neuroimaging

Autism spectrum disorder (ASD) is associated with heterogeneous alterations across distributed brain systems, creating challenges for whole-connectome inference. The difficulty arises not only from the large number of connections, but also from dependence among effects indexed by anatomically and functionally related region pairs. We...

💬 0 commentsarXiv:2608.20243v1PDF
0

Posted in stat.ME · 2026-08-20 · Manish Gupta, Dipanjan De

Multi-Method Causal Evidence Synthesis: Ranking Candidate Drivers by Convergent Cross-Method Evidence from Observational Data

Practitioners inferring causality from observational data usually rely on a single method and treat its output as causal truth. Recent tools select an optimal method for a dataset, and recent ensembles aggregate multiple causal-discovery algorithms into one graph, but little work pools evidence across different mathematical...

💬 0 commentsarXiv:2608.20187v1PDF
0

Posted in stat.ML · 2026-08-20 · Lohithsai Yadala Chanchu, Hany Abdulsamad, Christian A. Naesseth

Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo

We study inference-time control for text generation in discrete diffusion language models, where the goal is to steer sampling toward sequence-level rewards without retraining. Prior work in this domain has focused on particle-based methods such as best-of-$n$ sampling and bootstrap sequential Monte Carlo, which may suffer from...

💬 0 commentsarXiv:2608.20123v1PDF
0

Posted in stat.ME · 2026-08-20 · Yan Liu, Anita Koushik, Philippe Boileau, Cong Jiang, Miceline Mésidor, Claudia Waddingham, Denis Talbot, Mireille E. Schnitzer

Causal inference via propensity scores for case-control studies

Propensity score methods for causal inference are increasingly being used in cohort and experimental designs, but their development and uptake in outcome-dependent sampling schemes, such as case-control studies, remains limited. Case-control studies involve the sampling of individuals with and without an outcome of interest with the...

💬 0 commentsarXiv:2608.20080v1PDF
0

Posted in stat.AP · 2026-08-20 · Alejandro Rozo Posada, Maxime Fajgenblat, Christel Faes, James Colborn, Emanuele Giorgi, Baltazar Candrinho, Thomas Neyens

Integrating Temporal Disaggregation and Distributed Lag Nonlinear Models for Bayesian Spatio-Temporal Disease Mapping with High-Resolution Environmental Exposures

Environmental conditions are major drivers of malaria transmission, but epidemiological analyses are often constrained by temporal misalignment between health outcomes reported at coarse time scales and environmental exposures available at finer resolutions. Conventional approaches aggregate environmental data to match health...

💬 0 commentsarXiv:2608.20046v1PDF
0

Posted in stat.ME · 2026-08-20 · Rianne de Heide

Where Does the Union Bound Go? Best-Arm Identification and Strong FWER Control

In fixed-confidence best-arm identification, proofs often use a union bound across the competing arms. From a multiple-testing point of view this can look puzzling: if the best arm is unique, only one hypothesis of the form ``arm $i$ is best'' can be true. Why then should there be a Bonferroni-type factor of $K-1$? The answer is that...

💬 0 commentsarXiv:2608.19903v1PDF
0

Posted in stat.ME · 2026-08-20 · Mark Cary, Charles Bokor

A Repeated Measurements Approach to $SoH$ Battery Modelling of Cyclic Aged Data in a Laboratory Environment

This document describes the application of a first order linearised nonlinear repeated measurements approach to the analysis of battery cell ageing profiles generated under controlled conditions in a laboratory. The primary advantage of the model is it reflects the obvious structure in the data. Consequently, it is a two-component of...

💬 0 commentsarXiv:2608.19879v1PDF
0

Posted in stat.ME · 2026-08-20 · Johannes Hruza, Paweł Morzywołek, Jakob Zeitler, Samir Bhatt, Michael C Sachs

Partial Identification Learning with Categorical Treatments for Individualized Treatment Rules

We develop a partial identification learning framework for individualized treatment rules (ITRs) with categorical treatments, outcomes, and instrumental variables. Rather than relying on strong causal assumptions required for point identification, our framework leverages causal bounds to characterize the optimal treatment decision....

💬 0 commentsarXiv:2608.19853v1PDF
0

Posted in stat.ME · 2026-08-20 · Marin Šola, Xinwei Shen, Peter Bühlmann

Distributional Extrapolation for Interactions

Predicting combinatorial effects from limited-range observations is a fundamental challenge in many scientific domains, including drug discovery and hyperparameter optimization. We study combinatorial extrapolation, where training data consists of axis-aligned samples with only one active covariate, while test-time inputs involve...

💬 0 commentsarXiv:2608.19849v1PDF
0

Posted in stat.ME · 2026-08-20 · Xichen Guo, Feng Xie, Bingbing Tang, Yan Zeng, Zhang Hao, Zhi Geng, Ruichu Cai, Kun Zhang

Testing the Validity of Instrumental Variable Sets in Causal Additive Models with Non-Constant Effects

Instrumental variable (IV) methods are powerful for causal effect estimation with unmeasured confounding, but in practice researchers often face a set of candidate IVs whose validity is difficult to determine from observational data. This paper studies the problem of testing the validity of IV sets under Causal Additive Models with...

💬 0 commentsarXiv:2608.19771v1PDF
0

Posted in stat.CO · 2026-08-20 · Martin Tveten, Johannes Voll Kolstø, Per August Jarval Moen

skchange: Fast and Flexible Algorithms for Changepoint Detection

Skchange is an open-source Python library for detecting structural changes in time series. It implements modern change detection algorithms within a unified and extensible framework. The algorithms are modular and composable, and they include changepoint search methods based on both cost minimisation and statistical tests. Key...

💬 0 commentsarXiv:2608.19767v1PDF
0

Posted in stat.ME · 2026-08-20 · Zijun Gao, Kyounggeui Hong, Leyi Ma, Qianli Wu, Zachary Izzo, Ruishan Liu

Causal Survival Forests with Negative Controls

We study heterogeneous treatment-effect (HTE) estimation in observational survival studies commonly associated with both censored outcomes and unmeasured confounding. We integrate causal survival forests (CSF) with negative controls (NC) from proximal causal inference and introduce Negative Control Causal Survival Forests (NC-CSF), a...

💬 0 commentsarXiv:2608.19749v1PDF
0

Posted in stat.ME · 2026-08-20 · Laura M. Guzmán-Rincón, George R. E. Bradley, Joel Kandiah, Kyriakos Flouris, Pietro Liò, Paul J. Birrell, Alexander E. Zarebski, Daniela De Angelis

GENIE: Generative Neural Inference for Epidemics

The SARS-CoV-2 pandemic highlighted the ongoing risk infectious diseases pose to society and the value of reliable information on the likely future burden. When forecasting an epidemic at fine spatial resolution, traditionally used mechanistic compartmental model struggle to capture highly complex granular transmission dynamics,...

💬 0 commentsarXiv:2608.20253v1PDF
0

Posted in stat.ME · 2026-08-20 · Masahiro Tanaka

Curvature-Calibrated Quasi-Bayesian Updating for Moment-Restricted Models

Moment restrictions provide a flexible basis for quasi-Bayesian inference when a full likelihood is unavailable, but the weighting matrix in a quadratic moment criterion determines both the relative importance of the moments and the information scale of posterior updating. We propose curvature-calibrated quasi-Bayesian updating, which...

💬 0 commentsarXiv:2608.19634v1PDF
0

Posted in stat.ME · 2026-08-20 · Sayan Das, Debraj Das, Subhajit Dutta

Fast high-dimensional mean testing via logistic regression

We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common...

💬 0 commentsarXiv:2608.20286v1PDF
0

Posted in stat.ML · 2026-08-19 · Dalia Chakrabarty, Kangrui Wang, Chuqiao Zhang, Ye Liu

Learning Random Geometric Graphs Drawn in Probabilistic Metric Spaces

We present a new data-driven learning of a Random Geometric Graph (RGG) of a multivariate dataset, where the graph is drawn in a probabilistic metric space. This graph learning works for generic datasets, irrespective of the type of the observables; their probability distributions; or size of the data. We identify a metric of the...

💬 0 commentsarXiv:2608.19082v1PDF
0

Posted in stat.ML · 2026-08-19 · Yuga Iguchi, Paul Fearnhead

Diffusion Models for High-Dimensional Clustered Data: Intrinsic-Dimension Adaptivity via Bayesian Classification

The empirical success of diffusion models in generative modelling has motivated theoretical work, including quantitative error bounds and qualitative analyses that characterise the different phases of denoising. We bring these two areas together by studying the adaptivity of diffusion models to the structured geometry of multimodal...

💬 0 commentsarXiv:2608.19067v1PDF
0

Posted in stat.AP · 2026-08-19 · Sulagna Ghosh, Aaron Schein

Scalable Amortized Variational Inference for Non-Poisson Buy-'Til-You-Die Models

Despite the wide variety of existing Buy-`Til-You-Die (BTYD) models, nearly all rely upon the convenient assumption of transactions following a Poisson process. As modern customer bases grow larger and more diverse, a major gap in the marketing literature is BTYD models that can account for heterogeneity in timing patterns across...

💬 0 commentsarXiv:2608.19022v1PDF
0

Posted in stat.ME · 2026-08-19 · David Snider, Zhongyuan Lyu, Jian Kang, Yuqi Gu

Mixed Membership Model of Low-rank Matrices with Multimodal Extension

Matrix-valued observations arise in multiplex networks, neuroimaging, and other domains where population-level patterns are often low-rank and subjects may express several latent patterns simultaneously. Existing tensor PCA methods provide continuous subject scores but their loading matrices can be difficult to interpret as population...

💬 0 commentsarXiv:2608.18953v1PDF
0

Posted in stat.ME · 2026-08-19 · Lena Schemet, Andreas Groll, Sarah Friedrich-Welz

Model-based bootstrap inference for Cox models after Lasso selection

Inference after variable selection in Cox regression is difficult because simple Wald-type intervals after selection can have poor finite-sample conditional coverage. We study a model-based bootstrap for inference after Cox-Lasso variable selection. The Cox-Lasso is fitted once to the original data to select a set of variables, after...

💬 0 commentsarXiv:2608.18893v1PDF
0

Posted in stat.ML · 2026-08-19 · Matthias Mandl, Hanne Kekkonen

Sharper Regret Bounds for Time-Varying Gaussian Process Bandits with Constant Exploration

We study Bayesian optimization in a time-varying environment where the unknown reward function evolves according to a Gaussian process drift model. Existing GP-UCB analyses in this setting typically require the exploration parameter to grow with the horizon to maintain uniform confidence bounds. Using per-round local confidence...

💬 0 commentsarXiv:2608.18863v1PDF