Qwen Councils

All arXiv

arXiv preprints from January 1, 2026 through September 8, 2026 — 19:31:13 EST

0

Posted in eess.SY · 2026-08-27 · Loizos Hadjiloizou, Michael C. Welle, Hang Yin, Danica Kragic

Towards Safe Reinforcement Learning with Reduced Conservativeness: A Case Study on Drone Flight Control

Incorporating formal methods into reinforcement learning (RL) has the potential to result in the best of both worlds, combining the robustness of formal guarantees with the adaptability and learning capabilities of RL, though careful design is needed to balance safety and exploration. In this work, we propose a framework to mitigate...

💬 0 commentsarXiv:2608.26852v1PDF
0

Posted in stat.CO · 2026-08-27 · Mingcan Wang, Xiangjun Wang

Deep-Control BSDE: Layerwise Brownian-Weighted Regression for High-Dimensional Semilinear PDEs

High-dimensional semilinear parabolic partial differential equations arise in stochastic control, financial engineering, and uncertainty quantification, but classical spatial discretizations suffer from the curse of dimensionality. Motivated by Gaussian perturbation and conditional regression in denoising score matching, we propose...

💬 0 commentsarXiv:2608.27369v1PDF
0

Posted in stat.AP · 2026-08-27 · Manuele Leonelli

How exceptional was the Big Three era? Extremes and persistence in men's professional tennis

Three players won 66 of the 81 Grand Slam titles contested between 2003 and 2023, and their era is widely held to be the most dominant in the history of tennis. Assessing it means comparing players who never met, so that every comparison passes through the opponents each did face. We measure dominance by how far a player stands above...

💬 0 commentsarXiv:2608.27362v1PDF
0

Posted in stat.ME · 2026-08-27 · Subir Hait

Evidence, Calibration, and Stability: A Triadic Framework for Hypothesis Testing Under Model Uncertainty

Statistical tests are often asked to do too much. A single reported result is expected to describe what the observed data say, reassure readers about repeated-sampling behavior, and remain convincing when the working model is perturbed. Those tasks are connected, but they are not equivalent. Fisherian inductive inference and...

💬 0 commentsarXiv:2608.27320v1PDF
0

Posted in stat.ML · 2026-08-27 · Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$...

💬 0 commentsarXiv:2608.27313v1PDF
0

Posted in stat.ML · 2026-08-27 · Elena Badillo-Goicoechea, Fengfeng He

Recovering Expert Critic-Sourced Network Adjacency between Musical Artists from Acoustic Distributions: A Construct-Validity Approach

Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start regime, and intrinsic musical content, available for any recording. We argue that a third, largely untapped signal is both richer and more principled: critical adjacency, the pairwise relation established when an expert critic...

💬 0 commentsarXiv:2608.27291v1PDF
0

Posted in stat.ME · 2026-08-27 · Jack M. Wolf, Joseph S. Koopmeiners, David M. Vock

Combining covariate adjustment with information from secondary endpoints to improve precision in randomized trials

Background/Aims: Adjustment for prognostic baseline covariates can improve precision in randomized trials. Previous work has shown that jointly modeling primary and secondary endpoints can yield additional precision by borrowing information across endpoints. We investigated whether these approaches can be combined to achieve...

💬 0 commentsarXiv:2608.27289v1PDF
0

Posted in cs.AI · 2026-08-27 · Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller

Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation

Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through \textit{de novo} generation of product molecules or through heuristic graph edits that operate directly on molecular topology. We introduce MAELLE (\textbf{M}ech\textbf{A}nistic \textbf{E}dit...

💬 0 commentsarXiv:2608.27429v1PDF
0

Posted in cs.CL · 2026-08-27 · Vésteinn Snæbjarnarson, Samuel Kiegeland, Manuel de Prada Corral, Ryan Cotterell, Tim Vieira

Stochastic Estimation of Transduced Language Models

Transduced language models (TLMs) compose a pretrained \emph{source} language model with a functional finite-state transducer to induce a language model over \emph{target} strings. Computing the probability of a target prefix under a TLM amounts to summing the source-model probabilities of all source strings that the transducer maps...

💬 0 commentsarXiv:2608.27428v1PDF
0

Posted in cs.SE · 2026-08-27 · Yisen Xi

Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both cheaply. We present Persona-Execution Separation (PES): persona and execution reside in...

💬 0 commentsarXiv:2608.27427v1PDF
0

Posted in cs.CR · 2026-08-27 · Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evaluate ModelScan, ModelAudit, and Fickling using a controlled, artifact-backed benchmark on a...

💬 0 commentsarXiv:2608.27424v1PDF
0

Posted in cs.IR · 2026-08-27 · Edgar Chavez

misi: a Metric Inverted Sample Index

We present misi, an inverted index for approximate nearest-neighbor search over general metric spaces whose vocabulary is a random sample of the database, of size proportional to $n$. Each object is represented by its $k_b$ nearest sample points, found by a pluggable inner index over the sample; queries are answered by an idf-weighted...

💬 0 commentsarXiv:2608.27422v1PDF
0

Posted in cs.AI · 2026-08-27 · Kevin Zhu, Ryan Zhang, Baraa Abed, Tilendra Choudhary, Malvern Madondo, Mehak Arora, Yixuan Yang, Alasdair Gent, Aditya Nagori, Omer T. Inan, Krista L. Haines, Patrick Georgoff, Suresh M. Agarwal, Vijay Krishnamoorthy, Tetsu Ohnuma, Mihai V. Podgoreanu, Michael R. Pinsky, Gilles Clermont, Craig M. Coopersmith, Craig S. Jabaley, Rishikesan Kamaleswaran

Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study

Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsely discretized and calibrated to a cohort that no longer reflects contemporary critical care. No alternative learned directly from patient trajectories is in routine use. We conducted a retrospective two-cohort study on...

💬 0 commentsarXiv:2608.27421v1PDF
0

Posted in cs.CL · 2026-08-27 · Xingyu Shen, Huishuai Zhang, Peng Li, Yinchun Wang, Dongyan Zhao

Boosting LLM Exploration via Weak-Model Guidance in RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric...

💬 0 commentsarXiv:2608.27420v1PDF
0

Posted in cs.GT · 2026-08-27 · Léonard Brice, F. Thomas Bruss, Anirban Majumdar, Jean-François Raskin

Algorithms for Robbins' Problem using Markov Decision Processes

In this paper, we consider Robbins' problem, which is a full information variant of the well-known secretary selection problem. In this version of the problem, the goal is to minimize the expected rank of the selected candidate among $n$ that are interviewed sequentially, and a decision to select or not the $m^{th}$ candidate needs to...

💬 0 commentsarXiv:2608.27419v1PDF
0

Posted in cs.CV · 2026-08-27 · Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for visual...

💬 0 commentsarXiv:2608.27417v1PDF
0

Posted in math.CO · 2026-08-27 · Hermann Wilhelm

Refutation of the Non-Cancelling-Intersections Conjecture

The Non-Cancelling Intersections (NCI) conjecture of Amarilli, Monet and Suciu [arXiv:2401.16210] states that the union of a finite family of sets can always be built from its algebraically non-cancelling intersections using only disjoint unions and subset complements. In Wilhelm [arXiv:2608.19414] the conjecture was shown to fail...

💬 0 commentsarXiv:2608.27416v1PDF
0

Posted in cs.IR · 2026-08-27 · Maksim Utushkin, Andrei Ovsiannikov, Alexander D'yakonov

Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling

Friend recommendation is inherently graph-structured: the relevance of a potential connection depends on multi-hop social context rather than user attributes alone. However, deploying message-passing GNNs on a production-scale social graph with hundreds of millions of users and tens of billions of edges requires addressing numerous...

💬 0 commentsarXiv:2608.27413v1PDF
0

Posted in cs.CL · 2026-08-27 · Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, Yanghua Xiao

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artefacts they reuse: Merge combines expert task vectors, Mix...

💬 0 commentsarXiv:2608.27409v1PDF
0

Posted in cs.CV · 2026-08-27 · Agniv Chatterjee, Georgios Pavlakos

Reconstructing Humans and Objects in Interaction using Large Reconstruction Models

Estimation of Human-Object Interactions in 3D (3D HOI) is a fundamental problem in 3D computer vision with applications in AR/VR, robotics, and embodied AI. However, reconstructing these interactions in 3D remains challenging due to depth ambiguities, occlusions, and object shape variability. Existing approaches are primarily...

💬 0 commentsarXiv:2608.27407v1PDF
0

Posted in cs.RO · 2026-08-27 · Kechen Liu, Ola Shorinwa

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned...

💬 0 commentsarXiv:2608.27406v1PDF
0

Posted in cs.CL · 2026-08-27 · Orion Reblitz-Richardson

How Language Models Organize and Structure Moral Knowledge

How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language...

💬 0 commentsarXiv:2608.27402v1PDF
0

Posted in math.PR · 2026-08-27 · Shuoqing Deng, Daxin Huang

Distribution-constrained optimal multiple stopping: the Root-type solution

We consider the distribution-constrained optimal stopping problem introduced by Bayraktar and Miller (Mathematical Finance, 2019) and Beiglbock et al. (PTRF, 2018). Motivated by the multi-marginal Skorokhod embedding problems (SEP) and applications in mathematical finance, we generalize (a class of) its solution to the multi-marginal...

💬 0 commentsarXiv:2608.27374v1PDF
0

Posted in q-fin.CP · 2026-08-27 · Nneka Umeorah, Tolulope Fadina

A Temporal Multiplex Graph Neural Network for Systemic Risk Transmission in Global Banking

This paper develops a unified framework for assessing systemic risk and identifying contagion channels in the global banking system using a Temporal Heterogeneous Multiplex Graph Neural Network. We construct a harmonised quarterly panel combining bank fundamentals, CDS spreads, and macroeconomic indicators, and represent these data as...

💬 0 commentsarXiv:2608.27295v1PDF
0

Posted in q-fin.RM · 2026-08-27 · Aleksandar Arandjelovic, Pavel V. Shevchenko, George Tzougas

On the approximation of posterior laws in compound loss models by conditional Wasserstein GANs

Bayesian inference in compound loss models must often be repeated across policies, market scenarios, and prior specifications. Outside conjugate cases, this may require repeated numerical integration or Markov chain Monte Carlo (MCMC). We formulate this problem as amortized posterior approximation and construct a conditional...

💬 0 commentsarXiv:2608.27229v1PDF