Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 7, 2026 — 03:14:52 EST

0

Posted in cs.DS · 2026-08-17 · Yunbum Kook, Santosh S. Vempala

Spectral Gaps of Hit-and-Run and Coordinate Hit-and-Run

For any convex body $\mathcal{K}\subset\mathbb{R}^{n}$ containing a unit ball, the spectral gap of Hit-and-Run is $Ω(1/(n^2 C_{\mathsf{PI}}))$, where $C_{\mathsf{PI}}$ is the Poincaré constant of the uniform distribution $π$ over $\mathcal{K}$. This implies that Hit-and-Run converges to a distribution within $χ^2$-divergence...

💬 0 commentsarXiv:2608.16878v1PDF
0

Posted in cs.SC · 2026-08-17 · Kejia Zhang, Youran Sun, Xinyu Ren, Chugang Yi, Haizhao Yang

AutoSR: Automatic Symbolic Regression by Searching Research States

We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searching persistent scientific investigations rather than isolated equations. Finite, noisy data often yield numerically competitive expressions that imply very different behavior outside the observed...

💬 0 commentsarXiv:2608.16876v1PDF
0

Posted in cs.LG · 2026-08-17 · Jiaming Li

An Analytical-Prior Framework for Data-Efficient Prediction of Sound-Reduction Frequencies in Rectangular Side-Branch Helmholtz Resonators

High-fidelity finite-element simulations can provide accurate numerical predictions for side-branch resonators, but large simulation datasets are expensive to generate and purely data-driven surrogates may become unreliable when simulation-labelled data are scarce. This study develops an analytical-prior learning framework that reuses...

💬 0 commentsarXiv:2608.16873v1PDF
0

Posted in cs.IR · 2026-08-17 · Mohsen Malmir, Houssam Nassif, Danish Nasir Shaikh, Taher Rahgooy, Murat Ali Bayir

Impression Share Prediction: An Offline Evaluation Task for Ranking Systems

Offline evaluation is a major gateway before online evaluation of ranking models in A/B testing. Standard offline metrics measure predictive accuracy, but are only a surrogate for downstream utility: a model can improve them while redistributing impressions across objective buckets in ways that degrade downstream utility. No offline...

💬 0 commentsarXiv:2608.16872v1PDF
0

Posted in cs.LG · 2026-08-17 · Serena Su, Yifan Wang, Senwei Liang

Data-Efficient and Interpretable Classification of Circulating Tumor Cell Phenotypes in Microfluidic Devices via Deep Learning

Accurate classification of circulating tumor cell (CTC) phenotypes can provide valuable information for assessing metastatic potential. Label free microfluidic devices provide a hydrodynamic obstacle course that transforms subtle biophysical characteristics of CTCs, including size and deformability, into distinct kinematic...

💬 0 commentsarXiv:2608.16870v1PDF
0

Posted in cs.GT · 2026-08-17 · Bailey Flanigan, Ismar Volic

The New Mathematics of Democracy

This article surveys emerging directions in the mathematics of democracy. It uses three case studies --- voting theory, participatory budgeting, and deliberative democracy --- to highlight how contemporary challenges motivate rigorous mathematical research that incorporates real-world data, institutional constraints, and...

💬 0 commentsarXiv:2608.16869v1PDF
0

Posted in cs.CL · 2026-08-17 · Benjamin Belay

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal state occurred. We test a bounded form of this idea in two controlled architectures: a...

💬 0 commentsarXiv:2608.16868v1PDF
0

Posted in cs.AI · 2026-08-14 · Taenyun Kim, Edyta Bogucka, Daniele Quercia

Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers

As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices...

💬 0 commentsarXiv:2608.14522v1PDF
0

Posted in cs.HC · 2026-08-14 · Kai Nylund, Michael Correll, Lace Padilla

Visualizing Uncertainty in Non-linear Projections with Ensembles

Widely used non-linear dimensionality reduction (NLDR) methods such as UMAP and t-SNE are stochastic--repeated runs on the same data can produce different low-dimensional projections. In this paper, we explore two problems related to projection variability: on some datasets clusters, structure, and outliers may change run-to-run, and...

💬 0 commentsarXiv:2608.14513v1PDF
0

Posted in cs.IT · 2026-08-14 · Yubo Zhang, Yiyao Liu, Xiaodong Wang

Learning-to-Transition for Large-scale and High-Order MIMO Detection

High-order multiple-input multiple-output (MIMO) detection requires efficient search over a large discrete symbol space while producing reliable soft information for channel decoding. This paper develops a learning-to-transition (L2T) framework that formulates MIMO detection as a stochastic sequence of complete-vector transitions. At...

💬 0 commentsarXiv:2608.14511v1PDF
0

Posted in cs.AI · 2026-08-14 · Zhelun Wu

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context. Combining interpretations rewards fixed arithmetic, comparability across instances, and the option to return...

💬 0 commentsarXiv:2608.14509v1PDF
0

Posted in cs.LG · 2026-08-14 · Pin-Yen Huang, Sachin Chhabra, Prasanth Sai Gouripeddi, Abhinav Kumar, Baoxin Li

RecipeNet: A Hierarchical Transformer for Recipe Data

Recipe data arises in domains such as materials synthesis, pharmaceutical formulation, and industrial manufacturing, where procedures are represented as ordered sequences of steps containing heterogeneous structured fields. Existing tabular learning methods typically flatten this structure into fixed-schema representations, limiting...

💬 0 commentsarXiv:2608.14505v1PDF
0

Posted in cs.SI · 2026-08-14 · Emily J Evans, Weihong Guo, Carlotta Domenicon

RegRole: Regularized Role Detection and Prediction in Temporal Dynamic Networks

This paper introduces a dynamic role discovery technique in temporal dynamic networks, utilizing temporally regularized Non-negative Matrix Factorization (NMF). Our technique differs from existing dynamic role analysis techniques by creating a consistent set of roles across all time periods, as well as a universal transition matrix...

💬 0 commentsarXiv:2608.14504v1PDF
0

Posted in cs.CR · 2026-08-14 · Bar Alon, Itai Dinur, Muthuramakrishnan Venkitasubramaniam

Lower Bounds on Black-Box Constructions of Pseudorandom Functions

In their seminal work, Goldreich, Goldwasser, and Micali [CRYPTO 1984] constructed a pseudorandom function (PRF) using a black-box access to a pseudorandom generator (PRG). When combined with Levin's domain extension technique, the GGM construction invokes the PRG $ω(\log n)$ times, where $n$ denotes the input length to the PRG. To...

💬 0 commentsarXiv:2608.14501v1PDF
0

Posted in cs.GT · 2026-08-14 · Zohar Barak, Inbal Talgam-Cohen

Ex-ante versus Ex-post: Egalitarian Facility Location Mechanism Design

We study the facility location mechanism design problem where $n$ strategic agents report locations in Euclidean space and the mechanism outputs a single facility location. Each agent's cost is its distance from the facility, and our objective is to minimize the egalitarian cost, i.e., the maximum agent cost, in a strategyproof way. ...

💬 0 commentsarXiv:2608.14499v1PDF
0

Posted in cs.LG · 2026-08-14 · Hanfeng Lu, Tianyu Feng, Suyi Li, Yuheng Zhao, Wei Gao, Shaopan Xiong, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Wei Wang

Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training

Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using task feedback, but current on-policy RL runtimes execute rollout, reference scoring, and actor training in strict serial phases. While...

💬 0 commentsarXiv:2608.14498v1PDF
0

Posted in cs.LG · 2026-08-14 · Hao Yan, Lisa Pilgram, Dan Liu, Linglong Kong, Fida Dankar, Khaled El Emam

Generating Benchmark Health Data Using a Tabular Diffusion Transformer

Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets. However, existing synthetic tabular data generation methods are largely restricted to single-input-table scenarios and struggle to effectively handle multiple heterogeneous tables with...

💬 0 commentsarXiv:2608.14496v1PDF
0

Posted in cs.IT · 2026-08-14 · Galen Reeves, Ramji Venkataramanan

Lossy Compression via Sparse Regression Codes: Generalized Construction and Finite-length Bounds

We study sparse regression codes (SPARCs) for lossy compression under simple greedy encoding rules, including both correlation-based and distance-based methods. We generalize the SPARC construction, and consider the class of \emph{additive orthogonal} regression codes, of which standard SPARCs are a special case. For this class of...

💬 0 commentsarXiv:2608.14494v1PDF
0

Posted in cs.LG · 2026-08-14 · Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He

Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled...

💬 0 commentsarXiv:2608.14430v1PDF
0

Posted in cs.LG · 2026-08-14 · Junichiro Niimi

Revisiting Energy-based Tabular Anomaly Detection: Energy and Reconstruction are Complementary

Tabular anomaly detection is dominated by classical density-proxy methods (Isolation Forest, OCSVM, LOF), reconstruction-based detectors (Autoencoders, VAEs), and modern non-parametric scorers (COPOD, ECOD, Deep SVDD), all of which approximate the inlier distribution only indirectly; explicit energy-based models are largely absent....

💬 0 commentsarXiv:2608.14186v1PDF
0

Posted in cs.LG · 2026-08-14 · Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu

Forecast Collapse in Time-Series Foundation Models

When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We...

💬 0 commentsarXiv:2608.14106v1PDF
0

Posted in cs.LG · 2026-08-14 · Joseph Sankoorikal Johny

When Does More Correct Data Hurt? Insertion-Stability and the Limits of Dimension-Based Theory

Adding data known to be correct ought to be safe. Not always. Larsen, Pabbaraju and Shetty model the failure with a monotone adversary, which reads an i.i.d. training sample and may append as many further examples as it likes, provided the target hypothesis labels them all. Mehrotra has since settled the cost, showing that for classes...

💬 0 commentsarXiv:2608.14020v1PDF
0

Posted in cs.LG · 2026-08-14 · Vincent Counathe, Ben Athiwaratkun, Christopher De Sa, Tianyi Zhang

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. However, QAT computes the loss and surrogate gradients using a lossy reconstruction of latent full-precision weights, while...

💬 0 commentsarXiv:2608.13966v1PDF
0

Posted in cs.CV · 2026-08-14 · Jing-Cheng Yang, Hao-Jung Wang, Jinhao Du, Yang Hu, Ming-shan Tsai, Jens Rittscher, Bin Li

Spatial Message Passing in Language Space for Pathology Image Interpretation

Multimodal Large Language Models (MLLMs) can generate pathological descriptions from histological images, but gigapixel Whole Slide Images (WSIs) exceed their visual context limits. The standard tiling workaround makes WSIs tractable yet severs the tissue neighborhoods that define tumor-stroma interfaces and morphology. We introduce...

💬 0 commentsarXiv:2608.14309v1PDF
0

Posted in cs.AI · 2026-08-14 · Alireza Kargarzadeh, Nariman Khaledian, Navid Parvini, Sid Ghatak, Arman Khaledian

Buy the Rumor, Sell the News: When Is News Priced In?

Two old market sayings hold that news is already priced in by the time it is published, and that the rumor is bought while the news is sold. Both place the price move associated with a piece of news before and at publication rather than after it. Whether the claims hold, for which kinds of news, and by how much are basic questions...

💬 0 commentsarXiv:2608.14014v1PDF