Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 7, 2026 — 09:22:00 EST

0

Posted in cs.CV · 2026-07-28 · Yuan Yin, Elias Ramzi, Marc Lafon, Valentin Charraut, Victor Bares, Yihong Xu, Éloi Zablocki, Alexandre Boulch, Thibault Buhet, Andrei Bursuc, Matthieu Cord

Pictura: Perspective-View Self-Play at Scale for Driving

Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed...

💬 0 commentsarXiv:2607.26005v1PDF
0

Posted in cs.CV · 2026-07-28 · Neta Shaul, Chao Liu, Arash Vahdat, Julius Berner

Parallel Decoding Distillation for Fast Image and Video Generation

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality...

💬 0 commentsarXiv:2607.26004v1PDF
0

Posted in cs.LG · 2026-07-28 · Wenzhi Zhong, Edward Milsom, Michael Murray

Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a "small" perturbation is inherently geometry-dependent: while existing SAM variants have explored a wide range of choices, a clear perspective on which geometries are most...

💬 0 commentsarXiv:2607.26001v1PDF
0

Posted in cs.CL · 2026-07-27 · Zhen Huang, Yikun Wang, Shijie Xia, Pengfei Liu

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each example. We propose DataOrchestra, a framework that...

💬 0 commentsarXiv:2607.24717v1PDF
0

Posted in cs.AI · 2026-07-27 · Ali Ansari, Yasmin Mohammadi, Farnoush Nili, Parsa Esmaeilkhani, Longin Jan Latecki, Eduard Dragut

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the first large-scale benchmark for structured understanding of ER diagrams, comprising...

💬 0 commentsarXiv:2607.24707v1PDF
0

Posted in cs.CV · 2026-07-27 · Hang Xing, Guangjun Liu, Yan Xia, Xueming Ding

SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations

Few-shot segmentation (FSS) commonly assumes clean pixel-level support masks, yet practical support supervision often uses boxes, scribbles, coarse masks, or pseudo-masks. These weak annotations may include texture-similar distractors and background context alongside the target, contaminating class prototypes or visual prompts before...

💬 0 commentsarXiv:2607.24706v1PDF
0

Posted in cs.CV · 2026-07-27 · Anika Knupfer, Maximilian Lindholz, Johanna Paula Müller, Jordina Aviles Verdera, Smiti Tripathy, Susanne Schulz-Heise, Jana Hutter

Panda: Unsupervised Pelvic Anomaly Detection for Real-Time MR Imaging

Female pelvic diseases remain an under researched area characterized by often delayed diagnosis. While pelvic MRI offers superior soft-tissue contrast for diagnosis and image-guided procedures, real-time anomaly detection remains challenging due to physiological motion, tissue deformation, and instrument artifacts. Existing supervised...

💬 0 commentsarXiv:2607.24703v1PDF
0

Posted in cs.CV · 2026-07-27 · Andong Lu, Ziyi Zha, Jiandong Jin, Shihao Li, Chenglong Li, Jin Tang, Bin Luo

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be unsatisfactory. In addition,...

💬 0 commentsarXiv:2607.24701v1PDF
0

Posted in cs.SI · 2026-07-27 · Dini Wang, Ho-Chun Herbert Chang

Modest Algorithmic Mediation can Maximize Topical Diversity in Hybrid Human-AI Systems

In the artificial intelligence (AI) era, the rise of algorithmic feeds has fundamentally transformed information diffusion on social media. While early platforms organized visibility through explicit social networks, contemporary systems mediate exposure through intelligent recommender algorithms that personalize attention. This paper...

💬 0 commentsarXiv:2607.24698v1PDF
0

Posted in cs.NI · 2026-07-27 · Jhonatan Tavori, Gur-Eyal Sela, Ion Stoica, Gil Zussman

Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines

Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on time and combined with the fast path predictions. Across several application...

💬 0 commentsarXiv:2607.24692v1PDF
0

Posted in cs.DB · 2026-07-27 · Zeyu Zhang, Xue Li, Iacer Calixto, Paul Groth, Sebastian Schelter

Beyond Scale and Generation: Understanding Language Model-based Entity Matching

Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model variant(reflecting different pretraining...

💬 0 commentsarXiv:2607.24688v1PDF
0

Posted in cs.CV · 2026-07-27 · Francisco Mena, Dino Ienco, Roberto Interdonato, Cassio F. Dantas, Simon Besnard

Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification

Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to inconsistent modality availability between training and inference times. To handle...

💬 0 commentsarXiv:2607.24683v1PDF
0

Posted in cs.AI · 2026-07-27 · Stefan G. Creadore

Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We developed Plato-Bio, a biology-routed extension of the open Plato/Denario architecture that couples explicit workflow states with provenance records, citation...

💬 0 commentsarXiv:2607.23975v1PDF
0

Posted in cs.AI · 2026-07-27 · Saurabh Ranjan, Konstantina Sokratous, Brian Odegaard

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two...

💬 0 commentsarXiv:2607.23927v1PDF
0

Posted in cs.HC · 2026-07-26 · Ruyi Cao, Lily M. Turkstra, Adyah Rastogi, Michael Beyeler

Evaluating Closed-Loop EEG Feedback for Simulated Prosthetic Vision in Immersive VR: A Sham-Controlled Feasibility Study

Visual prostheses require users to interpret sparse and distorted artificial percepts through active visual search. We developed an EEG-guided neuroadaptive training platform for simulated prosthetic vision in immersive virtual reality and evaluated its feasibility in a sham-controlled object-localization task. Twenty-two sighted...

💬 0 commentsarXiv:2607.23889v1PDF
0

Posted in cs.LG · 2026-07-26 · Shuyu Chen, Chen Zhu, Ye Zhang, Yang Li, Qiqi Xie, Haohan Wang

SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing

Identifying therapeutic target genes from single-cell RNA sequencing (scRNA-seq) data remains a fundamental challenge in translational biology. Unlike bulk assays, scRNA-seq captures heterogeneous cellular states and rare subpopulations, but this same heterogeneity makes target discovery highly sensitive to analytical choices...

💬 0 commentsarXiv:2607.23821v1PDF
0

Posted in cs.LG · 2026-07-26 · Hengyuan Cao, Shizhuo Cheng, Mingxuan Liu, Weicheng Huang, Yunhong Lu, Chenxi Cai, Yan Zhang, Min Zhang

Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling

The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However, existing approaches are predominantly implemented under a single-target, single-state assumption,...

💬 0 commentsarXiv:2607.23518v1PDF
0

Posted in cs.RO · 2026-07-27 · Yifan Ye, Yankai Fu, Yaoxu Lv, Bohan Hou, Jun Cen, Lingdong Kong, Duo Zheng, Tianxing Chen, Jiaming Liu, Ziang Cao, Yunfan Lou, Wei Chow, Xian Sun, Yingshuo Wang, Kuangzhi Ge, Xiaowei Chi, Xidong Zhang, Zhibo Pang, Yiwu Zhong, Sirui Han, Zhihe Lu, Weihao Yuan, Qifeng Chen, Michael Yu Wang, Yao Mu, Ziwei Liu, Jianfei Yang, Ping Luo, Shanghang Zhang

Data Pyramid for Embodied Manipulation

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data...

💬 0 commentsarXiv:2607.24744v1PDF
0

Posted in cs.CV · 2026-07-27 · Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and...

💬 0 commentsarXiv:2607.24743v1PDF
0

Posted in cs.DC · 2026-07-27 · Xinyang Wen

Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal Transport

Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed Sinkhorn processes frames sequentially and synchronizes after every iteration. We present TemporalSinkhorn, a parallel-in-time executor that batches future candidates and their...

💬 0 commentsarXiv:2607.24741v1PDF
0

Posted in cs.HC · 2026-07-27 · Helen Weixu Chen, Victoria Sakhnini, Lesley Istead

Make or Take: How Students Navigate Self-Created and Instructor-Provided Cheat Sheets

The use of cheat sheets in exams is often framed as a way to reduce cognitive load and support student performance. However, little is known about how students choose between self-created and instructor-provided cheat sheets, or how these choices relate to their broader approaches to exam preparation. We conducted a longitudinal study...

💬 0 commentsarXiv:2607.24736v1PDF
0

Posted in cs.DS · 2026-07-27 · Jon Kleinberg, Amin Saberi, Xizhi Tan, Grigoris Velegkas

Learning Distributions from Multiple Data Providers

Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples. The goal is to learn an unknown distribution $p$ on a finite domain $[n]$. The learner is given a fixed family of queryable sets $\mathscr{S} \subseteq 2^{[n]}$, and each...

💬 0 commentsarXiv:2607.24732v1PDF
0

Posted in cs.CV · 2026-07-27 · Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the...

💬 0 commentsarXiv:2607.24731v1PDF
0

Posted in cs.CV · 2026-07-27 · Krithi Shailya, Ananya Lakshmi Ravi, Venkatanathan K. V., Sowmya S. Sundaram, Gokul S. Krishnan, Aditi Anand, Balaraman Ravindran

KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In clinical workflows, chest X-ray classifiers are increasingly paired with Vision-Language Models (VLMs) to generate natural-language explanations. However, these systems add linguistic...

💬 0 commentsarXiv:2607.24730v1PDF
0

Posted in cs.CV · 2026-07-27 · Huy Huynh, Jingwei Ma, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz

MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale

We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of consumer-grade microscope close-ups, MicroZoom synthesizes a seamless, gigapixel-resolution image grounded in the material character of the real references, enabling exploratory...

💬 0 commentsarXiv:2607.24729v1PDF