Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 8, 2026 — 15:31:56 EST

0

Posted in cs.CV · 2026-01-20 · Matthew Gwilliam, Xiao Wang, Xuefeng Hu, Zhenheng Yang

Implicit Neural Representation Facilitates Unified Universal Vision Encoding

Models for image representation learning are typically designed for either recognition or generation. Various forms of contrastive learning help models learn to convert images to embeddings that are useful for classification, detection, and segmentation. On the other hand, models can be trained to reconstruct images with pixel-wise,...

💬 0 commentsarXiv:2601.14256v1PDF
0

Posted in cs.CV · 2026-01-20 · Sangbeom Lim, Seoung Wug Oh, Jiahui Huang, Heeji Yoon, Seungryong Kim, Joon-Young Lee

VideoMaMa: Mask-Guided Video Matting via Generative Prior

Generalizing video matting models to real-world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask-to-Matte Model (VideoMaMa) that converts coarse segmentation masks into pixel accurate alpha mattes, by leveraging pretrained video diffusion models. VideoMaMa demonstrates...

💬 0 commentsarXiv:2601.14255v1PDF
0

Posted in cs.CV · 2026-01-20 · Hongyuan Chen, Xingyu Chen, Youjia Zhang, Zexiang Xu, Anpei Chen

Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis

We present Motion 3-to-4, a feed-forward framework for synthesising high-quality 4D dynamic objects from a single monocular video and an optional 3D reference mesh. While recent advances have significantly improved 2D, video, and 3D content generation, 4D synthesis remains difficult due to limited training data and the inherent...

💬 0 commentsarXiv:2601.14253v1PDF
0

Posted in cs.IT · 2026-01-20 · Tristan Simas

Semantic Identity Compression: Zero-Error Laws, Rate-Distortion, and Neurosymbolic Necessity

Symbolic systems operate over precise identities: variables denote specific objects, pointers target precise memory locations, and database keys refer to singular records. Neural embeddings generalize by compressing away semantic detail, but this compression creates collision ambiguity: multiple distinct entities can share the same...

💬 0 commentsarXiv:2601.14252v6PDF
0

Posted in cs.CV · 2026-01-20 · Said Taghadouini, Adrien Cavaillès, Baptiste Aubertin

LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR

We present LightOnOCR-2-1B, a 1B-parameter end-to-end multilingual vision--language model that converts document images (e.g., PDFs) into clean, naturally ordered text without brittle OCR pipelines. Trained on a large-scale, high-quality distillation mix with strong coverage of scans, French documents, and scientific PDFs,...

💬 0 commentsarXiv:2601.14251v2PDF
0

Posted in cs.CV · 2026-01-20 · Pengze Zhang, Yanze Wu, Mengtian Li, Xu Bai, Songtao Zhao, Fulong Ye, Chong Mou, Xinghui Li, Zhuowei Chen, Qian He, Mingyuan Gao

OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer

Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the rich spatio-temporal information inherent in videos, thereby limiting flexibility and...

💬 0 commentsarXiv:2601.14250v1PDF
0

Posted in cs.CL · 2026-01-20 · Yuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment

Long chain-of-thought (CoT) trajectories provide rich supervision signals for distilling reasoning from teacher to student LLMs. However, both prior work and our experiments show that trajectories from stronger teachers do not necessarily yield better students, highlighting the importance of data-student suitability in distillation....

💬 0 commentsarXiv:2601.14249v5PDF
0

Posted in cs.CV · 2026-01-20 · Zeyuan Chen, Kai Zhang, Zhuowen Tu, Yuanjun Xiong

Soft Tail-dropping for Adaptive Visual Tokenization

We present Soft Tail-dropping Adaptive Tokenizer (STAT), a 1D discrete visual tokenizer that adaptively chooses the number of output tokens per image according to its structural complexity and level of detail. STAT encodes an image into a sequence of discrete codes together with per-token keep probabilities. Beyond standard...

💬 0 commentsarXiv:2601.14246v1PDF
0

Posted in cs.IR · 2026-01-20 · Zhongyu Yang, Wei Pang, Yingfang Yuan

XR: Cross-Modal Agents for Composed Image Retrieval

Retrieval is being redefined by agentic AI, demanding multimodal reasoning beyond conventional similarity-based paradigms. Composed Image Retrieval (CIR) exemplifies this shift as each query combines a reference image with textual modifications, requiring compositional understanding across modalities. While embedding-based CIR methods...

💬 0 commentsarXiv:2601.14245v2PDF
0

Posted in cs.LG · 2026-01-20 · Haocheng Xi, Charlie Ruan, Peiyuan Liao, Yujun Lin, Han Cai, Yilong Zhao, Shuo Yang, Kurt Keutzer, Song Han, Ligeng Zhu

Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow

Reinforcement learning (RL) is essential for enhancing the complex reasoning capabilities of large language models (LLMs). However, existing RL training pipelines are computationally inefficient and resource-intensive, with the rollout phase accounting for over 70% of total training time. Quantized RL training, particularly using FP8...

💬 0 commentsarXiv:2601.14243v2PDF
0

Posted in cs.CL · 2026-01-20 · Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, Julien Benchek, David Ostrofsky, Anirudh Ravichandran, Debnil Sur, Neel Venugopal, Alannah Hsia, Isaac Robinson, Calix Huang, Olivia Varones, Daniyal Khan, Michael Haines, Austin Bridges, Jesse Boyle, Koby Twist, Zach Richards, Chirag Mahapatra, Brendan Foody, Osvald Nitski

APEX-Agents

We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers. APEX-Agents requires agents to navigate realistic work environments with files and tools....

💬 0 commentsarXiv:2601.14242v3PDF
0

Posted in cs.SD · 2026-01-20 · Aafiya Hussain, Gaurav Srivastava, Alvi Ishmam, Zaber Hakim, Chris Thomas

SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models

Multimodal foundation models that integrate audio, vision, and language achieve strong performance on reasoning and generation tasks, yet their robustness to adversarial manipulation remains poorly understood. We study a realistic and underexplored threat model: untargeted, audio-only adversarial attacks on trimodal...

💬 0 commentsarXiv:2601.16231v1PDF
0

Posted in cs.LG · 2026-01-20 · Shaurya Mathur, Shreyas Bellary Manjunath, Nitin Kulkarni, Alina Vereshchaka

Spatiotemporal Wildfire Prediction and Reinforcement Learning for Helitack Suppression

Wildfires are growing in frequency and intensity, devastating ecosystems and communities while causing billions of dollars in suppression costs and economic damage annually in the U.S. Traditional wildfire management is mostly reactive, addressing fires only after they are detected. We introduce \textit{FireCastRL}, a proactive...

💬 0 commentsarXiv:2601.14238v1PDF
0

Posted in cs.IT · 2026-01-20 · Giulio Pech, Mert Gökduman, Hanwen Yao, Henry D. Pfister

Stabilizer-Assisted Inactivation Decoding of Quantum Error-Correcting Codes with Erasures

In this work, we develop a reduced complexity maximum likelihood (ML) decoder for quantum low-density parity-check (QLDPC) codes over erasures. Our decoder combines classical inactivation decoding, which integrates peeling with symbolic guessing, with a new dual peeling procedure. In the dual peeling stage, we perform row operations...

💬 0 commentsarXiv:2601.14236v1PDF
0

Posted in cs.LG · 2026-01-20 · Qiyang Li, Sergey Levine

Q-learning with Adjoint Matching

We propose Q-learning with Adjoint Matching (QAM), a novel TD-based reinforcement learning (RL) algorithm that tackles a long-standing challenge in continuous-action RL: efficient optimization of an expressive diffusion or flow-matching policy with respect to a parameterized Q-function. Effective optimization requires exploiting the...

💬 0 commentsarXiv:2601.14234v4PDF
0

Posted in cs.LG · 2026-01-20 · Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov

KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are unchanged, but existing benchmarks entangle multiple sources of shift and hinder systematic analysis. We introduce KAGE-Env, a JAX-native 2D platformer that factorizes the observation process into...

💬 0 commentsarXiv:2601.14232v2PDF
0

Posted in cs.CL · 2026-01-20 · Yiyang Wang, Yiqiao Jin, Alex Cabral, Josiah Hester

MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems

Multi-agent systems (MAS) are emerging as promising socio-collaborative companions for emotional and cognitive support. However, existing systems frequently suffer from persona collapse, where agents revert to generic, homogenized assistant behaviors, and social sycophancy, where agents produce redundant, non-constructive dialogue. We...

💬 0 commentsarXiv:2601.14230v2PDF
0

Posted in cs.LG · 2026-01-20 · Punit Kumar, Vaibhav Saran, Divyesh Patel, Nitin Kulkarni, Alina Vereshchaka

Attention-Based Offline Reinforcement Learning and Clustering for Interpretable Sepsis Treatment

Sepsis remains one of the leading causes of mortality in intensive care units, where timely and accurate treatment decisions can significantly impact patient outcomes. In this work, we propose an interpretable decision support framework. Our system integrates four core components: (1) a clustering-based stratification module that...

💬 0 commentsarXiv:2601.14228v1PDF
0

Posted in cs.SD · 2026-01-20 · Theodore Aptekarev, Vladimir Sokolovsky, Gregory Furman

Transformer Architectures for Respiratory Sound Analysis and Multimodal Diagnosis

Respiratory sound analysis is a crucial tool for screening asthma and other pulmonary pathologies, yet traditional auscultation remains subjective and experience-dependent. Our prior research established a CNN baseline using DenseNet201, which demonstrated high sensitivity in classifying respiratory sounds. In this work, we (i) adapt...

💬 0 commentsarXiv:2601.14227v1PDF
0

Posted in cs.IR · 2026-01-20 · Sahel Sharifymoghaddam, Jimmy Lin

Rerank Before You Reason: Analyzing Reranking Tradeoffs through Effective Token Cost in Deep Search Agents

Deep research agents rely on iterative retrieval and reasoning to answer complex queries, but scaling test-time computation raises significant efficiency concerns. We study how to allocate reasoning budget in deep search pipelines, focusing on the role of listwise reranking. Using the BrowseComp-Plus benchmark, we analyze tradeoffs...

💬 0 commentsarXiv:2601.14224v2PDF
0

Posted in cs.SI · 2026-01-20 · Mohak Goyal, Lodewijk Gelauff, Naman Gupta, Ashish Goel, Kamesh Munagala

Beyond Polarization: Opinion Mixing and Social Influence in Deliberation

Deliberative processes are often discussed as increasing or decreasing polarization. This approach misses a different, and arguably more diagnostic, dimension of opinion change: whether deliberation reshuffles who agrees with whom, or simply moves everyone in parallel while preserving the pre-deliberation rank ordering. We introduce...

💬 0 commentsarXiv:2601.14221v1PDF
0

Posted in cs.SD · 2026-01-20 · Carlos Hernandez-Olivan, Hendrik Vincent Koops, Hao Hao Tan, Elio Quinton

Single-step Controllable Music Bandwidth Extension With Flow Matching

Audio restoration consists in inverting degradations of a digital audio signal to recover what would have been the pristine quality signal before the degradation occurred. This is valuable in contexts such as archives of music recordings, particularly those of precious historical value, for which a clean version may have been lost or...

💬 0 commentsarXiv:2601.14356v1PDF
0

Posted in cs.NE · 2026-01-20 · Daniel Loscos, Narciso Marti-Oliet, Ismael Rodriguez

Generalization and Completeness of Stochastic Local Search Algorithms

We generalize Stochastic Local Search (SLS) heuristics into a unique formal model. This model has two key components: a common structure designed to be as large as possible and a parametric structure intended to be as small as possible. Each heuristic is obtained by instantiating the parametric part in a different way. Particular...

💬 0 commentsarXiv:2601.14212v1PDF
0

Posted in cs.LO · 2026-01-20 · Johannes Niederhauser, Aart Middeldorp

Unification of Deterministic Higher-Order Patterns (Full Version)

We present a sound and complete unification procedure for deterministic higher-order patterns, a class of simply-typed lambda terms introduced by Yokoyama et al. which comes with a deterministic matching problem. Our unification procedure can be seen as a special case of full higher-order unification where flex-flex pairs can be...

💬 0 commentsarXiv:2601.14211v4PDF
0

Posted in cs.CL · 2026-01-20 · Rohan Bhatnagar, Youran Sun, Chi Andrew Zhang, Yixin Wen, Haizhao Yang

DRIFT: Detecting Representational Inconsistencies for Factual Truthfulness

LLMs often produce fluent but incorrect answers, yet detecting such hallucinations typically requires multiple sampling passes or post-hoc verification, adding significant latency and cost. We hypothesize that intermediate layers encode confidence signals that are lost in the final output layer, and propose a lightweight probe to read...

💬 0 commentsarXiv:2601.14210v2PDF