Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 14, 2026 — 05:31:55 EST

0

Posted in cs.AI · 2026-01-08 · Nuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen, Furong Huang, Shuchao Bi, Lizhu Zhang, Zhuokai Zhao

Token-Level LLM Collaboration via FusionRoute

Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a single general-purpose model typically requires scaling to sizes that are prohibitively expensive to train and deploy. On the other hand, while smaller domain-specialized models are much more...

💬 0 commentsarXiv:2601.05106v5PDF
0

Posted in cs.RO · 2026-01-08 · Oumaima Barhoumi, Mohamed H Zaki, Sofiène Tahar

Formal Safety Guarantees for Autonomous Vehicles using Barrier Certificates

Modern AI technologies enable autonomous vehicles to perceive complex scenes, predict human behavior, and make real-time driving decisions. However, these data-driven components often operate as black boxes, lacking interpretability and rigorous safety guarantees. Autonomous vehicles operate in dynamic, mixed-traffic environments...

💬 0 commentsarXiv:2601.09740v1PDF
0

Posted in cs.IT · 2026-01-08 · Roxana Smarandache, David G. M. Mitchell

The Number of Cycles of Bi-regular Tanner Graphs in Terms of the Eigenvalues of the Adjacency Matrix

In this paper, we explore new connections between the cycles in the graph of low-density parity-check (LDPC) codes and the eigenvalues of the corresponding adjacency matrix. The resulting observations are used to derive fast, simple, recursive formulas for the number of cycles $N_{2k}$ of length $2k$, $k<g$, in a bi-regular graph of...

💬 0 commentsarXiv:2601.05340v1PDF
0

Posted in cs.CR · 2026-01-08 · Badhan Chandra Das, Md Tasnim Jawad, Joaquin Molto, M. Hadi Amini, Yanzhao Wu

Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models

In recent years, the security vulnerabilities of Multi-modal Large Language Models (MLLMs) have become a serious concern in the Generative Artificial Intelligence (GenAI) research. These highly intelligent models, capable of performing multi-modal tasks with high accuracy, are also severely susceptible to carefully launched security...

💬 0 commentsarXiv:2601.05339v1PDF
0

Posted in cs.AR · 2026-01-08 · Yuval Harary, Almog Sharoni, Esteban Garzón, Marco Lanuzza, Adam Teman, Leonid Yavits

PiC-BNN: A 128-kbit 65 nm Processing-in-CAM-Based End-to-End Binary Neural Network Accelerator

Binary Neural Networks (BNNs), where weights and activations are constrained to binary values (+1, -1), are a highly efficient alternative to traditional neural networks. Unfortunately, typical BNNs, while binarizing linear layers (matrix-vector multiplication), still implement other network layers (batch normalization, softmax,...

💬 0 commentsarXiv:2601.19920v1PDF
0

Posted in cs.RO · 2026-01-08 · Tracey Yee Hsin Tay, Xu Yan, Jonathan Ouyang, Daniel Wu, William Jiang, Jonathan Kao, Yuchen Cui

Intent at a Glance: Gaze-Guided Robotic Manipulation via Foundation Models

Designing intuitive interfaces for robotic control remains a central challenge in enabling effective human-robot interaction, particularly in assistive care settings. Eye gaze offers a fast, non-intrusive, and intent-rich input modality, making it an attractive channel for conveying user goals. In this work, we present GAMMA (Gaze...

💬 0 commentsarXiv:2601.05336v1PDF
0

Posted in cs.LG · 2026-01-08 · Fang Wu, Stan Z. Li

Dynamics-inspired Structure Hallucination for Protein-protein Interaction Modeling

Protein-protein interaction (PPI) represents a central challenge within the biology field, and accurately predicting the consequences of mutations in this context is crucial for drug design and protein engineering. Deep learning (DL) has shown promise in forecasting the effects of such mutations, but is hindered by two primary...

💬 0 commentsarXiv:2601.06214v1PDF
0

Posted in cs.CR · 2026-01-08 · Keerthi Kumar. M, Swarun Kumar Joginpelly, Sunil Khemka, Lakshmi. S R, Navin Chhibber

Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model

Background: Cyber-attacks have evolved rapidly in recent years, many individuals and business owners have been affected by cyber-attacks in various ways. Cyber-attacks include various threats such as ransomware, malware, phishing, and Denial of Service (DoS)-related attacks. Challenges: Traditional models such as Generative Artificial...

💬 0 commentsarXiv:2601.06213v1PDF
0

Posted in cs.AI · 2026-01-08 · Tengwei Song, Long Yin, Zhen Han, Zhiqiang Xu

Improving Enzyme Prediction with Chemical Reaction Equations by Hypergraph-Enhanced Knowledge Graph Embeddings

Predicting enzyme-substrate interactions has long been a fundamental problem in biochemistry and metabolic engineering. While existing methods could leverage databases of expert-curated enzyme-substrate pairs for models to learn from known pair interactions, the databases are often sparse, i.e., there are only limited and incomplete...

💬 0 commentsarXiv:2601.05330v1PDF
0

Posted in cs.SD · 2026-01-08 · Junyang Chen, Yuhang Jia, Hui Wang, Jiaming Zhou, Yong Qin

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing. To address these limitations, we propose CosyEdit, an end-to-end speech editing model adapted from CosyVoice through task-specific post-training and a...

💬 0 commentsarXiv:2601.05329v2PDF
0

Posted in cs.CV · 2026-01-08 · Fenil R. Doshi, Thomas Fel, Talia Konkle, George Alvarez

Bi-Orthogonal Factor Decomposition for Vision Transformers

Self-attention is the central computational primitive of Vision Transformers, yet we lack a principled understanding of what information attention mechanisms exchange between tokens. Attention maps describe where weight mass concentrates; they do not reveal whether queries and keys trade position, content, or both. We introduce...

💬 0 commentsarXiv:2601.05328v1PDF
0

Posted in cs.CV · 2026-01-08 · Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, Andrea Vedaldi

Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video

We propose Mesh4D, a feed-forward model for monocular 4D mesh reconstruction. Given a monocular video of a dynamic object, our model reconstructs the object's complete 3D shape and motion, represented as a deformation field. Our key contribution is a compact latent space that encodes the entire animation sequence in a single pass....

💬 0 commentsarXiv:2601.05251v1PDF
0

Posted in cs.CV · 2026-01-08 · Yuan-Kang Lee, Kuan-Lin Chen, Chia-Che Chang, Yu-Lun Liu

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes

Nighttime color constancy still remains a challenging problem in computational photography due to low-light noise and complex illumination conditions. We present RL-AWB, a novel framework combining statistical methods with deep reinforcement learning for nighttime white balance. Our method begins with a statistical algorithm tailored...

💬 0 commentsarXiv:2601.05249v4PDF
0

Posted in cs.CV · 2026-01-08 · Daniele Lizzio Bosco, Shuteng Wang, Giuseppe Serra, Vladislav Golyanik

QNeRF: Neural Radiance Fields on a Simulated Gate-Based Quantum Computer

Recently, Quantum Visual Fields (QVFs) have shown promising improvements in model compactness and convergence speed for learning the provided 2D or 3D signals. Meanwhile, novel-view synthesis has seen major advances with Neural Radiance Fields (NeRFs), where models learn a compact representation from 2D images to render 3D scenes,...

💬 0 commentsarXiv:2601.05250v1PDF
0

Posted in cs.RO · 2026-01-08 · Zhuoyang Liu, Jiaming Liu, Hao Chen, Jiale Yu, Ziyu Guo, Chengkai Hou, Chenyang Gu, Xiangju Mi, Renrui Zhang, Kun Wu, Zhengping Che, Jian Tang, Pheng-Ann Heng, Shanghang Zhang

LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future observations prior to execution. However, explicit reasoning typically incurs non-negligible inference latency, which constrains the temporal resolution...

💬 0 commentsarXiv:2601.05248v4PDF
0

Posted in cs.LO · 2026-01-08 · Oskar Fiuk

Random Models and the Guarded Fragment

Building on ideas of Gurevich and Shelah for the Gödel Class, we present a new probabilistic proof of the finite model property for the Guarded Fragment of First-Order Logic. Our proof is conceptually simple and yields the optimal doubly-exponential upper bound on the size of minimal models. We precisely analyse the obtained bound, up...

💬 0 commentsarXiv:2601.05247v2PDF
0

Posted in cs.CV · 2026-01-08 · Gangwei Xu, Haotong Lin, Hongcheng Luo, Haiyang Sun, Bing Wang, Guang Chen, Sida Peng, Hangjun Ye, Xin Yang

Pixel-Perfect Visual Geometry Estimation

Recovering clean and accurate geometry from images is essential for robotics and augmented reality. However, existing geometry foundation models still suffer severely from flying pixels and the loss of fine details. In this paper, we present pixel-perfect visual geometry models that can predict high-quality, flying-pixel-free point...

💬 0 commentsarXiv:2601.05246v1PDF
0

Posted in cs.CY · 2026-01-08 · Ritwik Gupta, Andrew W. Reddie

The LLM Mirage: Economic Interests and the Subversion of Weaponization Controls

U.S. AI security policy is increasingly shaped by an $\textit{LLM Mirage}$, the belief that national security risks scale in proportion to the compute used to train frontier language models. That premise fails in two ways. It miscalibrates strategy because adversaries can obtain weaponizable capabilities with task-specific systems...

💬 0 commentsarXiv:2601.05307v2PDF
0

Posted in cs.LG · 2026-01-08 · Natalie Collina, Jiuyao Lu, Georgy Noarov, Aaron Roth

Optimal Lower Bounds for Online Multicalibration

We prove tight lower bounds for online multicalibration, establishing an information-theoretic separation from marginal calibration. In the general setting where group functions can depend on both context and the learner's predictions, we prove an $Ω(T^{2/3})$ lower bound on expected multicalibration error using just three disjoint...

💬 0 commentsarXiv:2601.05245v2PDF
0

Posted in cs.CV · 2026-01-08 · Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Yu-Gang Jiang

GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation

Referring Expression Segmentation (RES) and Comprehension (REC) respectively segment and detect the object described by an expression, while Referring Expression Generation (REG) generates an expression for the selected object. Existing datasets and methods commonly support single-target expressions only, i.e., one expression refers...

💬 0 commentsarXiv:2601.05244v1PDF
0

Posted in cs.RO · 2026-01-08 · Xingyi He, Adhitya Polavaram, Yunhao Cao, Om Deshmukh, Tianrui Wang, Xiaowei Zhou, Kuan Fang

Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration

Functional grasping with dexterous robotic hands is a key capability for enabling tool use and complex manipulation, yet progress has been constrained by two persistent bottlenecks: the scarcity of large-scale datasets and the absence of integrated semantic and geometric reasoning in learned models. In this work, we present CorDex, a...

💬 0 commentsarXiv:2601.05243v1PDF
0

Posted in cs.CL · 2026-01-08 · Shih-Yang Liu, Xin Dong, Ximing Lu, Shizhe Diao, Peter Belcak, Mingjie Liu, Min-Hung Chen, Hongxu Yin, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Yejin Choi, Jan Kautz, Pavlo Molchanov

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

As language models become increasingly capable, users expect them to provide not only accurate responses but also behaviors aligned with diverse human preferences across a variety of scenarios. To achieve this, Reinforcement learning (RL) pipelines have begun incorporating multiple rewards, each capturing a distinct preference, to...

💬 0 commentsarXiv:2601.05242v1PDF
0

Posted in cs.CV · 2026-01-08 · Boyang Wang, Haoran Zhang, Shujie Zhang, Jinkun Hao, Mingda Jia, Qi Lv, Yucheng Mao, Zhaoyang Lyu, Jia Zeng, Xudong Xu, Jiangmiao Pang

RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation

The diversity, quantity, and quality of manipulation data are critical for training effective robot policies. However, due to hardware and physical setup constraints, collecting large-scale real-world manipulation data remains difficult to scale across diverse environments. Recent work uses text-prompt conditioned image diffusion...

💬 0 commentsarXiv:2601.05241v1PDF
0

Posted in cs.LG · 2026-01-08 · Ilmo Sung

Robust Reasoning as a Symmetry-Protected Topological Phase

Large language models suffer from "hallucinations"-logical inconsistencies induced by semantic noise. We propose that current architectures operate in a "Metric Phase," where causal order is vulnerable to spontaneous symmetry breaking. Here, we identify robust inference as an effective Symmetry-Protected Topological phase, where...

💬 0 commentsarXiv:2601.05240v1PDF
0

Posted in cs.CV · 2026-01-08 · Xiao Fu, Shitao Tang, Min Shi, Xian Liu, Jinwei Gu, Ming-Yu Liu, Dahua Lin, Chen-Hsuan Lin

Plenoptic Video Generation

Camera-controlled generative video re-rendering methods, such as ReCamMaster, have achieved remarkable progress. However, despite their success in single-view setting, these works often struggle to maintain consistency across multi-view scenarios. Ensuring spatio-temporal coherence in hallucinated regions remains challenging due to...

💬 0 commentsarXiv:2601.05239v1PDF