Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 9, 2026 — 21:50:29 EST

0

Posted in cs.CV · 2026-01-19 · Takaki Yamamoto, Chihiro Noguchi, Toshihiro Tanizawa

Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data

Spatial understanding remains a key challenge in vision-language models. Yet it is still unclear whether such understanding is truly acquired, and if so, through what mechanisms. We present a controllable 1D image-text testbed to probe how left-right relational understanding emerges in Transformer-based vision and text encoders...

💬 0 commentsarXiv:2601.12809v2PDF
0

Posted in cs.IT · 2026-01-19 · Tong Wu, Zhiyong Chen, Guo Lu, Li Song, Feng Yang, Meixia Tao, Wenjun Zhang

Joint Source-Channel-Generation Coding: From Distortion-oriented Reconstruction to Semantic-consistent Generation

Conventional communication systems, including both separation-based coding and AI-driven joint source-channel coding (JSCC), are largely guided by Shannon's rate-distortion theory. However, relying on generic distortion metrics fails to capture complex human visual perception, often resulting in blurred or unrealistic reconstructions....

💬 0 commentsarXiv:2601.12808v1PDF
0

Posted in cs.LG · 2026-01-19 · Zixing Song, Irwin King

Semi-supervised Instruction Tuning for Large Language Models on Text-Attributed Graphs

The emergent reasoning capabilities of Large Language Models (LLMs) offer a transformative paradigm for analyzing text-attributed graphs. While instruction tuning is the prevailing method for adapting pre-trained LLMs to graph learning tasks like node classification, it requires a substantial volume of annotated (INSTRUCTION, OUTPUT)...

💬 0 commentsarXiv:2601.12807v1PDF
0

Posted in cs.CR · 2026-01-19 · Nay Myat Min, Long H. Pham, Hongyu Zhang, Jun Sun

CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models

Single-pass hallucination detectors rely on internal telemetry (e.g., uncertainty, hidden-state geometry, and attention) of large language models, implicitly assuming hallucinations leave separable traces in these signals. We study a white-box, model-side adversary that fine-tunes lightweight LoRA adapters on the model while keeping...

💬 0 commentsarXiv:2601.14310v1PDF
0

Posted in cs.AI · 2026-01-19 · Hanwei Zhang, Luo Cheng, Rui Wen, Yang Zhang, Lijun Zhang, Holger Hermanns

SL-CBM: Enhancing Concept Bottleneck Models with Semantic Locality for Better Interpretability

Explainable AI (XAI) is crucial for building transparent and trustworthy machine learning systems, especially in high-stakes domains. Concept Bottleneck Models (CBMs) have emerged as a promising ante-hoc approach that provides interpretable, concept-level explanations by explicitly modeling human-understandable concepts. However,...

💬 0 commentsarXiv:2601.12804v1PDF
0

Posted in cs.SD · 2026-01-19 · Jihoo Jung, Ji-Hoon Kim, Doyeop Kwak, Junwon Lee, Juhan Nam, Joon Son Chung

UNMIXX: Untangling Highly Correlated Singing Voices Mixtures

We introduce UNMIXX, a novel framework for multiple singing voices separation (MSVS). While related to speech separation, MSVS faces unique challenges: data scarcity and the highly correlated nature of singing voices mixture. To address these issues, we propose UNMIXX with three key components: (1) musically informed mixing strategy...

💬 0 commentsarXiv:2601.12802v1PDF
0

Posted in cs.RO · 2026-01-19 · Peng Li, Zihan Zhuang, Yangfan Gao, Yi Dong, Sixian Li, Changhao Jiang, Shihan Dou, Zhiheng Xi, Enyu Zhou, Jixuan Huang, Hui Li, Jingjing Gong, Xingjun Ma, Tao Gui, Zuxuan Wu, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Xipeng Qiu

FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions

Humanoid robots are capable of performing various actions such as greeting, dancing and even backflipping. However, these motions are often hard-coded or specifically trained, which limits their versatility. In this work, we present FRoM-W1, an open-source framework designed to achieve general humanoid whole-body motion control using...

💬 0 commentsarXiv:2601.12799v1PDF
0

Posted in cs.CV · 2026-01-19 · Zhihan Zeng, Yang Zhao, Kaihe Wang, Dusit Niyato, Yue Xiu, Lu Chen, Zhongpei Zhang, Ning Wei

PhyG-MoE: A Physics-Guided Mixture-of-Experts Framework for Energy-Efficient GNSS Interference Recognition

Complex electromagnetic interference increasingly compromises Global Navigation Satellite Systems (GNSS), threatening the reliability of Space-Air-Ground Integrated Networks (SAGIN). Although deep learning has advanced interference recognition, current static models suffer from a \textbf{fundamental limitation}: they impose a fixed...

💬 0 commentsarXiv:2601.12798v1PDF
0

Posted in cs.CL · 2026-01-19 · Jiawei Xu, Zhenyu Yu, Ziqian Bi, Minh Duc Pham, Xiaoyi Qu, Danyang Zhang

PRIME: Policy-Reinforced Iterative Multi-agent Execution for Algorithmic Reasoning in Large Language Models

Large language models have demonstrated remarkable capabilities across diverse reasoning tasks, yet their performance on algorithmic reasoning remains limited. To handle this limitation, we propose PRIME (Policy-Reinforced Iterative Multi-agent Execution), a framework comprising three specialized agents, an executor for step-by-step...

💬 0 commentsarXiv:2602.11170v1PDF
0

Posted in cs.RO · 2026-01-19 · Changwei Jing, Jai Krishna Bandi, Jianglong Ye, Yan Duan, Pieter Abbeel, Xiaolong Wang, Sha Yi

Contact-Aware Neural Dynamics

High-fidelity physics simulation is essential for scalable robotic learning, but the sim-to-real gap persists, especially for tasks involving complex, dynamic, and discontinuous interactions like physical contacts. Explicit system identification, which tunes explicit simulator parameters, is often insufficient to align the intricate,...

💬 0 commentsarXiv:2601.12796v1PDF
0

Posted in cs.CV · 2026-01-19 · Zeren Sun, Yazhou Yao, Tongliang Liu, Zechao Li, Fumin Shen, Jinhui Tang

Combating Noisy Labels through Fostering Self- and Neighbor-Consistency

Label noise is pervasive in various real-world scenarios, posing challenges in supervised deep learning. Deep networks are vulnerable to such label-corrupted samples due to the memorization effect. One major stream of previous methods concentrates on identifying clean data for training. However, these methods often neglect imbalances...

💬 0 commentsarXiv:2601.12795v1PDF
0

Posted in cs.CV · 2026-01-19 · Zhihan Zeng, Yang Zhao, Kaihe Wang, Dusit Niyato, Hongyuan Shu, Junchu Zhao, Yanjun Huang, Yue Xiu, Zhongpei Zhang, Ning Wei

SKANet: A Cognitive Dual-Stream Framework with Adaptive Modality Fusion for Robust Compound GNSS Interference Classification

As the electromagnetic environment becomes increasingly complex, Global Navigation Satellite Systems (GNSS) face growing threats from sophisticated jamming interference. Although Deep Learning (DL) effectively identifies basic interference, classifying compound interference remains difficult due to the superposition of diverse jamming...

💬 0 commentsarXiv:2601.12791v1PDF
0

Posted in cs.RO · 2026-01-19 · Yang Zhang, Jianming Ma, Liyun Yan, Zhanxiang Cao, Yazhou Zhang, Haoyang Li, Yue Gao

FocusNav: Spatial Selective Attention with Waypoint Guidance for Humanoid Local Navigation

Robust local navigation in unstructured and dynamic environments remains a significant challenge for humanoid robots, requiring a delicate balance between long-range navigation targets and immediate motion stability. In this paper, we propose FocusNav, a spatial selective attention framework that adaptively modulates the robot's...

💬 0 commentsarXiv:2601.12790v1PDF
0

Posted in cs.CR · 2026-01-19 · Suyang Sun, Weifei Jin, Yuxin Cao, Wei Song, Jie Hao

DUAP: Dual-task Universal Adversarial Perturbations Against Voice Control Systems

Modern Voice Control Systems (VCS) rely on the collaboration of Automatic Speech Recognition (ASR) and Speaker Recognition (SR) for secure interaction. However, prior adversarial attacks typically target these tasks in isolation, overlooking the coupled decision pipeline in real-world scenarios. Consequently, single-task attacks often...

💬 0 commentsarXiv:2601.12786v2PDF
0

Posted in cs.LG · 2026-01-19 · Yuqi Li, Kuiye Ding, Chuanguang Yang, Szu-Yu Chen, Yingli Tian

Distilling Time Series Foundation Models for Efficient Forecasting

Time Series foundation models (TSFMs) deliver strong forecasting performance through large-scale pretraining, but their large parameter sizes make deployment costly. While knowledge distillation offers a natural and effective approach for model compression, techniques developed for general machine learning tasks are not directly...

💬 0 commentsarXiv:2601.12785v1PDF
0

Posted in cs.DC · 2026-01-19 · Haoyang Li, Sheng Lin, Fangcheng Fu, Yuming Zhou, Xiaodong Ji, Yanfeng Zhao, Lefeng Wang, Jie Jiang, Bin Cui

Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination

Reinforcement learning (RL) post-training has become pivotal for enhancing the capabilities of modern large models. A recent trend is to develop RL systems with a fully disaggregated architecture, which decouples the three RL phases (rollout, reward, and training) onto separate resources and executes them asynchronously. However, two...

💬 0 commentsarXiv:2601.12784v1PDF
0

Posted in cs.AI · 2026-01-19 · Hyejin Park, Junhyuk Kwon, Suha Kwak, Jungseul Ok

VIRO: Robust and Efficient Neuro-Symbolic Reasoning with Verification for Referring Expression Comprehension

Referring Expression Comprehension (REC) aims to localize the image region corresponding to a natural language query. Recent neuro-symbolic REC approaches leverage large language models (LLMs) and vision-language models (VLMs) to perform compositional reasoning, decomposing queries into structured programs and executing them...

💬 0 commentsarXiv:2601.12781v2PDF
0

Posted in cs.IT · 2026-01-19 · Zhe Sun, Terry Shue Chien Lau, Mengying Zhao, Zimeng Zhou, Fang-Wei Fu

Extended Gabidulin-Kronecker Product Codes and Their Application to Cryptosystems

In this paper, we initiate the study of Extended Gabidulin codes with a Kronecker product structure and propose three enhanced variants of the Rank Quasi-Cyclic (RQC) (Melchor et.al., IEEE IT, 2018) cryptosystem. First, we establish precise bounds on the minimum rank distance of Gabidulin-Kronecker product codes under two distinct...

💬 0 commentsarXiv:2601.12780v1PDF
0

Posted in cs.CV · 2026-01-19 · Nafis Sadeq, Qingfeng Liu, Mostafa El-Khamy

Open Vocabulary Panoptic Segmentation With Retrieval Augmentation

Given an input image and set of class names, panoptic segmentation aims to label each pixel in an image with class labels and instance labels. In comparison, Open Vocabulary Panoptic Segmentation aims to facilitate the segmentation of arbitrary classes according to user input. The challenge is that a panoptic segmentation system...

💬 0 commentsarXiv:2601.12779v1PDF
0

Posted in cs.LG · 2026-01-19 · Yuta Hirabayashi, Daisuke Matusoka, Konobu Kimura

Eddy-Resolving Global Ocean Forecasting with Multi-Scale Graph Neural Networks

Research on data-driven ocean models has progressed rapidly in recent years; however, the application of these models to global eddy-resolving ocean forecasting remains limited. The accurate representation of ocean dynamics across a wide range of spatial scales remains a major challenge in such applications. This study proposes a...

💬 0 commentsarXiv:2601.12775v1PDF
0

Posted in cs.NI · 2026-01-19 · Yulu Han, Ziye Jia, Jingjing Zhao, Lijun He, Yao Wu, Qihui Wu

SDN-Blockchain Based Security Routing for UAV Communication via Reinforcement Learning

The unmanned aerial vehicle (UAV) network plays important roles in emergency communications. However, it is challenging to design reliable routing strategies that ensure low latency, energy efficiency, and security in the dynamic and attack-prone environments. To this end, we design a secure routing architecture integrating...

💬 0 commentsarXiv:2601.12774v1PDF
0

Posted in cs.CL · 2026-01-19 · Keito Inoshita

Who Does This Name Remind You of ? Nationality Prediction via Large Language Model Associative Memory

Large language models (LLMs) possess extensive world knowledge, yet methods for effectively eliciting this knowledge remain underexplored. Nationality and region prediction tasks require understanding of not only linguistic features but also cultural and historical background, making LLM world knowledge particularly valuable. However,...

💬 0 commentsarXiv:2601.12771v2PDF
0

Posted in cs.CV · 2026-01-19 · Shuling Zhao, Dan Xu

One-Shot Feed-Forward 360$^{\circ}$ Animatable Avatar via Inpainted UV-Space Gaussian Modeling

Building one-shot 3D animatable head avatars is an important yet challenging problem. Existing methods generally collapse under large camera pose variations, compromising the realism of 3D avatars. In this work, we propose a new framework to tackle the novel setting of one-shot 3D full-head animatable avatar reconstruction in a single...

💬 0 commentsarXiv:2601.12770v2PDF
0

Posted in cs.CV · 2026-01-19 · Zequn Xie, Boyun Zhang, Yuxiao Lin, Tao Jin

Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval

Video-text retrieval (VTR) aims to locate relevant videos using natural language queries. Current methods, often based on pre-trained models like CLIP, are hindered by video's inherent redundancy and their reliance on coarse, final-layer features, limiting matching accuracy. To address this, we introduce the HVP-Net (Hierarchical...

💬 0 commentsarXiv:2601.12768v1PDF
0

Posted in cs.CV · 2026-01-19 · Lu Yue, Yue Fan, Shiwei Lian, Yu Zhao, Jiaxin Yu, Liang Xie, Feitian Zhang

Spatial-VLN: Zero-Shot Vision-and-Language Navigation With Explicit Spatial Perception and Exploration

Zero-shot Vision-and-Language Navigation (VLN) agents leveraging Large Language Models (LLMs) excel in generalization but suffer from insufficient spatial perception. Focusing on complex continuous environments, we categorize key perceptual bottlenecks into three spatial challenges: door interaction,multi-room navigation, and...

💬 0 commentsarXiv:2601.12766v1PDF