Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 14, 2026 — 13:40:50 EST

0

Posted in cs.CL · 2026-01-07 · Jakob Schuster, Vagrant Gautam, Katja Markert

Whose Facts Win? LLM Source Preferences under Knowledge Conflicts

As large language models (LLMs) are more frequently used in retrieval-augmented generation pipelines, it is increasingly relevant to study their behavior under knowledge conflicts. Thus far, the role of the source of the retrieved information has gone unexamined. We address this gap with a novel framework to investigate how source...

💬 0 commentsarXiv:2601.03746v3PDF
0

Posted in cs.CL · 2026-01-07 · Yi Yao, He Zhu, Piaohong Wang, Jincheng Ren, Xinlong Yang, Qianben Chen, Xiaowan Li, Dingfeng Shi, Jiaxian Li, Qiexiang Wang, Sinuo Wang, Xinpeng Liu, Jiaqi Wu, Minghao Liu, Wangchunshu Zhou

O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL

The performance gap between closed-source and open-source large language models (LLMs) is largely attributed to disparities in access to high-quality training data. To bridge this gap, we introduce a novel framework for the automated synthesis of sophisticated, research-grade instructional data. Our approach centers on a multi-agent...

💬 0 commentsarXiv:2601.03743v1PDF
0

Posted in cs.CV · 2026-01-07 · Jinghan Yu, Junhao Xiao, Chenyu Zhu, Jiaming Li, Jia Li, HanMing Deng, Xirui Wang, Guoli Jia, Jianjun Li, Xiang Bai, Bowen Zhou, Zhiyuan Ma

I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing

Existing text-guided image editing methods primarily rely on end-to-end pixel-level inpainting paradigm. Despite its success in simple scenarios, this paradigm still significantly struggles with compositional editing tasks that require precise local control and complex multi-object spatial reasoning. This paradigm is severely limited...

💬 0 commentsarXiv:2601.03741v2PDF
0

Posted in cs.CV · 2026-01-07 · Shuyan Bai, Tingfa Xu, Peifu Liu, Yuhao Qiu, Huiyan Bai, Huan Chen, Yanyan Peng, Jianan Li

HyperCOD: The First Challenging Benchmark and Baseline for Hyperspectral Camouflaged Object Detection

RGB-based camouflaged object detection struggles in real-world scenarios where color and texture cues are ambiguous. While hyperspectral image offers a powerful alternative by capturing fine-grained spectral signatures, progress in hyperspectral camouflaged object detection (HCOD) has been critically hampered by the absence of a...

💬 0 commentsarXiv:2601.03736v1PDF
0

Posted in cs.CV · 2026-01-07 · Xiaoxian Shen, Yuhui Zhang, Sahithi Ankireddy, Xiaohan Wang, Maya Varma, Henry Guo, Curtis Langlotz, Serena Yeung-Levy

RadDiff: Describing Differences in Radiology Image Sets with Natural Language

Understanding how two radiology image sets differ is critical for generating clinical insights and for interpreting medical AI systems. We introduce RadDiff, a multimodal agentic system that performs radiologist-style comparative reasoning to describe clinically meaningful differences between paired radiology studies. RadDiff builds...

💬 0 commentsarXiv:2601.03733v1PDF
0

Posted in cs.SE · 2026-01-07 · Jia Li, Yuxin Su, Michael R. Lyu

From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level

As large language models (LLMs) evolve into autonomous agents, evaluating repository-level reasoning, the ability to maintain logical consistency across massive, real-world, interdependent file systems, has become critical. Current benchmarks typically fluctuate between isolated code snippets and black-box evaluations. We present...

💬 0 commentsarXiv:2601.03731v3PDF
0

Posted in cs.IR · 2026-01-07 · Fabian Haak, Philipp Schaer

Perception-Aware Bias Detection for Query Suggestions

Bias in web search has been in the spotlight of bias detection research for quite a while. At the same time, little attention has been paid to query suggestions in this regard. Awareness of the problem of biased query suggestions has been raised. Likewise, there is a rising need for automatic bias detection approaches. This paper adds...

💬 0 commentsarXiv:2601.03730v1PDF
0

Posted in cs.CV · 2026-01-07 · Donghwan Lee, Byeongjin Kim, Geunhee Kim, Hyukjin Kwon, Nahyeon Maeng, Wooju Kim

MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species

Fine-grained recognition of marine organisms is important for ecological research, biodiversity monitoring, and habitat conservation. However, existing methods often focus on the target organism alone, which can overlook informative cues from the surrounding environment. Moreover, biological taxonomy is often underused during model...

💬 0 commentsarXiv:2601.03729v3PDF
0

Posted in cs.CV · 2026-01-07 · Zhipeng Qian, Zihan Liang, Yufei Ma, Ben Chen, Huangyu Dai, Yiwei Ma, Jiayi Ji, Chenyi Lei, Han Li, Xiaoshuai Sun

CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval

Composed Image Retrieval (CIR) enables users to search for target images using both a reference image and manipulation text, offering substantial advantages over single-modality retrieval systems. However, existing CIR methods suffer from representation space fragmentation: queries and targets comprise heterogeneous modalities and are...

💬 0 commentsarXiv:2601.03728v3PDF
0

Posted in cs.CL · 2026-01-07 · Fadhil Muhammad, Alwin Djuliansah, Adrian Aryaputra Hamzah, Kurniawati Azizah

Stuttering-Aware Automatic Speech Recognition for Indonesian Language

Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like Indonesian where specialized datasets are virtually non-existent. To overcome this scarcity, we...

💬 0 commentsarXiv:2601.03727v2PDF
0

Posted in cs.LG · 2026-01-07 · Jing-Cheng Pang, Liu Sun, Chang Zhou, Xian Tang, Haichuan Ma, Kun Jiang, Jianlong Wang, Kai Zhang, Sijie Wu, Haoran Cai, Chenwei Wu, Xubin Li, Xin Chen

EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuning

Domain-specific large language models (LLMs), typically developed by fine-tuning a pre-trained general-purpose LLM on specialized datasets, represent a significant advancement in applied AI. A common strategy in LLM fine-tuning is curriculum learning, which pre-orders training samples based on metrics like difficulty to improve...

💬 0 commentsarXiv:2601.03725v1PDF
0

Posted in cs.LG · 2026-01-07 · Shijie Zhang, Kevin Zhang, Zheyuan Gu, Xiang Guo, Rujun Guo, Shaoyu Liu, Guanjun Jiang, Xiaozhao Wang

ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an important paradigm for unlocking reasoning capabilities in large language models, exemplified by the success of OpenAI o1 and DeepSeek-R1. Currently, Group Relative Policy Optimization (GRPO) stands as the dominant algorithm in this domain due to its stable...

💬 0 commentsarXiv:2601.03723v1PDF
0

Posted in cs.CV · 2026-01-07 · Wenyong Li, Qi Jiang, Weijian Hu, Kailun Yang, Zhanjun Zhang, Wenjun Tian, Kaiwei Wang, Jian Bai

Towards Real-world Lens Active Alignment with Unlabeled Data via Domain Adaptation

Active Alignment (AA) is a key technology for the large-scale automated assembly of high-precision optical systems. Compared with labor-intensive per-model on-device calibration, a digital-twin pipeline built on optical simulation offers a substantial advantage in generating large-scale labeled data. However, complex imaging...

💬 0 commentsarXiv:2601.03718v2PDF
0

Posted in cs.CY · 2026-01-07 · Mark Theby

A Mixed Methods Systematic Analysis of Issues and Factors Influencing Organizational Cloud Computing Adoption and Usage in the Public Sector: Initial Findings

Cloud computing has been shown to be an essential enabling technology for public sector organizations PSOs and offers numerous potential benefits, including reduced information technology infrastructure costs, increased innovation potential, and improved resource resilience and scalability. Despite governments' intensifying efforts to...

💬 0 commentsarXiv:2601.06175v1PDF
0

Posted in cs.SD · 2026-01-07 · Benedikt Mayrhofer, Franz Pernkopf, Philipp Aichinger, Martin Hagmüller

Lightweight and perceptually-guided voice conversion for electro-laryngeal speech

Electro-laryngeal (EL) speech is characterized by constant pitch, limited prosody, and mechanical noise, reducing naturalness and intelligibility. We propose a lightweight adaptation of the state-of-the-art StreamVC framework to this setting by removing pitch and energy modules and combining self-supervised pretraining with supervised...

💬 0 commentsarXiv:2601.03892v2PDF
0

Posted in cs.LG · 2026-01-07 · Ibrahim Delibasoglu

Spectral Manifold Regularization for Stable and Modular Routing in Deep MoE Architectures

Mixture of Experts (MoE) architectures enable efficient scaling of neural networks but suffer from expert collapse, where routing converges to a few dominant experts. This reduces model capacity and causes catastrophic interference during adaptation. We propose the Spectrally-Regularized Mixture of Experts (SR-MoE), which imposes...

💬 0 commentsarXiv:2601.03889v1PDF
0

Posted in cs.CR · 2026-01-07 · Aakash Singh, Kuldeep Singh Yadav, V. Anil Kumar, Samiran Ghosh, Pranita Baro, Basavala Bhanu Prasanth

A Longitudinal Measurement Study of Log4Shell Exploitation from a Reactive Network Telescope

The disclosure of the Log4Shell vulnerability in December 2021 led to an unprecedented wave of global scanning and exploitation activity. A recent study provided important initial insights, but was largely limited in duration and geography, focusing primarily on European and U.S. network telescope deployments and covering the...

💬 0 commentsarXiv:2601.04281v2PDF
0

Posted in cs.SD · 2026-01-07 · Yunpei Li, Xun Zhou, Jinchao Wang, Lu Wang, Yong Wu, Siyi Zhou, Yiquan Zhou, Jingchen Shu

IndexTTS 2.5 Technical Report

In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M) module, which together enable faithful emotion replication and establish the first autoregressive...

💬 0 commentsarXiv:2601.03888v3PDF
0

Posted in cs.CV · 2026-01-07 · Sanidhya Ghosal, Anurag Sharma, Sushil Ghildiyal, Mukesh Saini

FLNet: Flood-Induced Agriculture Damage Assessment using Super Resolution of Satellite Images

Distributing government relief efforts after a flood is challenging. In India, the crops are widely affected by floods; therefore, making rapid and accurate crop damage assessment is crucial for effective post-disaster agricultural management. Traditional manual surveys are slow and biased, while current satellite-based methods face...

💬 0 commentsarXiv:2601.03884v1PDF
0

Posted in cs.CR · 2026-01-07 · Liangbo Xie, Mude Cai, Xiaolong Yang, Mu Zhou, Jiacheng Wang, Dusit Niyato

A Privacy-Preserving Localization Scheme with Node Selection in Mobile Networks

Localization in mobile networks has been widely applied in many scenarios. However, an entity responsible for location estimation exposes both the target and anchors to potential location leakage at any time, creating serious security risks. Although existing studies have proposed privacy-preserving localization algorithms, they still...

💬 0 commentsarXiv:2601.04280v1PDF
0

Posted in cs.CL · 2026-01-07 · Yitong Qiao, Licheng Pan, Yu Mi, Lei Liu, Yue Shen, Fei Sun, Zhixuan Chu

Lowest Span Confidence: A Zero-Shot Metric for Efficient and Black-Box Hallucination Detection in LLMs

Hallucinations in Large Language Models (LLMs), i.e., the tendency to generate plausible but non-factual content, pose a significant challenge for their reliable deployment in high-stakes environments. However, existing hallucination detection methods generally operate under unrealistic assumptions, i.e., either requiring expensive...

💬 0 commentsarXiv:2601.19918v1PDF
0

Posted in cs.LG · 2026-01-07 · Shudong Liu, Hanwen Zhang, Xiuling Wang, Yuesheng Zhu, Guibo Luo

Feature-Aware One-Shot Federated Learning via Hierarchical Token Sequences

One-shot federated learning (OSFL) reduces the communication cost and privacy risks of iterative federated learning by constructing a global model with a single round of communication. However, most existing methods struggle to achieve robust performance on real-world domains such as medical imaging, or are inefficient when handling...

💬 0 commentsarXiv:2601.03882v1PDF
0

Posted in cs.SE · 2026-01-07 · Giovanni Rosa, David Moreno-Lumbreras, Gregorio Robles, Jesús M. González-Barahona

Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design

Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their behavior in structured, specification-driven processes remains poorly understood. This paper presents an empirical study design using CURRANTE, a Visual Studio Code extension that enables a human-in-the-loop workflow for...

💬 0 commentsarXiv:2601.03878v1PDF
0

Posted in cs.LG · 2026-01-07 · Pau Esteve, Massimiliano Zanin

Generation of synthetic delay time series for air transport applications

The generation of synthetic data is receiving increasing attention from the scientific community, thanks to its ability to solve problems like data scarcity and privacy, and is starting to find applications in air transport. We here tackle the problem of generating synthetic, yet realistic, time series of delays at airports, starting...

💬 0 commentsarXiv:2601.04279v1PDF
0

Posted in cs.CL · 2026-01-07 · Xiaoyu Xu, Minxin Du, Zitong Li, Zi Liang, Zhibiao Guo, Shiyu Zhang, Peizhao Hu, Qingqing Ye, Haibo Hu

From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning

Although machine unlearning is essential for removing private, harmful, or copyrighted content from LLMs, current benchmarks often fail to faithfully represent the true ``forgetting scope'' learned by the model. We formalize two distinct unlearning granularities, domain-level and instance-level, and propose \BiForget, an automated...

💬 0 commentsarXiv:2601.04278v2PDF