Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 10, 2026 — 22:45:55 EST

0

Posted in cs.CV · 2026-01-17 · Haipeng Zhou, Zhaohu Xing, Hongqiu Wang, Jun Ma, Ping Li, Lei Zhu

Toward Real-World High-Precision Image Matting and Segmentation

High-precision scene parsing tasks, including image matting and dichotomous segmentation, aim to accurately predict masks with extremely fine details (such as hair). Most existing methods focus on salient, single foreground objects. While interactive methods allow for target adjustment, their class-agnostic design restricts...

💬 0 commentsarXiv:2601.12080v1PDF
0

Posted in cs.CV · 2026-01-17 · Jing Zhang, Bingjie Fan, Jixiang Zhu, Zhe Wang

EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space

We propose EmoLat, a novel emotion latent space that enables fine-grained, text-driven image sentiment transfer by modeling cross-modal correlations between textual semantics and visual emotion features. Within EmoLat, an emotion semantic graph is constructed to capture the relational structure among emotions, objects, and visual...

💬 0 commentsarXiv:2601.12079v1PDF
0

Posted in cs.CL · 2026-01-17 · Linfeng Du, Ye Yuan, Zichen Zhao, Fuyuan Lyu, Emiliano Penaloza, Xiuying Chen, Zipeng Sun, Jikun Kang, Laurent Charlin, Xue Liu, Haolun Wu

Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization

Large language models (LLMs) excel at general-purpose tasks, yet adapting their responses to individual users remains challenging. Retrieval augmentation provides a lightweight alternative to fine-tuning by conditioning LLMs on user history records, and existing approaches typically select these records based on semantic relevance. We...

💬 0 commentsarXiv:2601.12078v2PDF
0

Posted in cs.CV · 2026-01-17 · H. Jiang, Y. Sun, Z. Dong, T. Liu, Y. Gu

CroBIM-V: Memory-Quality Controlled Remote Sensing Referring Video Object Segmentation

Remote sensing video referring object segmentation (RS-RVOS) is challenged by weak target saliency and severe visual information truncation in dynamic scenes, making it extremely difficult to maintain discriminative target representations during segmentation. Moreover, progress in this field is hindered by the absence of large-scale...

💬 0 commentsarXiv:2601.12076v1PDF
0

Posted in cs.LG · 2026-01-17 · Zoha Azimi, Reza Farahani, Radu Prodan, Christian Timmerer

ELLMPEG: An Edge-based Agentic LLM Video Processing Tool

Large language models (LLMs), the foundation of generative AI systems like ChatGPT, are transforming many fields and applications, including multimedia, enabling more advanced content generation, analysis, and interaction. However, cloud-based LLM deployments face three key limitations: high computational and energy demands, privacy...

💬 0 commentsarXiv:2602.00028v1PDF
0

Posted in cs.CL · 2026-01-17 · Mehrdad Farahani, Franziska Penzkofer, Richard Johansson

To Copy or Not to Copy: Copying Is Easier to Induce Than Recall

Language models used in retrieval-augmented settings must arbitrate between parametric knowledge stored in their weights and contextual information in the prompt. This work presents a mechanistic study of that choice by extracting an \emph{arbitration vector} from model activations on a curated dataset designed to disentangle (i)...

💬 0 commentsarXiv:2601.12075v1PDF
0

Posted in cs.LG · 2026-01-17 · Zhenyu Pu, Yu Yang, Lun Yang, Qing-Shan Jia, Xiaohong Guan, Costas J. Spanos

Representation Learning Enhanced Deep Reinforcement Learning for Optimal Operation of Hydrogen-based Multi-Energy Systems

Hydrogen-based multi-energy systems (HMES) have emerged as a promising low-carbon and energy-efficient solution, as it can enable the coordinated operation of electricity, heating and cooling supply and demand to enhance operational flexibility, improve overall energy efficiency, and increase the share of renewable integration....

💬 0 commentsarXiv:2602.00027v1PDF
0

Posted in cs.CL · 2026-01-17 · Rowzatul Zannat, Abdullah Al Shafi, Abdul Muntakim

Bridging the Gap in Bangla Healthcare: Machine Learning Based Disease Prediction Using a Symptoms-Disease Dataset

Increased access to reliable health information is essential for non-English-speaking populations, yet resources in Bangla for disease prediction remain limited. This study addresses this gap by developing a comprehensive Bangla symptoms-disease dataset containing 758 unique symptom-disease relationships spanning 85 diseases. To...

💬 0 commentsarXiv:2601.12068v1PDF
0

Posted in cs.CV · 2026-01-17 · VSS Tejaswi Abburi, Ananya Singhal, Saurabh J. Shigwan, Nitin Kumar

ARMARecon: An ARMA Convolutional Filter based Graph Neural Network for Neurodegenerative Dementias Classification

Early detection of neurodegenerative diseases such as Alzheimer's Disease (AD) and Frontotemporal Dementia (FTD) is essential for reducing the risk of progression to severe disease stages. As AD and FTD propagate along white-matter regions in a global, graph-dependent manner, graph-based neural networks are well suited to capture...

💬 0 commentsarXiv:2601.12067v1PDF
0

Posted in cs.CV · 2026-01-17 · Zijie Lou, Xiangwei Feng, Jiaxin Wang, Jiangtao Yao, Fei Che, Tianbao Liu, Chengjing Wu, Xiaochao Qu, Luoqi Liu, Ting Liu

Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninformative Gaussian noise. This approach discards the rich structural and contextual priors present in the original input video. Consequently, such methods often lack sufficient guidance,...

💬 0 commentsarXiv:2601.12066v4PDF
0

Posted in cs.CV · 2026-01-17 · Honglin Lin, Chonghan Qin, Zheng Liu, Qizhi Pei, Yu Li, Zhanping Zhong, Xin Gao, Yanfeng Wang, Conghui He, Lijun Wu

Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to-Image (T2I) models often produce outputs that are visually plausible yet scientifically incorrect, resulting in a...

💬 0 commentsarXiv:2601.17027v1PDF
0

Posted in cs.CV · 2026-01-17 · Xiaomei Yang, Antai Liu, Xizhan Gao, Fa Zhu, Sijie Niu, Giancarlo Fortino

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across different modalities. Recent research tends to use modality-shared language prompts generated by CLIP to guide the learning of modal-invariant representations. Despite achieving optimal...

💬 0 commentsarXiv:2601.12062v2PDF
0

Posted in cs.CL · 2026-01-17 · Jinsook Lee, Kirk Vanacore, Zhuqian Zhou, Bakhtawar Ahtisham, Jeanine Grutter, Rene F. Kizilcec

Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation

Dialogue Act (DA) annotation typically treats communicative or pedagogical intent as localized to individual utterances or turns. This leads annotators to agree on the underlying action while disagreeing on segment boundaries, reducing apparent reliability. We propose codebook-injected segmentation, which conditions boundary decisions...

💬 0 commentsarXiv:2601.12061v2PDF
0

Posted in cs.CC · 2026-01-17 · Ismael Rodriguez, David Rubio, Fernando Rubio

Complexity of adaptive testing in scenarios defined extensionally

In this paper we consider a testing setting where the set of possible definitions of the Implementation Under Test (IUT), as well as the behavior of each of these definitions in all possible interactions, are extensionally defined, i.e., on an element-by-element and case-by-case basis. Under this setting, the problem of finding the...

💬 0 commentsarXiv:2601.12056v1PDF
0

Posted in cs.CV · 2026-01-17 · Lina Meyer, Felix Wissel, Tobias Knopp, Susanne Pfefferle, Ralf Fliegert, Maximilian Sandmann, Liana Uebler, Franziska Möckl, Björn-Philipp Diercks, David Lohr, René Werner

Automating Parameter Selection in Deep Image Prior for Fluorescence Microscopy Image Denoising via Similarity-Based Parameter Transfer

Unsupervised deep image prior (DIP) addresses shortcomings of training data requirements and limited generalization associated with supervised deep learning. The performance of DIP depends on the network architecture and the stopping point of its iterative process. Optimizing these parameters for a new image requires time, restricting...

💬 0 commentsarXiv:2601.12055v1PDF
0

Posted in cs.CV · 2026-01-17 · Zaiyan Zhang, Jie Li, Shaowei Shi, Qiangqiang Yuan

Task-Driven Prompt Learning: A Joint Framework for Multi-modal Cloud Removal and Segmentation

Optical remote sensing imagery is indispensable for Earth observation, yet persistent cloud occlusion limits its downstream utility. Most cloud removal (CR) methods are optimized for low-level fidelity and can over-smooth textures and boundaries that are critical for analysis-ready data (ARD), leading to a mismatch between visually...

💬 0 commentsarXiv:2601.12052v2PDF
0

Posted in cs.CV · 2026-01-17 · Weixin Ye, Wei Wang, Yahui Liu, Yue Song, Bin Ren, Wei Bi, Rita Cucchiara, Nicu Sebe

A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models

In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient...

💬 0 commentsarXiv:2601.12051v1PDF
0

Posted in cs.IT · 2026-01-17 · Saeed Razavikia, Mohammad Kazemi, Deniz Gündüz, Carlo Fischione

Function Computation Over Multiple Access Channels via Hierarchical Constellations

We study function computation over a Gaussian multiple-access channel (MAC), where multiple transmitters aim at computing a function of their values at a common receiver. To this end, we propose a novel coded-modulation framework for over-the-air computation (OAC) based on hierarchical constellation design, which supports reliable...

💬 0 commentsarXiv:2601.12050v1PDF
0

Posted in cs.CV · 2026-01-17 · Chenchen Zhao, Muxi Chen, Qiang Xu

\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions

Interpretability of modern visual models is crucial, particularly in high-stakes applications. However, existing interpretability methods typically suffer from either reliance on white-box model access or insufficient quantitative rigor. To address these limitations, we introduce FocaLogic, a novel model-agnostic framework designed to...

💬 0 commentsarXiv:2601.12049v1PDF
0

Posted in cs.CR · 2026-01-17 · Xiaomei Zhang, Zhaoxi Zhang, Leo Yu Zhang, Yanjun Zhang, Guanhong Tao, Shirui Pan

Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models

Visual token compression is widely adopted to improve the inference efficiency of Large Vision-Language Models (LVLMs), enabling their deployment in latency-sensitive and resource-constrained scenarios. However, existing work has mainly focused on efficiency and performance, while the security implications of visual token compression...

💬 0 commentsarXiv:2601.12042v1PDF
0

Posted in cs.AI · 2026-01-17 · Murilo da Luz, Bruno Brandão, Luana Martins, Gustavo Oliveira, Bryan de Oliveira, Luckeciano Melo, Telma Soares

Partial Reasoning in Language Models: Search and Refinement Guided by Uncertainty

The use of Large Language Models (LLMs) for reasoning and planning tasks has drawn increasing attention in Artificial Intelligence research. Despite their remarkable progress, these models still exhibit limitations in multi-step inference scenarios, particularly in mathematical and logical reasoning. We introduce PREGU (Partial...

💬 0 commentsarXiv:2601.12040v1PDF
0

Posted in cs.AI · 2026-01-17 · Beishui Liao

Subargument Argumentation Frameworks: Separating Direct Conflict from Structural Dependency

Dung's abstract argumentation frameworks model acceptability solely in terms of an attack relation, thereby conflating two conceptually distinct aspects of argumentative reasoning: direct conflict between arguments and the structural dependencies that arise from their internal composition. While this abstraction preserves...

💬 0 commentsarXiv:2601.12038v3PDF
0

Posted in cs.CE · 2026-01-17 · Gustavo Delazeri, Marcus Ritt

Wildfire Suppression: Complexity, Models, and Instances

Wildfires cause major losses worldwide, and the frequency of fire-weather conditions is likely to increase in many regions. We study the allocation of suppression resources over time on a graph-based representation of a landscape to slow down fire propagation. Our contributions are theoretical and methodological. First, we prove that...

💬 0 commentsarXiv:2603.29865v1PDF
0

Posted in cs.HC · 2026-01-17 · Yue Yang, Christoph Leuze, Brian Hargreaves, Bruce Daniel, Fred M Baik

Multimodal Feedback for Handheld Tool Guidance: Combining Wrist-Based Haptics with Augmented Reality

We investigate how vibrotactile wrist feedback can enhance spatial guidance for handheld tool movement in optical see-through augmented reality (AR). While AR overlays are widely used to support surgical tasks, visual occlusion, lighting conditions, and interface ambiguity can compromise precision and confidence. To address these...

💬 0 commentsarXiv:2601.12037v1PDF
0

Posted in cs.SI · 2026-01-17 · Qitong Liu, Hao Peng, Zuchen Li, Xihang Meng, Ziyu Yang, Jiting Li, Li Sun, Philip S. Yu

Effective and Unsupervised Social Event Detection and Evolution via RAG and Structural Entropy

With the growing scale of social media, social event detection and evolution modeling have attracted increasing attention. Graph neural networks (GNNs) and transformer-based pre-trained language models (PLMs) have become mainstream approaches in this area. However, existing methods still face three major challenges. First, the sheer...

💬 0 commentsarXiv:2601.12035v1PDF