Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 7, 2026 — 18:52:22 EST

0

Posted in cs.AI · 2026-01-21 · Amaury Guichard, Laurent Michel, Hélène Verhaeghe, Pierre Schaus

Towards Bound Consistency for the No-Overlap Constraint Using MDDs

Achieving bound consistency for the no-overlap constraint is known to be NP-complete. Therefore, several polynomial-time tightening techniques, such as edge finding, not-first-not-last reasoning, and energetic reasoning, have been introduced for this constraint. In this work, we derive the first bound-consistent algorithm for the...

💬 0 commentsarXiv:2601.14784v1PDF
0

Posted in cs.CL · 2026-01-21 · Anqi Li, Yuqian Chen, Yu Lu, Zhaoming Chen, Yuan Xie, Zhenzhong Lan

RECAP: Resistance Capture in Text-based Mental Health Counseling with Large Language Models

Recognizing and navigating client resistance is critical for effective mental health counseling, yet detecting such behaviors is particularly challenging in text-based interactions. Existing NLP approaches oversimplify resistance categories, ignore the sequential dynamics of therapeutic interventions, and offer limited...

💬 0 commentsarXiv:2601.14780v1PDF
0

Posted in cs.CR · 2026-01-21 · Yuang Qi, Na Zhao, Qiyi Yao, Benlong Wu, Weiming Zhang, Nenghai Yu, Kejiang Chen

STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model

Recent provably secure linguistic steganography (PSLS) methods rely on mainstream autoregressive language models (ARMs) to address historically challenging tasks, that is, to disguise covert communication as ``innocuous'' natural language communication. However, due to the characteristic of sequential generation of ARMs, the stegotext...

💬 0 commentsarXiv:2601.14778v1PDF
0

Posted in cs.CV · 2026-01-21 · Jiaxuan Liu, Yang Xiang, Han Zhao, Xiangang Li, Zhenhua Ling

FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

Movie dubbing is the task of synthesizing speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. However, existing methods face two major limitations: (1) high-quality multimodal dubbing datasets are limited in scale, suffer from...

💬 0 commentsarXiv:2601.14777v1PDF
0

Posted in cs.CV · 2026-01-21 · Xiaofan Yang, Yubin Liu, Wei Pan, Guoqing Chu, Junming Zhang, Jie Zhao, Zhuoqi Man, Xuanming Cao

M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention

Recent advances in multi-modal detection have significantly improved detection accuracy in challenging environments (e.g., low light, overexposure). By integrating RGB with modalities such as thermal and depth, multi-modal fusion increases data redundancy and system robustness. However, significant challenges remain in effectively...

💬 0 commentsarXiv:2601.14776v3PDF
0

Posted in cs.CV · 2026-01-21 · Keita Takeda, Tomoya Sakai

Does medical specialization of VLMs enhance discriminative power?: A comprehensive investigation through feature distribution analysis

This study investigates the feature representations produced by publicly available open source medical vision-language models (VLMs). While medical VLMs are expected to capture diagnostically relevant features, their learned representations remain underexplored, and standard evaluations like classification accuracy do not fully reveal...

💬 0 commentsarXiv:2601.14774v1PDF
0

Posted in cs.AI · 2026-01-21 · Haizhou Liu, Haodong Jin, Yiming Wang, Hui Yu

Semantic-Guided Unsupervised Video Summarization

Video summarization is a crucial technique for social understanding, enabling efficient browsing of massive multimedia content and extraction of key information from social platforms. Most existing unsupervised summarization methods rely on Generative Adversarial Networks (GANs) to enhance keyframe selection and generate coherent,...

💬 0 commentsarXiv:2601.14773v1PDF
0

Posted in cs.CV · 2026-01-21 · Puneet Sharma, Kristian Dalsbø Hindberg, Eibe Frank, Benedicte Schelde-Olesen, Ulrik Deding

Using Multi-Instance Learning to Identify Unique Polyps in Colon Capsule Endoscopy Images

Identifying unique polyps in colon capsule endoscopy (CCE) images is a critical yet challenging task for medical personnel due to the large volume of images, the cognitive load it creates for clinicians, and the ambiguity in labeling specific frames. This paper formulates this problem as a multi-instance learning (MIL) task, where a...

💬 0 commentsarXiv:2601.14771v1PDF
0

Posted in cs.GR · 2026-01-21 · Chun Chen, Minseok Chae, Seung-Woo Nam, Myeong-Ho Choi, Minseong Kim, Eunbi Lee, Yoonchan Jeong, Jae-Hyeung Park

PAColorHolo: A Perceptually-Aware Color Management Framework for Holographic Displays

Holographic displays offer significant potential for augmented and virtual reality applications by reconstructing wavefronts that enable continuous depth cues and natural parallax without vergence-accommodation conflict. However, despite advances in pixel-level image quality, current systems struggle to achieve perceptually accurate...

💬 0 commentsarXiv:2601.14766v1PDF
0

Posted in cs.LG · 2026-01-21 · Harold Kiossou, Pierre Schaus, Siegfried Nijssen

Anytime Optimal Decision Tree Learning with Continuous Features

In recent years, significant progress has been made on algorithms for learning optimal decision trees, primarily in the context of binary features. Extending these methods to continuous features remains substantially more challenging due to the large number of potential splits for each feature. Recently, an elegant exact algorithm was...

💬 0 commentsarXiv:2601.14765v1PDF
0

Posted in cs.AI · 2026-01-21 · Thomas Eiter, Tobias Geibinger, Zeynep G. Saribatur

An XAI View on Explainable ASP: Methods, Systems, and Perspectives

Answer Set Programming (ASP) is a popular declarative reasoning and problem solving approach in symbolic AI. Its rule-based formalism makes it inherently attractive for explainable and interpretive reasoning, which is gaining importance with the surge of Explainable AI (XAI). A number of explanation approaches and tools for ASP have...

💬 0 commentsarXiv:2601.14764v2PDF
0

Posted in cs.LG · 2026-01-21 · Injin Kong, Hyoungjoon Lee, Yohan Jo

Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models

Post-training pretrained autoregressive models (ARMs) into masked diffusion models (MDMs) has emerged as a cost-effective way to overcome the limitations of sequential generation. Yet it remains unclear whether post-trained MDMs acquire genuinely new computational mechanisms or merely re-express autoregressive computation in a...

💬 0 commentsarXiv:2601.14758v4PDF
0

Posted in cs.CV · 2026-01-21 · Kangcheng Zhou, Jun Jiang, Qing Zhang, Shuang Zheng, Qingli Li, Shugong Xu

ReinPath: A Multimodal Reinforcement Learning Approach for Pathology

Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited interpretability due to the lack of high-quality dataset that support explicit reasoning and inference...

💬 0 commentsarXiv:2601.14757v1PDF
0

Posted in cs.IT · 2026-01-21 · Lei Xie, Peilan Wang, Guanxiong Shen, Guyue Li, Weidong Mei, Liquan Chen

Secure Communication in MIMOME Movable-Antenna Systems with Statistical Eavesdropper CSI

This paper investigates the potential of movable antennas (MAs) to enhance physical layer security within a multiple-input multiple-output multiple-antenna eavesdropper (MIMOME) system. We consider a practical scenario where the transmitter operates with imperfect eavesdropper channel state information (ECSI), knowing only the...

💬 0 commentsarXiv:2601.14755v1PDF
0

Posted in cs.DL · 2026-01-21 · Marilena Daquino, Francesca Mambelli, Artem Kozlov

Many-to-many. Usability challenges of entity reconciliation in art history and photographic studies

This article investigates challenges in reconciling heterogeneous records across cultural institutions, focusing on art historical photo archives within the PHAROS consortium. Through case studies, the study analyses reconciliation workflows and cataloguing traditions, with attention to institutional contexts, data granularities, and...

💬 0 commentsarXiv:2601.14753v1PDF
0

Posted in cs.CL · 2026-01-21 · Yifan Wang, Shiyu Li, Peiming Li, Xiaochen Yang, Yang Tang, Zheng Wei

Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning

Chain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs). Although CoT prompting enhances reasoning, its verbosity imposes substantial computational overhead. Recent works often focus exclusively on outcome alignment and lack supervision on the...

💬 0 commentsarXiv:2601.14750v4PDF
0

Posted in cs.LG · 2026-01-21 · Hongyue Wu, Hangyu Li, Guodong Fan, Haoran Zhu, Shizhan Chen, Zhiyong Feng

RefProtoFL: Communication-Efficient Federated Learning via External-Referenced Prototype Alignment

Federated learning (FL) enables collaborative model training without sharing raw data in edge environments, but is constrained by limited communication bandwidth and heterogeneous client data distributions. Prototype-based FL mitigates this issue by exchanging class-wise feature prototypes instead of full model parameters; however,...

💬 0 commentsarXiv:2601.14746v2PDF
0

Posted in cs.SD · 2026-01-21 · Hongfu Liu, Zhouying Cui, Xiangming Gu, Ye Wang

Unlocking Large Audio-Language Models for Interactive Language Learning

Achieving pronunciation proficiency in a second language (L2) remains a challenge, despite the development of Computer-Assisted Pronunciation Training (CAPT) systems. Traditional CAPT systems often provide unintuitive feedback that lacks actionable guidance, limiting its effectiveness. Recent advancements in audio-language models...

💬 0 commentsarXiv:2601.14744v1PDF
0

Posted in cs.SE · 2026-01-21 · Konstantin Poddubnyy, Igor Vozniak, Ivan Burmistrov, Nils Lipp, Davit Hovhannisyan, Christian Mueller, Philipp Slusallek

ARISE -- Adaptive Refinement and Iterative Scenario Engineering

The effectiveness of collision-free trajectory planners depends on the quality and diversity of training data, especially for rare scenarios. A widely used approach to improve dataset diversity involves generating realistic synthetic traffic scenarios. However, producing such scenarios remains difficult due to the precision required...

💬 0 commentsarXiv:2601.14743v3PDF
0

Posted in cs.CV · 2026-01-21 · Ami Pandat, Kanyala Muvva, Punna Rajasekhar, Gopika Vinod, Rohit Shukla

SimD3: A Synthetic drone Dataset with Payload and Bird Distractor Modeling for Robust Detection

Reliable drone detection is challenging due to limited annotated real-world data, large appearance variability, and the presence of visually similar distractors such as birds. To address these challenges, this paper introduces SimD3, a large-scale high-fidelity synthetic dataset designed for robust drone detection in complex aerial...

💬 0 commentsarXiv:2601.14742v1PDF
0

Posted in cs.CV · 2026-01-21 · Chongbin Yi, Yuxin Liang, Ziqi Zhou, Peng Yang

Enhancing Text-to-Image Generation via End-Edge Collaborative Hybrid Super-Resolution

Artificial Intelligence-Generated Content (AIGC) has made significant strides, with high-resolution text-to-image (T2I) generation becoming increasingly critical for improving users' Quality of Experience (QoE). Although resource-constrained edge computing adequately supports fast low-resolution T2I generations, achieving...

💬 0 commentsarXiv:2601.14741v1PDF
0

Posted in cs.CV · 2026-01-21 · Liqin Wang, Qianyue Hu, Wei Lu, Xiangyang Luo

Safeguarding Facial Identity against Diffusion-based Face Swapping via Cascading Pathway Disruption

The rapid evolution of diffusion models has democratized face swapping but also raises concerns about privacy and identity security. Existing proactive defenses, often adapted from image editing attacks, prove ineffective in this context. We attribute this failure to an oversight of the structural resilience and the unique static...

💬 0 commentsarXiv:2601.14738v1PDF
0

Posted in cs.DB · 2026-01-21 · Dildar Ali, Suman Banerjee, Rajibul Islam

Trajectory-Driven Multi-Product Influence Maximization in Billboard Advertising

Billboard Advertising has emerged as an effective out-of-home advertising technique, where the goal is to select a limited number of slots and play advertisement content there, with the hope that it will be observed by many people and, effectively, a significant number of them will be influenced towards the brand. Given a trajectory...

💬 0 commentsarXiv:2601.14737v1PDF
0

Posted in cs.DC · 2026-01-21 · Varad Kulkarni, Vaibhav Jha, Nikhil Reddy, Anand Eswaran, Praveen Jayachandran, Yogesh Simmhan

Optimizing FaaS Platforms for MCP-enabled Agentic Workflows

Agentic workflows that use autonomous AI Agents powered by Large Language Models (LLMs) and Model Context Protocol (MCP) servers is rapidly rising. This introduces challenges in scalable cloud deployment and state management. Traditional hosting on Virtual Machines (VMs) is resource-intensive and lacks elasticity....

💬 0 commentsarXiv:2601.14735v2PDF
0

Posted in cs.CV · 2026-01-21 · Jing Lan, Hexiao Ding, Hongzhao Chen, Yufeng Jiang, Nga-Chun Ng, Gwing Kei Yip, Gerald W. Y. Cheng, Yunlin Mao, Jing Cai, Liang-ting Lin, Jung Sun Yoo

DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling

AI models for drug discovery and chemical literature mining must interpret molecular images and generate outputs consistent with 3D geometry and stereochemistry. Most molecular language models rely on strings or graphs, while vision-language models often miss stereochemical details and struggle to map continuous 3D structures into...

💬 0 commentsarXiv:2601.14732v1PDF