Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 9, 2026 — 11:52:02 EST

0

Posted in cs.CV · 2026-01-18 · Eli Passov, Nathan S. Netanyahu, Yosi Keller

Multi-Sensor Matching with HyperNetworks

Hypernetworks are models that generate or modulate the weights of another network. They provide a flexible mechanism for injecting context and task conditioning and have proven broadly useful across diverse applications without significant increases in model size. We leverage hypernetworks to improve multimodal patch matching by...

💬 0 commentsarXiv:2601.12325v1PDF
0

Posted in cs.HC · 2026-01-18 · Yue Deng, Xiaowei Chen, Junxiang Liao, Bo Li, Yixin Zou

Experiencer, Helper, or Observer: Online Fraud Intervention for Older Adults Through Role-based Simulation

Online fraud is a critical global threat that disproportionately targets older adults. Prior anti-fraud education for older adults has largely relied on static, traditional instruction that limits engagement and real-world transfer, whereas role-based simulation offers realistic yet low-risk opportunities for practice. Moreover, most...

💬 0 commentsarXiv:2601.12324v2PDF
0

Posted in cs.AI · 2026-01-18 · Yin Cai, Zhouhong Gu, Juntao Zhang, Ping Chen

MARO: Learning Stronger Reasoning from Social Interaction

Humans face countless scenarios that require reasoning and judgment in daily life. However, existing large language model training methods primarily allow models to learn from existing textual content or solve predetermined problems, lacking experience in real scenarios involving interaction, negotiation, and competition with others....

💬 0 commentsarXiv:2601.12323v2PDF
0

Posted in cs.LG · 2026-01-18 · Chang-Wei Shi, Shi-Shang Wang, Wu-Jun Li

Ordered Local Momentum for Asynchronous Distributed Learning under Arbitrary Delays

Momentum SGD (MSGD) serves as a foundational optimizer in training deep models due to momentum's key role in accelerating convergence and enhancing generalization. Meanwhile, asynchronous distributed learning is crucial for training large-scale deep models, especially when the computing capabilities of the workers in the cluster are...

💬 0 commentsarXiv:2601.12322v1PDF
0

Posted in cs.AI · 2026-01-18 · Dehao Ying, Fengchang Yu, Haihua Chen, Changjiang Jiang, Yurong Li, Wei Lu

Beyond Human Annotation: Recent Advances in Data Generation Methods for Document Intelligence

The advancement of Document Intelligence (DI) demands large-scale, high-quality training data, yet manual annotation remains a critical bottleneck. While data generation methods are evolving rapidly, existing surveys are constrained by fragmented focuses on single modalities or specific tasks, lacking a unified perspective aligned...

💬 0 commentsarXiv:2601.12318v1PDF
0

Posted in cs.LG · 2026-01-18 · Yiming Huang

Explanova: Automatically Discover Data Insights in N \times M Table via XAI Combined LLM Workflow

Automation in data analysis has been a long-time pursuit. Current agentic LLM shows a promising solution towards it. Like DeepAnalyze, DataSage, and Datawise. They are all powerful agentic frameworks for automatic fine-grained analysis and are powered by LLM-based agentic tool calling ability. However, what about powered by a preset...

💬 0 commentsarXiv:2601.12317v1PDF
0

Posted in cs.CV · 2026-01-18 · Xinyuan Zhao, Xianrui Chen, Ahmad Chaddad

GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer

We present a semantics modulated, multi scale Transformer for 3D gaze estimation. Our model conditions CLIP global features with learnable prototype banks (illumination, head pose, background, direction), fuses these prototype-enriched global vectors with CLIP patch tokens and high-resolution CNN tokens in a unified attention space,...

💬 0 commentsarXiv:2601.12316v1PDF
0

Posted in cs.SD · 2026-01-18 · Yiwen Zhang, Hui Zhang, Fanqin Meng

A Similarity Network for Correlating Musical Structure to Military Strategy

Music perception, a multi-sensory process based on the synesthesia effect, is an essential component of music aesthetic education. Understanding music structure helps both perception and aesthetic education. Music structure incorporates a range of information, the coordination of which forms the melody, just as different military...

💬 0 commentsarXiv:2601.12314v1PDF
0

Posted in cs.CR · 2026-01-18 · Ashikuzzaman, Md. Shawkat Hossain, Jubayer Abdullah Joy, Md Zahid Akon, Md Manjur Ahmed, Md. Naimul Islam

An Optimized Decision Tree-Based Framework for Explainable IoT Anomaly Detection

The increase in the number of Internet of Things (IoT) devices has tremendously increased the attack surface of cyber threats thus making a strong intrusion detection system (IDS) with a clear explanation of the process essential towards resource-constrained environments. Nevertheless, current IoT IDS systems are usually traded off...

💬 0 commentsarXiv:2601.14305v1PDF
0

Posted in cs.CV · 2026-01-18 · Xiangyu Hu, Yicheng Hong, Hongchuang Zheng, Wenjun Zeng, Bingyao Liu

S^2F-Net:A Robust Spatial-Spectral Fusion Framework for Cross-Model AIGC Detection

The rapid development of generative models has imposed an urgent demand for detection schemes with strong generalization capabilities. However, existing detection methods generally suffer from overfitting to specific source models, leading to significant performance degradation when confronted with unseen generative architectures. To...

💬 0 commentsarXiv:2601.12313v1PDF
0

Posted in cs.CV · 2026-01-18 · Yongjun Jeon, Jongmin Shin, Kanggil Park, Seonmin Park, Soyoung Lim, Jung Yong Kim, Jinsoo Rhu, Jongman Kim, Gyu-Seong Choi, Namkee Oh, Kyu-Hwan Jung

CurConMix+: A Unified Spatio-Temporal Framework for Hierarchical Surgical Workflow Understanding

Surgical action triplet recognition aims to understand fine-grained surgical behaviors by modeling the interactions among instruments, actions, and anatomical targets. Despite its clinical importance for workflow analysis and skill assessment, progress has been hindered by severe class imbalance, subtle visual variations, and the...

💬 0 commentsarXiv:2601.12312v1PDF
0

Posted in cs.NI · 2026-01-18 · Xiaofeng Luo, Jiayi He, Jiawen Kang, Ruichen Zhang, Zhaoshui He, Ekram Hossain, Dong In Kim

Cross-reality Location Privacy Protection in 6G-enabled Vehicular Metaverses: An LLM-enhanced Hybrid Generative Diffusion Model-based Approach

The emergence of 6G-enabled vehicular metaverses enables Autonomous Vehicles (AVs) to operate across physical and virtual spaces through space-air-ground-sea integrated networks. The AVs can deploy AI agents powered by large AI models as personalized assistants, on edge servers to support intelligent driving decision making and...

💬 0 commentsarXiv:2601.12311v1PDF
0

Posted in cs.AI · 2026-01-18 · Jennifer Dodgson, Alfath Daryl Alhajir, Michael Joedhitya, Akira Rafhael Janson Pattirane, Surender Suresh Kumar, Joseph Lim, C. H. Peh, Adith Ramdas, Steven Zhang Zhexu

Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection

Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic drift. This paper provides a proof-of-concept system architecture for stable self-training under sparse external feedback and bounded memory, and empirically characterises its learning...

💬 0 commentsarXiv:2601.12310v1PDF
0

Posted in cs.CV · 2026-01-18 · Anurag Kaushish, Ayan Sar, Sampurna Roy, Sudeshna Chakraborty, Prashant Trivedi, Tanupriya Choudhury, Kanav Gupta

Adaptive Multi-Scale Correlation Meta-Network for Few-Shot Remote Sensing Image Classification

Few-shot learning in remote sensing remains challenging due to three factors: the scarcity of labeled data, substantial domain shifts, and the multi-scale nature of geospatial objects. To address these issues, we introduce Adaptive Multi-Scale Correlation Meta-Network (AMC-MetaNet), a lightweight yet powerful framework with three key...

💬 0 commentsarXiv:2601.12308v1PDF
0

Posted in cs.MA · 2026-01-18 · Jiawei Xu, Arief Koesdwiady, Sisong Bei, Yan Han, Baixiang Huang, Dakuo Wang, Yutong Chen, Zheshen Wang, Peihao Wang, Pan Li, Ying Ding

Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline

Recent advances in LLM-based multi-agent systems (MAS) show that workflows composed of multiple LLM agents with distinct roles, tools, and communication patterns can outperform single-LLM baselines on complex tasks. However, most frameworks are homogeneous, where all agents share the same base LLM and differ only in prompts, tools,...

💬 0 commentsarXiv:2601.12307v1PDF
0

Posted in cs.HC · 2026-01-18 · Patrick Tresset, Markus Wulfmeier

An Embodied Companion for Visual Storytelling

As artificial intelligence shifts from pure tool for delegation toward agentic collaboration, its use in the arts can shift beyond the exploration of machine autonomy toward synergistic co-creation. While our earlier robotic works utilized automation to distance the artist's intent from the final mark, we present Companion: an...

💬 0 commentsarXiv:2603.05511v1PDF
0

Posted in cs.LG · 2026-01-18 · Deepak Kanneganti, Sajib Mistry, Sheik Fattah, Joshua Boland, Aneesh Krishna

Machine Learning as a Service (MLaaS) Dataset Generator Framework for IoT Environments

We propose a novel MLaaS Dataset Generator (MDG) framework that creates configurable and reproducible datasets for evaluating Machine Learning as a Service (MLaaS) selection and composition. MDG simulates realistic MLaaS behaviour by training and evaluating diverse model families across multiple real-world datasets and data...

💬 0 commentsarXiv:2601.12305v1PDF
0

Posted in cs.CV · 2026-01-18 · Wutao Chen, Huaqin Zou, Chen Wan, Lifeng Huang

A Two-Stage Globally-Diverse Adversarial Attack for Vision-Language Pre-training Models

Vision-language pre-training (VLP) models are vulnerable to adversarial examples, particularly in black-box scenarios. Existing multimodal attacks often suffer from limited perturbation diversity and unstable multi-stage pipelines. To address these challenges, we propose 2S-GDA, a two-stage globally-diverse attack framework. The...

💬 0 commentsarXiv:2601.12304v1PDF
0

Posted in cs.CV · 2026-01-18 · Shizhan Gong, Xiaofan Zhang, Qi Dou

Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual Representations

Deep learning has achieved remarkable success in image recognition, yet their inherent opacity poses challenges for deployment in critical domains. Concept-based interpretations aim to address this by explaining model reasoning through human-understandable concepts. However, existing post-hoc methods and ante-hoc concept bottleneck...

💬 0 commentsarXiv:2601.12303v1PDF
0

Posted in cs.IT · 2026-01-18 · Kristiina Oksner, Henk D. L. Hollmann, Ago-Erik Riet, Vitaly Skachek

On the Minimum Length of Functional Batch Codes with Small Recovery Sets

Batch codes are of potential use for load balancing and private information retrieval in distributed data storage systems. Recently, a special case of batch codes, termed functional batch codes, was proposed in the literature. In functional batch codes, users can query linear combinations of the information symbols, and not only the...

💬 0 commentsarXiv:2601.12302v2PDF
0

Posted in cs.IR · 2026-01-18 · Mingrui Liu, Sixiao Zhang, Cheng Long

Facet-Aware Multi-Head Mixture-of-Experts Model with Text-Enhanced Pre-training for Sequential Recommendation

Sequential recommendation (SR) systems excel at capturing users' dynamic preferences by leveraging their interaction histories. Most existing SR systems assign a single embedding vector to each item to represent its features, adopting various models to combine these embeddings into a sequence representation that captures user intent....

💬 0 commentsarXiv:2601.12301v1PDF
0

Posted in cs.HC · 2026-01-18 · Yue Deng, Changyang He, Bo Li, Yixin Zou

"What If My Face Gets Scanned Without Consent": Understanding Older Adults' Experiences with Biometric Payment

Biometric payment, i.e., biometric authentication implemented in digital payment systems, can reduce memory demands and streamline payment for older adults. However, older adults' perceptions and practices regarding biometric payment remain underexplored. We conducted semi-structured interviews with 22 Chinese older adults, including...

💬 0 commentsarXiv:2601.12300v2PDF
0

Posted in cs.AR · 2026-01-18 · Ye Lin, Chao Fang, Xiaoyong Song, Qi Wu, Anying Jiang, Yichuan Bai, Li Du

CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device

Edge deployment of low-batch large language models (LLMs) faces critical memory bandwidth bottlenecks when executing memory-intensive general matrix-vector multiplications (GEMV) operations. While digital processing-in-memory (PIM) architectures promise to accelerate GEMV operations, existing PIM-equipped edge devices still suffer...

💬 0 commentsarXiv:2601.12298v1PDF
0

Posted in cs.LG · 2026-01-18 · Hong Zheng, Fei Teng

Distribution Shift Is Key to Learning Invariant Prediction

An interesting phenomenon arises: Empirical Risk Minimization (ERM) sometimes outperforms methods specifically designed for out-of-distribution tasks. This motivates an investigation into the reasons behind such behavior beyond algorithmic design. In this study, we find that one such reason lies in the distribution shift across...

💬 0 commentsarXiv:2601.12296v1PDF
0

Posted in cs.AI · 2026-01-18 · Dawei Li, Yuguang Yao, Zhen Tan, Huan Liu, Ruocheng Guo

ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents

Reward-guided search methods have demonstrated strong potential in enhancing tool-using agents by effectively guiding sampling and exploration over complex action spaces. As a core design, those search methods utilize process reward models (PRMs) to provide step-level rewards, enabling more fine-grained monitoring. However, there is a...

💬 0 commentsarXiv:2601.12294v1PDF