Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 10, 2026 — 15:19:37 EST

0

Posted in cs.AI · 2026-01-15 · Mathieu Cherpitel, Janne Luijten, Thomas Bäck, Camiel Verhamme, Martijn Tannemaat, Anna Kononova

How does downsampling affect needle electromyography signals? A generalisable workflow for understanding downsampling effects on high-frequency time series

Automated analysis of needle electromyography (nEMG) signals is emerging as a tool to support the detection of neuromuscular diseases (NMDs), yet the signals' high and heterogeneous sampling rates pose substantial computational challenges for feature-based machine-learning models, particularly for near real-time analysis. Downsampling...

💬 0 commentsarXiv:2601.10191v1PDF
0

Posted in cs.CL · 2026-01-15 · Ziang Cui, Mengran Yu, Tianjiao Li, Chenyu Shi, Yingxuan Shi, Lusheng Zhang, Hongwei Lin

HOMURA: Taming the Sand-Glass for Time-Constrained LLM Translation via Reinforcement Learning

Large Language Models (LLMs) have achieved remarkable strides in multilingual translation but are hindered by a systemic cross-lingual verbosity bias, rendering them unsuitable for strict time-constrained tasks like subtitling and dubbing. Current prompt-engineering approaches struggle to resolve this conflict between semantic...

💬 0 commentsarXiv:2601.10187v2PDF
0

Posted in cs.LG · 2026-01-15 · Kiattikun Chobtham

Reinforcement Learning to Discover a North-East Monsoon Index for Rainfall Prediction in Thailand

Accurately predicting long-term rainfall is challenging. Global climate indices, such as the El Niño-Southern Oscillation, are standard input features for machine learning. However, a significant gap persists regarding local-scale indices capable of improving predictive accuracy in specific regions of Thailand. This paper introduces a...

💬 0 commentsarXiv:2601.10181v5PDF
0

Posted in cs.LG · 2026-01-15 · Chuyi Wang, Xiaohui Xie, Tongze Wang, Yong Cui

Bias in the Shadows: Explore Shortcuts in Encrypted Network Traffic Classification

Pre-trained models operating directly on raw bytes have achieved promising performance in encrypted network traffic classification (NTC), but often suffer from shortcut learning-relying on spurious correlations that fail to generalize to real-world data. Existing solutions heavily rely on model-specific interpretation techniques,...

💬 0 commentsarXiv:2601.10180v1PDF
0

Posted in cs.DC · 2026-01-15 · Ziting Zhang, Kai Wan, Minquan Cheng, Shuo Shao, Giuseppe Caire

Distributed Linearly Separable Computation with Arbitrary Heterogeneous Data Assignment

Distributed linearly separable computation is a fundamental problem in large-scale distributed systems, requiring the computation of linearly separable functions over different datasets across distributed workers. This paper studies a heterogeneous distributed linearly separable computation problem, including one master and N...

💬 0 commentsarXiv:2601.10177v1PDF
0

Posted in cs.LG · 2026-01-15 · Mingyu Zhao, Haoran Bai, Yu Tian, Bing Zhu, Hengliang Luo

CC-OR-Net: A Unified Framework for LTV Prediction through Structural Decoupling

Customer Lifetime Value (LTV) prediction, a central problem in modern marketing, is characterized by a unique zero-inflated and long-tail data distribution. This distribution presents two fundamental challenges: (1) the vast majority of low-to-medium value users numerically overwhelm the small but critically important segment of...

💬 0 commentsarXiv:2601.10176v2PDF
0

Posted in cs.IT · 2026-01-15 · Ting Yang, Kai Wan, Minquan Cheng, Xinping Yi, Robert Caiming Qiu, Giuseppe Caire

A Low-Complexity Framework for Multi-access Coded Caching Systems with Arbitrary User-cache Access Topology

This paper studies the multi-access coded caching (MACC) problem with arbitrary user-cache access topology, which extends existing MACC models that rely on highly structured and combinatorially designed topologies. We consider a MACC system consisting of a single server, $Λ$ cache-nodes, and $K$ user-nodes. The server stores $N$...

💬 0 commentsarXiv:2601.10175v3PDF
0

Posted in cs.CR · 2026-01-15 · Hao Li, Yankai Yang, G. Edward Suh, Ning Zhang, Chaowei Xiao

ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack

Large Language Models (LLMs) have enabled the development of powerful agentic systems capable of automating complex workflows across various fields. However, these systems are highly vulnerable to indirect prompt injection attacks, where malicious instructions embedded in external data can hijack agent behavior. In this work, we...

💬 0 commentsarXiv:2601.10173v1PDF
0

Posted in cs.IT · 2026-01-15 · Guohua Zhang, Xiangya Liu, Jianhua Zhang, Yi Fang

On Existence of Girth-8 QC-LDPC Code with Large Column Weight: Combining Mirror-sequence with Classification Modulo Ten

Quasi-cyclic (QC) LDPC codes with large girths play a crucial role in several research and application fields, including channel coding, compressed sensing and distributed storage systems. A major challenge in respect of the code construction is how to obtain such codes with the shortest possible length (or equivalently, the smallest...

💬 0 commentsarXiv:2601.10170v1PDF
0

Posted in cs.AI · 2026-01-15 · Boaz Carmeli, Ron Meir, Yonatan Belinkov

CtD: Composition through Decomposition in Emergent Communication

Compositionality is a cognitive mechanism that allows humans to systematically combine known concepts in novel ways. This study demonstrates how artificial neural agents acquire and utilize compositional generalization to describe previously unseen images. Our method, termed "Composition through Decomposition", involves two sequential...

💬 0 commentsarXiv:2601.10169v1PDF
0

Posted in cs.CV · 2026-01-15 · Yue Chang, Rufeng Chen, Zhaofan Zhang, Yi Chen, Yifan Tian, Sihong Xie

RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation

Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image aggregation under occlusions and constrained viewpoints. To mitigate the impact of such...

💬 0 commentsarXiv:2601.10168v2PDF
0

Posted in cs.CL · 2026-01-15 · Nhung Nguyen Thi Hong, Cuong Nguyen Dang, Tri Le Ngoc

Credit C-GPT: A Domain-Specialized Large Language Model for Conversational Understanding in Vietnamese Debt Collection

Debt collection is a critical function within the banking, financial services, and insurance (BFSI) sector, relying heavily on large-scale human-to-human conversational interactions conducted primarily in Vietnamese contact centers. These conversations involve informal spoken language, emotional variability, and complex...

💬 0 commentsarXiv:2601.10167v1PDF
0

Posted in cs.CV · 2026-01-15 · Chao Huang, Benfeng Wang, Wei Wang, Jie Wen, Li Shen, Wenqi Ren, Yong Xu, Xiaochun Cao

Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method

Recent progress in reasoning capabilities of Multimodal Large Language Models(MLLMs) has highlighted their potential for performing complex video understanding tasks. However, in the domain of Video Anomaly Detection and Understanding (VAD&U), existing MLLM-based methods are largely limited to anomaly localization or post-hoc...

💬 0 commentsarXiv:2601.10165v1PDF
0

Posted in cs.SE · 2026-01-15 · Themistoklis Diamantopoulos, Dimosthenis Natsos, Andreas L. Symeonidis

Towards Online Malware Detection using Process Resource Utilization Metrics

The rapid growth of Cloud Computing and Internet of Things (IoT) has significantly increased the interconnection of computational resources, creating an environment where malicious software (malware) can spread rapidly. To address this challenge, researchers are increasingly utilizing Machine Learning approaches to identify malware...

💬 0 commentsarXiv:2601.10164v1PDF
0

Posted in cs.CL · 2026-01-15 · Prachuryya Kaushik, Ashish Anand

AWED-FiNER: Agents, Web applications, and Expert Detectors for Fine-grained Named Entity Recognition across 36 Languages for 6.6 Billion Speakers

Named Entity Recognition (NER) is a foundational task in Natural Language Processing (NLP) and Information Retrieval (IR), which facilitates semantic search and structured data extraction. We introduce \textbf{AWED-FiNER}, an open-source collection of agentic tool, web application, and 53 state-of-the-art expert models that provide...

💬 0 commentsarXiv:2601.10161v2PDF
0

Posted in cs.CL · 2026-01-15 · Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa, Kyle O'Brien

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may internalise corresponding behavioural priors, giving rise to self-fulfilling misalignment. This...

💬 0 commentsarXiv:2601.10160v2PDF
0

Posted in cs.CL · 2026-01-15 · Guimin Hu, Meng Li, Qiwei Peng, Lijie Hu, Boyan Xu, Ruichu Cai

What Gets Activated: Uncovering Domain and Driver Experts in MoE Language Models

Most interpretability work focuses on layer- or neuron-level mechanisms in Transformers, leaving expert-level behavior in MoE LLMs underexplored. Motivated by functional specialization in the human brain, we analyze expert activation by distinguishing domain and driver experts. In this work, we study expert activation in MoE models...

💬 0 commentsarXiv:2601.10159v2PDF
0

Posted in cs.AI · 2026-01-15 · Yusong Wang, Jialun Shen, Zhihao Wu, Yicheng Xu, Shiyin Tan, Mingkun Xu, Changshuo Wang, Zixing Song, Prayag Tiwari

MMPG: MoE-based Adaptive Multi-Perspective Graph Fusion for Protein Representation Learning

Graph Neural Networks (GNNs) have been widely adopted for Protein Representation Learning (PRL), as residue interaction networks can be naturally represented as graphs. Current GNN-based PRL methods typically rely on single-perspective graph construction strategies, which capture partial properties of residue interactions, resulting...

💬 0 commentsarXiv:2601.10157v1PDF
0

Posted in cs.CL · 2026-01-15 · Yutao Mou, Zhangchi Xue, Lijun Li, Peiyang Liu, Shikun Zhang, Wei Ye, Jing Shao

ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback

While LLM-based agents can interact with environments via invoking external tools, their expanded capabilities also amplify security risks. Monitoring step-level tool invocation behaviors in real time and proactively intervening before unsafe execution is critical for agent deployment, yet remains under-explored. In this work, we...

💬 0 commentsarXiv:2601.10156v1PDF
0

Posted in cs.LG · 2026-01-15 · Aryan Karmore

LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers

Compressing the KV cache is a required step to deploy large language models on edge devices. Current quantization methods compress storage but fail to reduce bandwidth as attention calculation requires dequantizing keys from INT4/INT8 to FP16 before use. We observe that attention scoring is mathematically equivalent to the inner...

💬 0 commentsarXiv:2601.10155v1PDF
0

Posted in cs.AI · 2026-01-15 · Leonard Nürnberg, Dennis Bontempi, Suraj Pai, Curtis Lisle, Steve Pieper, Ron Kikinis, Sil van de Leemput, Rahul Soni, Gowtham Murugesan, Cosmin Ciausu, Miriam Groeneveld, Felix J. Dorfner, Jue Jiang, Aneesh Rangnekar, Harini Veeraraghavan, Joeran S. Bosma, Keno Bressem, Raymond Mak, Andrey Fedorov, Hugo JWL Aerts

MHub.ai: A Simple, Standardized, and Reproducible Platform for AI Models in Medical Imaging

Artificial intelligence (AI) has the potential to transform medical imaging by automating image analysis and accelerating clinical research. However, research and clinical use are limited by the wide variety of AI implementations and architectures, inconsistent documentation, and reproducibility issues. Here, we introduce MHub$.$ai,...

💬 0 commentsarXiv:2601.10154v1PDF
0

Posted in cs.LG · 2026-01-15 · Qiang Yu, Xinran Cheng, Shiqiang Xu, Chuanyi Liu

Simple Network Graph Comparative Learning

The effectiveness of contrastive learning methods has been widely recognized in the field of graph learning, especially in contexts where graph data often lack labels or are difficult to label. However, the application of these methods to node classification tasks still faces a number of challenges. First, existing data enhancement...

💬 0 commentsarXiv:2601.10150v1PDF
0

Posted in cs.AI · 2026-01-15 · Xiaowei Lv, Zhilin Zhang, Yijun Li, Yusen Huo, Siyuan Ju, Xuyan Li, Chunxiang Hong, Tianyu Wang, Yongcai Wang, Peng Sun, Chuan Yu, Jian Xu, Bo Zheng

DecisionLLM: Large Language Models for Long Sequence Decision Exploration

Long-sequence decision-making, which is usually addressed through reinforcement learning (RL), is a critical component for optimizing strategic operations in dynamic environments, such as real-time bidding in computational advertising. The Decision Transformer (DT) introduced a powerful paradigm by framing RL as an autoregressive...

💬 0 commentsarXiv:2601.10148v1PDF
0

Posted in cs.AI · 2026-01-15 · Haochong Xia, Yao Long Teng, Regan Tan, Molei Qin, Xinrun Wang, Bo An

History Is Not Enough: An Adaptive Dataflow System for Financial Time-Series Synthesis

In quantitative finance, the gap between training and real-world performance-driven by concept drift and distributional non-stationarity-remains a critical obstacle for building reliable data-driven systems. Models trained on static historical data often overfit, resulting in poor generalization in dynamic markets. The mantra "History...

💬 0 commentsarXiv:2601.10143v1PDF
0

Posted in cs.CE · 2026-01-15 · Ruiran Su, Janet B. Pierrehumbert, Markus Leippold

Actors, Frames and Arguments: A Multi-Decade Computational Analysis of Climate Discourse in Financial News using Large Language Models

Financial news media shapes trillion-dollar climate investment decisions, yet discourse in this elite domain remains underexplored. We analyze two decades of climate-related articles (2000-2023) from Dow Jones Newswire using an Actor-Frame-Argument (AFA) pipeline that extracts who speaks, how issues are framed, and which arguments are...

💬 0 commentsarXiv:2601.10142v1PDF