Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 11, 2026 — 17:13:09 EST

0

Posted in cs.CV · 2026-01-13 · Phuoc-Nguyen Bui, Toan Duc Nguyen, Junghyun Bum, Duc-Tai Le, Hyunseung Choo

Representation Learning with Semantic-aware Instance and Sparse Token Alignments

Medical contrastive vision-language pre-training (VLP) has demonstrated significant potential in improving performance on downstream tasks. Traditional approaches typically employ contrastive learning, treating paired image-report samples as positives and unpaired ones as negatives. However, in medical datasets, there can be...

💬 0 commentsarXiv:2601.08165v2PDF
0

Posted in cs.CV · 2026-01-13 · Jing Tao, Banglei Guan, Pengju Sun, Taihang Lei, Yang Shang, Qifeng Yu

A Hardware-Algorithm Co-Designed Framework for HDR Imaging and Dehazing in Extreme Rocket Launch Environments

Quantitative optical measurement of critical mechanical parameters -- such as plume flow fields, shock wave structures, and nozzle oscillations -- during rocket launch faces severe challenges due to extreme imaging conditions. Intense combustion creates dense particulate haze and luminance variations exceeding 120 dB, degrading image...

💬 0 commentsarXiv:2601.08162v1PDF
0

Posted in cs.RO · 2026-01-13 · Jing Tao, Banglei Guan, Yang Shang, Shunkun Liang, Qifeng Yu

Robust Subpixel Localization of Diagonal Markers in Large-Scale Navigation via Multi-Layer Screening and Adaptive Matching

This paper proposes a robust, high-precision positioning methodology to address localization failures arising from complex background interference in large-scale flight navigation and the computational inefficiency inherent in conventional sliding window matching techniques. The proposed methodology employs a three-tiered framework...

💬 0 commentsarXiv:2601.08161v1PDF
0

Posted in cs.CL · 2026-01-13 · Anxin Tian, Yiming Li, Xing Li, Hui-Ling Zhen, Lei Chen, Xianzhi Yu, Zhenhua Dong, Mingxuan Yuan

SwiftMem: Fast Agentic Memory via Query-aware Indexing

Agentic memory systems have become critical for enabling LLM agents to maintain long-term context and retrieve relevant information efficiently. However, existing memory frameworks suffer from a fundamental limitation: they perform exhaustive retrieval across the entire storage layer regardless of query characteristics. This...

💬 0 commentsarXiv:2601.08160v1PDF
0

Posted in cs.CL · 2026-01-13 · Yuqing Zhou, Zhuoer Wang, Jie Yuan, Hong Wang, Samson Koelle, Ziwei Zhu, Wei Niu

WISE-Flow: Workflow-Induced Structured Experience for Self-Evolving Conversational Service Agents

Large language model (LLM)-based agents are widely deployed in user-facing services but remain error-prone in new tasks, tend to repeat the same failure patterns, and show substantial run-to-run variability. Fixing failures via environment-specific training or manual patching is costly and hard to scale. To enable self-evolving agents...

💬 0 commentsarXiv:2601.08158v1PDF
0

Posted in cs.AI · 2026-01-13 · Arin Gopalan Yadav, Varad Dherange, Kumar Shivam

Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions

This paper introduces Project Synapse, a novel agentic framework designed for the autonomous resolution of last-mile delivery disruptions. Synapse employs a hierarchical multi-agent architecture in which a central Resolution Supervisor agent performs strategic task decomposition and delegates subtasks to specialized worker agents...

💬 0 commentsarXiv:2601.08156v1PDF
0

Posted in cs.CV · 2026-01-13 · Inpyo Song, Minjun Joo, Joonhyung Kwon, Eunji Jeon, Jangwon Lee

Instance-Aligned Captions for Explainable Video Anomaly Detection

Explainable video anomaly detection (VAD) is crucial for safety-critical applications, yet even with recent progress, much of the research still lacks spatial grounding, making the explanations unverifiable. This limitation is especially pronounced in multi-entity interactions, where existing explainable VAD methods often produce...

💬 0 commentsarXiv:2601.08155v1PDF
0

Posted in cs.NI · 2026-01-13 · Thakshila Perera, Amine Mezghani, Ekram Hossain

Multi-Objective Optimization for Joint Communication and Sensing in Multi-user MIMO Systems: Characterizing the Pareto Boundary

This paper investigates the Pareto boundary performance of a joint communication and sensing (JCAS) system that addresses both sensing and communication functions at the same time. In this scenario, a multiple-antenna base station (BS) transmits information to multiple single-antenna communication users while concurrently estimating...

💬 0 commentsarXiv:2601.08152v1PDF
0

Posted in cs.CV · 2026-01-13 · Shezheng Song, Shasha Li, Jie Yu

Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet how they internally integrate visual and textual information remains poorly understood. To bridge this gap, we perform a systematic layer-wise masking analysis across multiple architectures, revealing how visual-text fusion...

💬 0 commentsarXiv:2601.08151v1PDF
0

Posted in cs.IR · 2026-01-13 · Seokho Ahn, Sungbok Shin, Young-Duk Seo

Enriching Semantic Profiles into Knowledge Graph for Recommender Systems Using Large Language Models

Rich and informative profiling to capture user preferences is essential for improving recommendation quality. However, there is still no consensus on how best to construct and utilize such profiles. To address this, we revisit recent profiling-based approaches in recommender systems along four dimensions: 1) knowledge base, 2)...

💬 0 commentsarXiv:2601.08148v1PDF
0

Posted in cs.LG · 2026-01-13 · Chaoqun Fei, Huanjiang Liu, Tinglve Zhou, Yangyang Li, Tianyong Hao

Dynamic Graph Structure Learning via Resistance Curvature Flow

Geometric Representation Learning (GRL) aims to approximate the non-Euclidean topology of high-dimensional data through discrete graph structures, grounded in the manifold hypothesis. However, traditional static graph construction methods based on Euclidean distance often fail to capture the intrinsic curvature characteristics of the...

💬 0 commentsarXiv:2601.08149v1PDF
0

Posted in cs.CL · 2026-01-13 · Khumaisa Nur'aini, Ayu Purwarianti, Alham Fikri Aji, Derry Wijaya

Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation

Existing circuit discovery methods rely on templated tasks with clean counterfactuals, limiting their use on diverse natural text. We adapt Contextual Decomposition for Transformers (CD-T) for unstructured settings via label-balanced activation means and task-directional relevance scoring, enabling counterfactual-free circuit...

💬 0 commentsarXiv:2601.08146v3PDF
0

Posted in cs.IT · 2026-01-13 · Junfeng Jia, Yanxun Chang

Cardinality-consistent flag codes with longer type vectors

Flag codes generalize constant dimension codes by considering sequences of nested subspaces with prescribed dimensions as codewords. A comprehensive construction, which unites cyclic orbit flag codes, yields two families of flag codes on $\mathbb{F}^n_q$ (where $n=sk+h$ with $s\geq 2$ and $0\leq h < k$): optimum distance flag codes of...

💬 0 commentsarXiv:2601.08144v1PDF
0

Posted in cs.RO · 2026-01-13 · Takuya Kato, Kentaro Uno, Kazuya Yoshida

A Pin-Array Structure for Gripping and Shape Recognition of Convex and Concave Terrain Profiles

This paper presents a gripper capable of grasping and recognizing terrain shapes for mobile robots in extreme environments. Multi-limbed climbing robots with grippers are effective on rough terrains, such as cliffs and cave walls. However, such robots may fall over by misgrasping the surface or getting stuck owing to the loss of...

💬 0 commentsarXiv:2601.08143v1PDF
0

Posted in cs.NI · 2026-01-13 · Dilki Wijekoon, Amine Mezghani, Ekram Hossain

Joint Communication and Sensing in RIS-Assisted MIMO System Under Mutual Coupling

This paper considers a downlink Reconfigurable Intelligent Surface (RIS)-assisted Joint Communication and Sensing (JCAS) system within a physically-consistent setting, accounting for the effect of mutual coupling between RIS elements arising due to sub-element spacing. The system features a multiple-input multiple-output (MIMO)...

💬 0 commentsarXiv:2601.08142v1PDF
0

Posted in cs.CL · 2026-01-13 · Muhammad Taimoor Hassan, Jawad Ahmed, Muhammad Awais

Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-training

Despite remarkable progress in large language models, Urdu-a language spoken by over 230 million people-remains critically underrepresented in modern NLP systems. Existing multilingual models demonstrate poor performance on Urdu-specific tasks, struggling with the language's complex morphology, right-to-left Nastaliq script, and rich...

💬 0 commentsarXiv:2601.08141v1PDF
0

Posted in cs.CV · 2026-01-13 · Zhichen Zeng, Wenxuan Bao, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Xuying Ning, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, Hanghang Tong

Subspace Alignment for Vision-Language Model Test-time Adaptation

Vision-language models (VLMs), despite their extraordinary zero-shot capabilities, are vulnerable to distribution shifts. Test-time adaptation (TTA) emerges as a predominant strategy to adapt VLMs to unlabeled test data on the fly. However, existing TTA methods heavily rely on zero-shot predictions as pseudo-labels for self-training,...

💬 0 commentsarXiv:2601.08139v1PDF
0

Posted in cs.CY · 2026-01-13 · Carole J. Lee

Critically Engaged Pragmatism: Scientific Norm and Social, Pragmatist Epistemology for AI Science Evaluation Tools

AI science evaluation tools aim to assess research credibility. As with traditional metrics such as impact factors, their edicts can be decontextualised and repurposed in problematic ways. To address this, I propose Critically-Engaged Pragmatism as a scientific norm enjoining scientific communities to scrutinise the purposes and...

💬 0 commentsarXiv:2601.09753v2PDF
0

Posted in cs.LG · 2026-01-13 · Zeyang Li, Sunbochen Tang, Navid Azizan

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies

Diffusion and flow policies are gaining prominence in online reinforcement learning (RL) due to their expressive power, yet training them efficiently remains a critical challenge. A fundamental difficulty that distinguishes online RL from standard generative modeling is the lack of direct samples from the target Boltzmann distribution...

💬 0 commentsarXiv:2601.08136v2PDF
0

Posted in cs.NI · 2026-01-13 · Zengzipeng Tang, Yuxuan Sun, Wei Chen, Jianwen Ding, Bo Ai, Yulin Shao

Hierarchical Online-Scheduling for Energy-Efficient Split Inference with Progressive Transmission

Device-edge collaborative inference with Deep Neural Networks (DNNs) faces fundamental trade-offs among accuracy, latency and energy consumption. Current scheduling exhibits two drawbacks: a granularity mismatch between coarse, task-level decisions and fine-grained, packet-level channel dynamics, and insufficient awareness of per-task...

💬 0 commentsarXiv:2601.08135v1PDF
0

Posted in cs.CL · 2026-01-13 · Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur, Ivan Brugere, Charese H. Smiley, Kundan Thind, Mohammad M. Ghassemi

How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains

The miscalibration of Large Reasoning Models (LRMs) undermines their reliability in high-stakes domains, necessitating methods to accurately estimate the confidence of their long-form, multi-step outputs. To address this gap, we introduce the Reasoning Model Confidence estimation Benchmark (RMCB), a public resource of 347,496...

💬 0 commentsarXiv:2601.08134v2PDF
0

Posted in cs.CV · 2026-01-13 · Yujian Lee, Peng Gao, Yongqi Xu, Wentao Fan

How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?

Audio-visual semantic segmentation (AVSS) represents an extension of the audio-visual segmentation (AVS) task, necessitating a semantic understanding of audio-visual scenes beyond merely identifying sound-emitting objects at the visual pixel level. Contrary to a previous methodology, by decomposing the AVSS task into two discrete...

💬 0 commentsarXiv:2601.08133v2PDF
0

Posted in cs.CL · 2026-01-13 · Jonathan Su

Attention Projection Mixing with Exogenous Anchors

Cross-layer reuse of early attention projections can improve optimization and data efficiency, but it creates a structural conflict: the first layer must simultaneously act as a stable, reusable anchor for all deeper layers and as an effective computational block. We demonstrate that this tension constrains the performance of...

💬 0 commentsarXiv:2601.08131v4PDF
0

Posted in cs.MA · 2026-01-13 · Roland Rodriguez

Emergent Coordination in Multi-Agent Systems via Pressure Fields and Temporal Decay

Current multi-agent LLM frameworks rely on explicit orchestration patterns borrowed from human organizational structures: planners delegate to executors, managers coordinate workers, and hierarchical control flow governs agent interactions. These approaches suffer from coordination overhead that scales poorly with agent count and task...

💬 0 commentsarXiv:2601.08129v3PDF
0

Posted in cs.AI · 2026-01-13 · Rahul Gupta, Stephen D. H. Hsu

Embedded AI Companion System on Edge Devices

Computational resource constraints on edge devices make it difficult to develop a fully embedded AI companion system with a satisfactory user experience. AI companion and memory systems detailed in existing literature cannot be directly used in such an environment due to lack of compute resources and latency concerns. In this paper,...

💬 0 commentsarXiv:2601.08128v1PDF