Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 8, 2026 — 02:04:42 EST

0

Posted in cs.CV · 2026-01-21 · Ying Yang, Zhengyao Lv, Tianlin Pan, Haofan Wang, Binxin Yang, Hubery Yin, Chen Li, Ziwei Liu, Chenyang Si

StableWorld: Towards Stable and Consistent Long Interactive Video Generation

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as camera movements and text prompts. Despite remarkable progress in world modeling, current methods still suffer from...

💬 0 commentsarXiv:2601.15281v1PDF
0

Posted in cs.HC · 2026-01-21 · Chloe Qianhui Zhao, Jie Cao, Jionghao Lin, Kenneth R. Koedinger

LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback

Providing timely, targeted, and multimodal feedback helps students quickly correct errors, build deep understanding and stay motivated, yet making it at scale remains a challenge. This study introduces a real-time AI-facilitated multimodal feedback system that integrates structured textual explanations with dynamic multimedia...

💬 0 commentsarXiv:2601.15280v1PDF
0

Posted in cs.LG · 2026-01-21 · Christoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl, Günter Klambauer, Sohvi Luukkonen

MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs

A molecule's properties are fundamentally determined by its composition and structure encoded in its molecular graph. Thus, reasoning about molecular properties requires the ability to parse and understand the molecular graph. Large Language Models (LLMs) are increasingly applied to chemistry, tackling tasks such as molecular name...

💬 0 commentsarXiv:2601.15279v1PDF
0

Posted in cs.MM · 2026-01-21 · Mingyue Zha, Ho-Chun Herbert Chang

Interpreting Multimodal Communication at Scale in Short-Form Video: Visual, Audio, and Textual Mental Health Discourse on TikTok

Short-form video platforms integrate text, visuals, and audio into complex communicative acts, yet existing research analyzes these modalities in isolation, lacking scalable frameworks to interpret their joint contributions. This study introduces a pipeline combining automated multimodal feature extraction with Shapley value-based...

💬 0 commentsarXiv:2601.15278v1PDF
0

Posted in cs.CL · 2026-01-21 · Sahar Tahmasebi, Eric Müller-Budack, Ralph Ewerth

Robust Fake News Detection using Large Language Models under Adversarial Sentiment Attacks

Misinformation and fake news have become a pressing societal challenge, driving the need for reliable automated detection methods. Prior research has highlighted sentiment as an important signal in fake news detection, either by analyzing which sentiments are associated with fake news or by using sentiment and emotion features for...

💬 0 commentsarXiv:2601.15277v1PDF
0

Posted in cs.CV · 2026-01-21 · Yu Wu, Minsik Jeon, Jen-Hao Rick Chang, Oncel Tuzel, Shubham Tulsiani

RayRoPE: Projective Ray Positional Encoding for Multi-view Attention

We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can adapt to the geometry of the underlying 3D scene. We find that prior (absolute or relative)...

💬 0 commentsarXiv:2601.15275v3PDF
0

Posted in cs.LG · 2026-01-21 · Maciej Kilian, Oleg Mkrtchyan, Luke Zettlemoyer, Akshat Shrivastava, Armen Aghajanyan

Improving MoE Compute Efficiency by Composing Weight and Data Sparsity

Mixture-of-Experts layers achieve compute efficiency through weight sparsity: each token activates only a subset of experts. Data sparsity, where each expert processes only a subset of tokens, offers a complementary axis. Expert-choice routing implements data sparsity directly but violates causality in autoregressive models, creating...

💬 0 commentsarXiv:2601.15370v1PDF
0

Posted in cs.RO · 2026-01-21 · Heng Zhang, Wei-Hsing Huang, Qiyi Tong, Gokhan Solak, Puze Liu, Kaidi Zhang, Sheng Liu, Jan Peters, Yu She, Arash Ajoudani

CompliantVLA-adaptor: VLM-Guided Variable Impedance Action for Safe Contact-Rich Manipulation

We propose a CompliantVLA-adaptor that augments the state-of-the-art Vision-Language-Action (VLA) models with vision-language model (VLM)-informed context-aware variable impedance control (VIC) to improve the safety and effectiveness of contact-rich robotic manipulation tasks. Existing VLA systems (e.g., RDT, Pi0.5, OpenVLA-oft)...

💬 0 commentsarXiv:2601.15541v2PDF
0

Posted in cs.LG · 2026-01-21 · Dongchen Huang

PRISM: Deriving a White-Box Transformer as a Signal-Noise Decomposition Operator via Maximum Coding Rate Reduction

Deep learning models, particularly Transformers, are often criticized as "black boxes" and lack interpretability. We propose Prism, a white-box attention-based architecture derived from the principles of Maximizing Coding Rate Reduction ($\text{MCR}^2$). By modeling the attention mechanism as a gradient ascent process on a distinct...

💬 0 commentsarXiv:2601.15540v2PDF
0

Posted in cs.LG · 2026-01-21 · Himanshu Mishra, Kanwal Mehreen

QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs

Machine unlearning aims to remove specific knowledge (e.g., copyrighted or private data) from a trained model without full retraining. In practice, models are often quantized (e.g., 4-bit) for deployment, but we find that quantization can catastrophically restore forgotten information [1]. In this paper, we (1) analyze why low-bit...

💬 0 commentsarXiv:2601.15538v1PDF
0

Posted in cs.AI · 2026-01-21 · Zhikang Chen, Tingting Zhu

From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models

A world model is an AI system that simulates how an environment evolves under actions, enabling planning through imagined futures rather than reactive perception. Current world models, however, suffer from visual conflation: the mistaken assumption that high-fidelity video generation implies an understanding of physical and causal...

💬 0 commentsarXiv:2601.15533v1PDF
0

Posted in cs.NI · 2026-01-21 · Abd Ullah Khan, Wali Ullah Khan, Haejoon Jung, Hyundong Shin

Resource Allocation and Sharing for UAV-Assisted Integrated TN-NTN with Multi-Connectivity

Unmanned aerial vehicles (UAVs) with multi-connectivity (MC) capabilities efficiently and reliably transfer data between terrestrial networks (TNs) and non-terrestrial networks (NTNs). However, optimally sharing and allocating spectrum and power resources to maintain MC while ensuring reliable connectivity and optimal performance...

💬 0 commentsarXiv:2601.15532v2PDF
0

Posted in cs.LG · 2026-01-21 · Megan A. Witherow, Michael L. Evans, Ahmed Temtam, Hamid R. Okhravi, Khan M. Iftekharuddin

Machine learning-enhanced non-amnestic Alzheimer's disease diagnosis from MRI and clinical features

Alzheimer's disease (AD), defined as an abnormal buildup of amyloid plaques and tau tangles in the brain can be diagnosed with high accuracy based on protein biomarkers via PET or CSF analysis. However, due to the invasive nature of biomarker collection, most AD diagnoses are made in memory clinics using cognitive tests and evaluation...

💬 0 commentsarXiv:2601.15530v2PDF
0

Posted in cs.CL · 2026-01-21 · Himanshu Gupta, Pratik Jayarao, Chaitanya Dwivedi, Neeraj Varshney

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

Code-mixing and code-switching (CSW) remain challenging phenomena for large language models (LLMs). Despite recent advances in multilingual modeling, LLMs often struggle in mixed-language settings, exhibiting systematic degradation in grammaticality, factuality, and safety behavior. This work provides a comprehensive overview of CSW...

💬 0 commentsarXiv:2602.11181v2PDF
0

Posted in cs.DC · 2026-01-21 · Jiazhu Xie, Bowen Li, Heyu Fu, Chong Gao, Ziqi Xu, Fengling Han

Securing LLM-as-a-Service for Small Businesses: An Industry Case Study of a Distributed Chatbot Deployment Platform

Large Language Model (LLM)-based question-answering systems offer significant potential for automating customer support and internal knowledge access in small businesses, yet their practical deployment remains challenging due to infrastructure costs, engineering complexity, and security risks, particularly in retrieval-augmented...

💬 0 commentsarXiv:2601.15528v1PDF
0

Posted in cs.CL · 2026-01-21 · Inwon Kang, Parikshit Ram, Yi Zhou, Horst Samulowitz, Oshani Seneviratne

Language Model Representations for Efficient Few-Shot Tabular Classification

The Web is a rich source of structured data in the form of tables, from product catalogs and knowledge bases to scientific datasets. However, the heterogeneity of the structure and semantics of these tables makes it challenging to build a unified method that can effectively leverage the information they contain. Meanwhile, Large...

💬 0 commentsarXiv:2602.15844v1PDF
0

Posted in cs.AI · 2026-01-21 · Zhichao Yang, Jiashu He, Jinxuan Fan, Cirillo Cinzia

TransportAgents: a multi-agents LLM framework for traffic accident severity prediction

Accurate prediction of traffic crash severity is critical for improving emergency response and public safety planning. Although recent large language models (LLMs) exhibit strong reasoning capabilities, their single-agent architectures often struggle with heterogeneous, domain-specific crash data and tend to generate biased or...

💬 0 commentsarXiv:2601.15519v2PDF
0

Posted in cs.IR · 2026-01-21 · Wenxin Zhou, Ritesh Mehta, Anthony Miyaguchi

DS@GT at TREC TOT 2025: Bridging Vague Recollection with Fusion Retrieval and Learned Reranking

We develop a two-stage retrieval system that combines multiple complementary retrieval methods with a learned reranker and LLM-based reranking, to address the TREC Tip-of-the-Tongue (ToT) task. In the first stage, we employ hybrid retrieval that merges LLM-based retrieval, sparse (BM25), and dense (BGE-M3) retrieval methods. We also...

💬 0 commentsarXiv:2601.15518v2PDF
0

Posted in cs.CV · 2026-01-21 · William Huang, Siyou Pei, Leyi Zou, Eric J. Gonzalez, Ishan Chatterjee, Yang Zhang

DeltaDorsal: Enhancing Hand Pose Estimation with Dorsal Features in Egocentric Views

The proliferation of XR devices has made egocentric hand pose estimation a vital task, yet this perspective is inherently challenged by frequent finger occlusions. To address this, we propose a novel approach that leverages the rich information in dorsal hand skin deformation, unlocked by recent advances in dense visual featurizers....

💬 0 commentsarXiv:2601.15516v2PDF
0

Posted in cs.CR · 2026-01-21 · Marcell Szakály, Martin Strohmeier, Ivan Martinovic, Sebastian Köhler

DCeption: Real-world Wireless Man-in-the-Middle Attacks Against CCS EV Charging

The adoption of Electric Vehicles (EVs) is happening at a rapid pace. To ensure fast and safe charging, complex communication is required between the vehicle and the charging station. In the globally used Combined Charging System (CCS), this communication is carried over the HomePlug Green PHY (HPGP) physical layer. However, HPGP is...

💬 0 commentsarXiv:2601.15515v1PDF
0

Posted in cs.CL · 2026-01-21 · Adam Szelestey, Sofie van Engelen, Tianhao Huang, Justin Snelders, Qintao Zeng, Songgaojun Deng

AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains

Hallucination in large language models (LLMs) remains an acute concern, contributing to the spread of misinformation and diminished public trust, particularly in high-risk domains. Among hallucination types, factuality is crucial, as it concerns a model's alignment with established world knowledge. Adversarial factuality, defined as...

💬 0 commentsarXiv:2601.15511v1PDF
0

Posted in cs.AI · 2026-01-21 · Prasanna Kumar

The Dark Side of AI Transformers: Sentiment Polarization & the Loss of Business Neutrality by NLP Transformers

The use of Transfer Learning & Transformers has steadily improved accuracy and has significantly contributed in solving complex computation problems. However, this transformer led accuracy improvement in Applied AI Analytics specifically in sentiment analytics comes with the dark side. It is observed during experiments that a lot of...

💬 0 commentsarXiv:2601.15509v1PDF
0

Posted in cs.SE · 2026-01-21 · Li Huang, Bertrand Meyer, Manuel Oriol

Combining Tests and Proofs for Better Software Verification

Test or prove? These two approaches to software verification have long been presented as opposites. One is dynamic, the other static: a test executes the program, a proof only analyzes the program text. A different perspective is emerging, in which testing and proving are complementary rather than competing techniques for producing...

💬 0 commentsarXiv:2601.16239v2PDF
0

Posted in cs.CL · 2026-01-21 · Haaris Mian, Melanie Subbiah, Sharon Marcus, Nora Shaalan, Kathleen McKeown

Computational Representations of Character Significance in Novels

Characters in novels have typically been modeled based on their presence in scenes in narrative, considering aspects like their actions, named mentions, and dialogue. This conception of character places significant emphasis on the main character who is present in the most scenes. In this work, we instead adopt a framing developed from...

💬 0 commentsarXiv:2601.15508v1PDF
0

Posted in cs.CV · 2026-01-21 · Jinrui Yang, Qing Liu, Yijun Li, Mengwei Ren, Letian Zhang, Zhe Lin, Cihang Xie, Yuyin Zhou

A Unified and Controllable Framework for Layered Image Generation with Visual Effects

Recent image generation models produce impressive composites, but often fail to preserve the identity of user-provided content when editing specific elements: the surrounding scene may shift, and even the edited object's appearance can drift from the original. Layered representation offer a natural remedy--they allow users to...

💬 0 commentsarXiv:2601.15507v2PDF