Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 14, 2026 — 10:49:42 EST

0

Posted in cs.CL · 2026-01-07 · Sangyub Lee, Heedou Kim, Hyeoncheol Kim

Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios

The use of Large Language Models (LLMs) in police operations is growing, yet an evaluation framework tailored to police operations remains absent. While LLM's responses may not always be legally incorrect, their unverified use still can lead to severe issues such as unlawful arrests and improper evidence collection. To address this,...

💬 0 commentsarXiv:2601.03553v1PDF
0

Posted in cs.SI · 2026-01-07 · Lujia Bo, Mingxuan Chen, Youduo Chen, Xiaofan Gui, Jiang Bian, Chunyan Wang, Yi Liu

From Risk Perception to Behavior Large Language Models-Based Simulation of Pandemic Prevention Behaviors

Individual prevention behaviors are a primary line of defense during the early stages of novel infectious disease outbreaks, yet their adoption is heterogeneous and difficult to forecast-especially when empirical data are scarce and epidemic-policy contexts evolve rapidly. To address this gap, we develop an LLM-based...

💬 0 commentsarXiv:2601.03552v1PDF
0

Posted in cs.HC · 2026-01-07 · Michael Yin, Angela Chiang, Robert Xiao

Dissolving a Digital Relationship: A Critical Examination of Digital Severance Behaviours in Close Relationships

Fulfilling social connections are crucial for human well-being and belonging, but not all relationships last forever. As interactions increasingly move online, the act of digitally severing a relationship - e.g. through blocking or unfriending - has become progressively more common as well. This study considers actions of "digital...

💬 0 commentsarXiv:2601.03551v2PDF
0

Posted in cs.AI · 2026-01-07 · Zhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu, Yue Zhang

ReEfBench: Quantifying the Reasoning Efficiency of LLMs

Test-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscures whether performance gains stem from genuine reasoning or mere verbosity. To address this, (1) we propose a novel neuro-symbolic framework for the non-intrusive, comprehensive...

💬 0 commentsarXiv:2601.03550v1PDF
0

Posted in cs.CV · 2026-01-07 · Guobin Tu, Di Weng

FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation

Sign Language Translation (SLT) is a challenging cross-modal task requiring joint modeling of manual articulations and non-manual signals. Existing gloss-free SLT methods effectively capture gestural dynamics but often underutilize facial expressions, which play crucial grammatical and disambiguating roles. This limitation can cause...

💬 0 commentsarXiv:2601.03549v2PDF
0

Posted in cs.CL · 2026-01-07 · Guanyu Chen, Chenxiao Yu, Xiyang Hu

Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concerns and prosocial motivations can push choices in opposite directions. Existing evaluations often measure privacy-related attitudes or sharing intentions in isolation, which makes it difficult to...

💬 0 commentsarXiv:2601.03546v2PDF
0

Posted in cs.CV · 2026-01-07 · Di Xu, Hengjie Liu, Yang Yang, Mary Feng, Jin Ning, Xin Miao, Jessica E. Scholey, Alexandra E. Hotca-cho, William C. Chen, Michael Ohliger, Martina Descovich, Huiming Dong, Wensha Yang, Ke Sheng

B-FIRE: Binning-Free Diffusion Implicit Neural Representation for Hyper-Accelerated Motion-Resolved MRI

Accelerated dynamic volumetric magnetic resonance imaging (4DMRI) is essential for applications relying on motion resolution. Existing 4DMRI produces acceptable artifacts of averaged breathing phases, which can blur and misrepresent instantaneous dynamic information. Recovery of such information requires a new paradigm to reconstruct...

💬 0 commentsarXiv:2601.06166v2PDF
0

Posted in cs.CL · 2026-01-07 · Ye Shen, Dun Pei, Yiqiu Guo, Junying Wang, Yijin Guo, Zicheng Zhang, Qi Jia, Jun Zhou, Guangtao Zhai

EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory

Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across diverse memory dimensions, particularly in multi-session settings. In this work, we propose EvolMem, a new benchmark for assessing multi-session memory...

💬 0 commentsarXiv:2601.03543v1PDF
0

Posted in cs.CL · 2026-01-07 · Xukai Liu, Ye Liu, Jipeng Zhang, Yanghai Zhang, Kai Zhang, Qi Liu

Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language Models

Large language models (LLMs) perform well on multi-hop reasoning, yet how they internally compose multiple facts remains unclear. Recent work proposes \emph{hop-aligned circuit hypothesis}, suggesting that bridge entities are computed sequentially across layers before later-hop answers. Through systematic analyses on real-world...

💬 0 commentsarXiv:2601.03542v1PDF
0

Posted in cs.CL · 2026-01-07 · Jin Cui, Jiaqi Guo, Jiepeng Zhou, Ruixuan Yang, Jiayi Lu, Jiajun Xu, Jiangcheng Song, Boran Zhao, Pengju Ren

MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT Distillation

While Large Language Models (LLMs) have emerged with remarkable capabilities in complex tasks through Chain-of-Thought reasoning, practical resource constraints have sparked interest in transferring these abilities to smaller models. However, achieving both domain performance and cross-domain generalization remains challenging....

💬 0 commentsarXiv:2601.03717v1PDF
0

Posted in cs.CY · 2026-01-07 · François Rottenberg, Thomas Feys, Liesbet Van der Perre

The environmental impact of ICT in the era of data and artificial intelligence

The technology industry promotes artificial intelligence (AI) as a key enabler to solve a vast number of problems, including the environmental crisis. However, when looking at the emissions of datacenters from worldwide service providers, we observe a rapid increase aligned with the advent of AI. Some actors justify it by claiming...

💬 0 commentsarXiv:2601.06174v1PDF
0

Posted in cs.LG · 2026-01-07 · Weijie Shi, Yanxi Chen, Zexi Li, Xuchen Pan, Yuchang Sun, Jiajie Xu, Xiaofang Zhou, Yaliang Li

R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification

Reinforcement learning drives recent advances in LLM reasoning and agentic capabilities, yet current approaches struggle with both exploration and exploitation. Exploration suffers from low success rates on difficult tasks and high costs of repeated rollouts from scratch. Exploitation suffers from coarse credit assignment and training...

💬 0 commentsarXiv:2601.03715v2PDF
0

Posted in cs.CL · 2026-01-07 · Yunhao Liang, Ruixuan Ying, Bo Li, Hong Li, Kai Yan, Qingwen Li, Min Yang, Okamoto Satoshi, Zhe Cui, Shiwen Ni

Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR

DeepSeek-OCR utilizes an optical 2D mapping approach to achieve high-ratio vision-text compression, claiming to decode text tokens exceeding ten times the input visual tokens. While this suggests a promising solution for the LLM long-context bottleneck, we investigate a critical question: "Visual merit or linguistic crutch - which...

💬 0 commentsarXiv:2601.03714v2PDF
0

Posted in cs.CV · 2026-01-07 · Qingyao Tian, Bingyu Yang, Huai Liao, Xinyan Huang, Junyong Li, Dong Yi, Hongbin Liu

BREATH-VL: Vision-Language-Guided 6-DoF Bronchoscopy Localization via Semantic-Geometric Fusion

Vision-language models (VLMs) have recently shown remarkable performance in navigation and localization tasks by leveraging large-scale pretraining for semantic understanding. However, applying VLMs to 6-DoF endoscopic camera localization presents several challenges: 1) the lack of large-scale, high-quality, densely annotated, and...

💬 0 commentsarXiv:2601.03713v1PDF
0

Posted in cs.CR · 2026-01-07 · Ji Guo, Wenbo Jiang, Yansong Lin, Yijing Liu, Ruichen Zhang, Guomin Lu, Aiguo Chen, Xinshuo Han, Hongwei Li

State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space

Vision-Language-Action (VLA) models are widely deployed in safety-critical embodied AI applications such as robotics. However, their complex multimodal interactions also expose new security vulnerabilities. In this paper, we investigate a backdoor threat in VLA models, where malicious inputs cause targeted misbehavior while preserving...

💬 0 commentsarXiv:2601.04266v2PDF
0

Posted in cs.CY · 2026-01-07 · Sarah Spiekermann-Hoff, Marc Langheinrich, Johannes Hoff, Christiane Wendehorst, Jürgen Pfeffer, Thomas Fuchs, Armin Grunwald

The Power of 10: New Rules for the Digital World

As artificial intelligence rapidly advances, society is increasingly captivated by promises of superhuman machines and seamless digital futures. Yet these visions often obscure mounting social, ethical, and psychological concerns tied to pervasive digital technologies - from surveillance to mental health crises. This article argues...

💬 0 commentsarXiv:2601.03709v1PDF
0

Posted in cs.PL · 2026-01-07 · Qingyun Zou, Jiahao Cui, Nuo Chen, Bingsheng He, Weng-Fai Wong

MHRC-Bench: A Multilingual Hardware Repository-Level Code Completion benchmark

Large language models (LLMs) have achieved strong performance on code completion tasks in general-purpose programming languages. However, existing repository-level code completion benchmarks focus almost exclusively on software code and largely overlook hardware description languages. In this work, we present \textbf{MHRC-Bench},...

💬 0 commentsarXiv:2601.03708v2PDF
0

Posted in cs.CL · 2026-01-07 · Hengxing Cai, Yijie Rao, Ligang Huang, Zanyang Zhong, Jinhan Dong, Jingjun Tan, Changhao Nai, Jue Hou, Wenhao Lu, Renxin Zhong

AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions

Existing UAV vision-and-language navigation (VLN) benchmarks rarely provide realistic aerial scenes, natural process-level instructions, and sufficient scale simultaneously, making it difficult to systematically train and evaluate UAV VLN agents under realistic settings. To address this, we propose \textbf{AirNav}, a large-scale...

💬 0 commentsarXiv:2601.03707v2PDF
0

Posted in cs.LG · 2026-01-07 · Gil Shabat

The Geometry of the Pivot: A Note on Lazy Pivoted Cholesky and Farthest Point Sampling

Low-rank approximations of large kernel matrices are ubiquitous in machine learning, particularly for scaling Gaussian Processes to massive datasets. The Pivoted Cholesky decomposition is a standard tool for this task, offering a computationally efficient, greedy low-rank approximation. While its algebraic properties are...

💬 0 commentsarXiv:2601.03706v3PDF
0

Posted in cs.LG · 2026-01-07 · Wajid Arshad Abbasi, Syed Ali Abbas, Maryum Bibi, Saiqa Andleeb, Muhammad Naveed Akhtar

Investigating Knowledge Distillation Through Neural Networks for Protein Binding Affinity Prediction

The trade-off between predictive accuracy and data availability makes it difficult to predict protein--protein binding affinity accurately. The lack of experimentally resolved protein structures limits the performance of structure-based machine learning models, which generally outperform sequence-based methods. In order to overcome...

💬 0 commentsarXiv:2601.03704v1PDF
0

Posted in cs.LG · 2026-01-07 · Lang Cao, Hui Ruan, Yongqian Li, Peng Chao, Wu Ning, Haonan Song, Renhong Chen, Yitong Li

TreeAdv: Tree-Structured Advantage Redistribution for Group-Based RL

Reinforcement learning with group-based objectives, such as Group Relative Policy Optimization (GRPO), is a common framework for aligning large language models on complex reasoning tasks. However, standard GRPO treats each rollout trajectory as an independent flat sequence and assigns a single sequence-level advantage to all tokens,...

💬 0 commentsarXiv:2601.03703v2PDF
0

Posted in cs.MA · 2026-01-07 · Zhilong Tang, Shaohua Wu, Xinyan Zhao, Yu Wang, Xingchu Gong

A Chromatographic Process Design and Optimization Platform Powered by Large Language Models: A Case Application on Extract of Ginkgo Biloba Leaf

Chromatographic separation technology has been widely applied in pharmaceutical, chemical, and food industries due to its high efficiency. However, traditional human-dependent chromatographic process development faces challenges such as reliance on expert experience, long development cycles, and labor intensity. ChromR, a large...

💬 0 commentsarXiv:2601.03702v1PDF
0

Posted in cs.LG · 2026-01-07 · Xiuling Wang, Xin Huang, Guibo Luo, Jianliang Xu

Inference Attacks Against Graph Generative Diffusion Models

Graph generative diffusion models have recently emerged as a powerful paradigm for generating complex graph structures, effectively capturing intricate dependencies and relationships within graph data. However, the privacy risks associated with these models remain largely unexplored. In this paper, we investigate information leakage...

💬 0 commentsarXiv:2601.03701v1PDF
0

Posted in cs.CL · 2026-01-07 · Sangmin Yoo, Srikanth Malla, Chiho Choi, Wei D. Lu, Joon Hee Choi

ADEPT: Adaptive Dynamic Early-Exit Process for Transformers

The inference of large language models imposes significant computational workloads, often requiring the processing of billions of parameters. Although early-exit strategies have proven effective in reducing computational demands by halting inference earlier, they apply either to only the first token in the generation phase or at the...

💬 0 commentsarXiv:2601.03700v1PDF
0

Posted in cs.CL · 2026-01-07 · Quy-Anh Dang, Chris Ngo, Truong-Son Hy

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

As large language models (LLMs) become integral to safety-critical applications, ensuring their robustness against adversarial prompts is paramount. However, existing red teaming datasets suffer from inconsistent risk categorizations, limited domain coverage, and outdated evaluations, hindering systematic vulnerability assessments. To...

💬 0 commentsarXiv:2601.03699v2PDF