Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 14, 2026 — 09:19:58 EST

0

Posted in cs.CL · 2026-01-07 · Hui Huang, Xuanxin Wu, Muyun Yang, Yuki Arase

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases

This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. Our empirical analysis yields four key findings: 1) LRMs outperform non-reasoning LLMs in terms of judgment accuracy, particularly on reasoning-intensive tasks; 2) LRMs demonstrate superior...

💬 0 commentsarXiv:2601.03630v2PDF
0

Posted in cs.LG · 2026-01-07 · Dmytro Matsypura, Yu Pan, Hanzhao Wang

Learning Shortest Paths When Data is Scarce

Digital twins and other simulators are increasingly used to support routing decisions in large-scale networks. However, simulator outputs often exhibit systematic bias, while ground-truth measurements are costly and scarce. We study a stochastic shortest-path problem in which a planner has access to abundant synthetic samples, limited...

💬 0 commentsarXiv:2601.03629v1PDF
0

Posted in cs.DL · 2026-01-07 · Muneer Ahmad, Undie Felicia Nkatv, Sajid Saleem

Global research trends and collaborations in Fibrodysplasia Ossificans Progressiva: A bibliometric analysis (1989-2023)

Fibrodysplasia Ossificans Progressiva (FOP) is a rare and debilitating genetic disorder characterized by the progressive formation of bone in muscles and connective tissues. This scientometric analysis examines the global research trends on FOP between 1989 and 2023 using bibliographic data from Web of Science. The study highlights...

💬 0 commentsarXiv:2601.03628v1PDF
0

Posted in cs.CL · 2026-01-07 · Jean Seo, Gibaeg Kim, Kihun Shin, Seungseop Lim, Hyunkyung Lee, Wooseok Han, Jongwon Lee, Eunho Yang

Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines

We introduce EPAG, a benchmark dataset and framework designed for Evaluating the Pre-consultation Ability of LLMs using diagnostic Guidelines. LLMs are evaluated directly through HPI-diagnostic guideline comparison and indirectly through disease diagnosis. In our experiments, we observe that small open-source models fine-tuned with a...

💬 0 commentsarXiv:2601.03627v3PDF
0

Posted in cs.AI · 2026-01-07 · Zoran Milosevic, Fethi Rabhi

Architecting Agentic Communities using Design Patterns

The rapid evolution of Large Language Models (LLM) and subsequent Agentic AI technologies requires systematic architectural guidance for building sophisticated, production-grade systems. This paper presents an approach for architecting such systems using design patterns derived from enterprise distributed systems standards, formal...

💬 0 commentsarXiv:2601.03624v3PDF
0

Posted in cs.LG · 2026-01-07 · Wang Cai, Yilin Wen, Jinchang Hou, Du Su, Guoqiu Wang, Zhonghou Lv, Chenfu Bao, Yunfang Wu

Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis

Safety alignment in Large Language Models (LLMs) inherently presents a multi-objective optimization conflict, often accompanied by an unintended degradation of general capabilities. Existing mitigation strategies typically rely on global gradient geometry to resolve these conflicts, yet they overlook Modular Heterogeneity within...

💬 0 commentsarXiv:2601.04262v1PDF
0

Posted in cs.CR · 2026-01-07 · Hang Fu, Wanli Peng, Yinghan Zhou, Jiaxuan Wu, Juan Wen, Yiming Xue

Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models

The widespread adoption of Large Language Model (LLM) in commercial and research settings has intensified the need for robust intellectual property protection. Backdoor-based LLM fingerprinting has emerged as a promising solution for this challenge. In practical application, the low-cost multi-model collaborative technique, LLM...

💬 0 commentsarXiv:2601.04261v1PDF
0

Posted in cs.SE · 2026-01-07 · Verya Monjezi, Ashish Kumar, Ashutosh Trivedi, Gang Tan, Saeid Tizpaz-Niari

On the Robustness of Fairness Practices: A Causal Framework for Systematic Evaluation

Machine learning (ML) algorithms are increasingly deployed to make critical decisions in socioeconomic applications such as finance, criminal justice, and autonomous driving. However, due to their data-driven and pattern-seeking nature, ML algorithms may develop decision logic that disproportionately distributes opportunities,...

💬 0 commentsarXiv:2601.03621v1PDF
0

Posted in cs.LG · 2026-01-07 · Anshum Rankawat

Parent-Guided Adaptive Reliability (PGAR): A Behavioural Meta-Learning Framework for Stable and Trustworthy AI

Parent-Guided Adaptive Reliability (PGAR) is a lightweight behavioural meta-learning framework that adds a supervisory "parent" layer on top of a standard learner to improve stability, calibration, and recovery under disturbances. PGAR computes three reflex-level signals (incident detection, overconfidence correction, and recovery...

💬 0 commentsarXiv:2601.06167v1PDF
0

Posted in cs.DB · 2026-01-07 · Muhammad Imam Luthfi Balaka, Raul Castro Fernandez

The Pneuma Project: Reifying Information Needs as Relational Schemas to Automate Discovery, Guide Preparation, and Align Data with Intent

Data discovery and preparation remain persistent bottlenecks in the data management lifecycle, especially when user intent is vague, evolving, or difficult to operationalize. The Pneuma Project introduces Pneuma-Seeker, a system that helps users articulate and fulfill information needs through iterative interaction with a language...

💬 0 commentsarXiv:2601.03618v1PDF
0

Posted in cs.CV · 2026-01-07 · Samson Oseiwe Ajadalu

Systematic Evaluation of Depth Backbones and Semantic Cues for Monocular Pseudo-LiDAR 3D Detection

Monocular 3D object detection offers a low-cost alternative to LiDAR, yet remains less accurate due to the difficulty of estimating metric depth from a single image. We systematically evaluate how depth backbones and feature engineering affect a monocular Pseudo-LiDAR pipeline on the KITTI validation split. Specifically, we compare...

💬 0 commentsarXiv:2601.03617v1PDF
0

Posted in cs.CL · 2026-01-07 · Binh Nguyen, Charles Fleming, Thai Le

SARA: Stress Test Reasoning in Audio Deepfake Detection

Audio Language Models (ALMs) offer a promising shift towards explainable audio deepfake detections (ADD), moving beyond \textit{black-box} classifiers by providing transparency to their predictions via reasoning traces. However, such reasoning may not support the model predictions, reflecting poor coherence, or, worse, may rationalize...

💬 0 commentsarXiv:2601.03615v2PDF
0

Posted in cs.LG · 2026-01-07 · Joonwon Seo

Mathematical Foundations of Polyphonic Music Generation via Structural Inductive Bias

This monograph addresses the "Missing Middle" problem in AI music generation - the challenge of producing coherent, phrase-level musical structure. Using Beethoven's piano sonatas as a case study, I introduce the Smart Embedding architecture, a factorized representation grounded in the empirically verified independence of pitch and...

💬 0 commentsarXiv:2601.03612v8PDF
0

Posted in cs.SD · 2026-01-07 · Nithinkumar K., Anand R

Investigation into respiratory sound classification for an imbalanced data set using hybrid LSTM-KAN architectures

Respiratory sounds captured via auscultation contain critical clues for diagnosing pulmonary conditions. Automated classification of these sounds faces challenges due to subtle acoustic differences and severe class imbalance in clinical datasets. This study investigates respiratory sound classification with a focus on mitigating...

💬 0 commentsarXiv:2601.03610v1PDF
0

Posted in cs.CV · 2026-01-07 · Pratyush Jena, Amal Joseph, Arnav Sharma, Ravi Kiran Sarvadevabhatla

Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization

Binarization is a popular first step towards text extraction in historical artifacts. Stone inscription images pose severe challenges for binarization due to poor contrast between etched characters and the stone background, non-uniform surface degradation, distracting artifacts, and highly variable text density and layouts. These...

💬 0 commentsarXiv:2601.03609v1PDF
0

Posted in cs.RO · 2026-01-07 · Tae Hoon Yang, Haochen Shi, Jiacheng Hu, Zhicong Zhang, Daniel Jiang, Weizhuo Wang, Yao He, Zhen Wu, Yuming Chen, Yifan Hou, Monroe Kennedy, Shuran Song, C. Karen Liu

Locomotion Beyond Feet

Most locomotion methods for humanoid robots focus on leg-based gaits, yet natural bipeds frequently rely on hands, knees, and elbows to establish additional contacts for stability and support in complex environments. This paper introduces Locomotion Beyond Feet, a comprehensive system for whole-body humanoid locomotion across...

💬 0 commentsarXiv:2601.03607v1PDF
0

Posted in cs.LG · 2026-01-07 · Sumedh Pendurkar, Guni Sharon

Policy-Guided Search on Tree-of-Thoughts for Efficient Problem Solving with Bounded Language Model Queries

Recent studies explored integrating state-space search algorithms with Language Models (LM) to perform look-ahead on the token generation process, the ''Tree-of-Thoughts'' (ToT), generated by LMs, thereby improving performance on problem-solving tasks. However, the affiliated search algorithms often overlook the significant...

💬 0 commentsarXiv:2601.03606v1PDF
0

Posted in cs.CL · 2026-01-07 · Hui Huang, Muyun Yang, Yuki Arase

DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier

Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, fueling growing interest in factuality verification. Existing research on factuality verification primarily conducts binary judgments (e.g., correct or incorrect), which fails to distinguish varying degrees of error...

💬 0 commentsarXiv:2601.03605v1PDF
0

Posted in cs.AI · 2026-01-07 · Chuanliu Fan, Zicheng Ma, Huanran Meng, Aijia Zhang, Wenjie Du, Jun Zhang, Yi Qin Gao, Ziqiang Cao, Guohong Fu

Interleaved Tool-Call Reasoning for Protein Function Understanding

Recent advances in large language models (LLMs) have highlighted the effectiveness of chain-of-thought reasoning in symbolic domains such as mathematics and programming. However, our study shows that directly transferring such text-based reasoning paradigms to protein function understanding is ineffective: reinforcement learning...

💬 0 commentsarXiv:2601.03604v2PDF
0

Posted in cs.LG · 2026-01-07 · Kaidong Feng, Zhu Sun, Roy Ka-Wei Lee, Xun Jiang, Yin-Leng Theng, Yi Ding

A Comparative Study of Traditional Machine Learning, Deep Learning, and Large Language Models for Mental Health Forecasting using Smartphone Sensing Data

Smartphone sensing offers an unobtrusive and scalable way to track daily behaviors linked to mental health, capturing changes in sleep, mobility, and phone use that often precede symptoms of stress, anxiety, or depression. While most prior studies focus on detection that responds to existing conditions, forecasting mental health...

💬 0 commentsarXiv:2601.03603v2PDF
0

Posted in cs.LG · 2026-01-07 · Xiao Lin, Philip Li, Zhichen Zeng, Tingwei Li, Tianxin Wei, Xuying Ning, Gaotang Li, Yuzhong Chen, Hanghang Tong

ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification

Despite rich safety alignment strategies, large language models (LLMs) remain highly susceptible to jailbreak attacks, which compromise safety guardrails and pose serious security risks. Existing detection methods mainly detect jailbreak status relying on jailbreak templates present in the training data. However, few studies address...

💬 0 commentsarXiv:2601.03600v1PDF
0

Posted in cs.CL · 2026-01-07 · Yingjian Chen, Haoran Liu, Yinhong Liu, Sherry T. Tong, Aosong Feng, Jinghui Lu, Juntao Zhang, Yusuke Iwasawa, Yutaka Matsuo, Irene Li

From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs

Large Language Models (LLMs) show strong reasoning ability in open-domain question answering, yet their reasoning processes are typically linear and often logically inconsistent. In contrast, real-world reasoning requires integrating multiple premises and solving subproblems in parallel. Existing methods, such as Chain-of-Thought...

💬 0 commentsarXiv:2601.03597v2PDF
0

Posted in cs.CV · 2026-01-07 · Qianyu Guo, Jingrong Wu, Jieji Ren, Weifeng Ge, Wenqiang Zhang

Adaptive Attention Distillation for Robust Few-Shot Segmentation under Environmental Perturbations

Few-shot segmentation (FSS) aims to rapidly learn novel class concepts from limited examples to segment specific targets in unseen images, and has been widely applied in areas such as medical diagnosis and industrial inspection. However, existing studies largely overlook the complex environmental factors encountered in real world...

💬 0 commentsarXiv:2601.03596v3PDF
0

Posted in cs.AI · 2026-01-07 · Yi Fang, Wenjie Wang, Mingfeng Xue, Boyi Deng, Fengli Xu, Dayiheng Liu, Fuli Feng

Controllable LLM Reasoning via Sparse Autoencoder-Based Steering

Large Reasoning Models (LRMs) exhibit human-like cognitive reasoning strategies (e.g. backtracking, cross-verification) during reasoning process, which improves their performance on complex tasks. Currently, reasoning strategies are autonomously selected by LRMs themselves. However, such autonomous selection often produces inefficient...

💬 0 commentsarXiv:2601.03595v1PDF