Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 14, 2026 — 13:00:40 EST

0

Posted in cs.LG · 2026-01-07 · Arpad Berta, Gabor Danner, Istvan Hegedus, Mark Jelasity

Detecting Semantic Backdoors in a Mystery Shopping Scenario

Detecting semantic backdoors in classification models--where some classes can be activated by certain natural, but out-of-distribution inputs--is an important problem that has received relatively little attention. Semantic backdoors are significantly harder to detect than backdoors that are based on trigger patterns due to the lack of...

💬 0 commentsarXiv:2601.03805v1PDF
0

Posted in cs.LG · 2026-01-07 · Rehan Ahmad, Muhammad Kashif, Nouhaila Innan, Muhammad Shafique

Quantum vs. Classical Machine Learning: A Benchmark Study for Financial Prediction

In this paper, we present a reproducible benchmarking framework that systematically compares QML models with architecture-matched classical counterparts across three financial tasks: (i) directional return prediction on U.S. and Turkish equities, (ii) live-trading simulation with Quantum LSTMs versus classical LSTMs on the S\&P 500,...

💬 0 commentsarXiv:2601.03802v1PDF
0

Posted in cs.CL · 2026-01-07 · Taisiia Tikhomirova, Dirk U. Wulff

Where meaning lives: Layer-wise accessibility of psycholinguistic features in encoder and decoder language models

Understanding where transformer language models encode psychologically meaningful aspects of meaning is essential for both theory and practice. We conduct a systematic layer-wise probing study of 58 psycholinguistic features across 10 transformer models, spanning encoder-only and decoder-only architectures, and compare three embedding...

💬 0 commentsarXiv:2601.03798v1PDF
0

Posted in cs.LG · 2026-01-07 · Sethupathy Parameswaran, Suresh Sundaram, Yuan Fang

Prompt Tuning without Labeled Samples for Zero-Shot Node Classification in Text-Attributed Graphs

Node classification is a fundamental problem in information retrieval with many real-world applications, such as community detection in social networks, grouping articles published online and product categorization in e-commerce. Zero-shot node classification in text-attributed graphs (TAGs) presents a significant challenge,...

💬 0 commentsarXiv:2601.03793v1PDF
0

Posted in cs.CL · 2026-01-07 · Huynh Trung Kiet, Dao Sy Duy Minh, Nguyen Dinh Ha Duong, Le Hoang Minh Huy, Long Nguyen, Dien Dinh

VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation

Large Language Models (LLMs) have demonstrated remarkable proficiency in general medical domains. However, their performance significantly degrades in specialized, culturally specific domains such as Vietnamese Traditional Medicine (VTM), primarily due to the scarcity of high-quality, structured benchmarks. In this paper, we introduce...

💬 0 commentsarXiv:2601.03792v2PDF
0

Posted in cs.CL · 2026-01-07 · Xiaoyu Luo, Yiyi Chen, Qiongxiu Li, Johannes Bjerva

Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework

Large Language Models (LLMs) have been reported to "leak" Personally Identifiable Information (PII), with successful PII reconstruction often interpreted as evidence of memorization. We propose a principled revision of memorization evaluation for LLMs, arguing that PII leakage should be evaluated under low lexical cue conditions,...

💬 0 commentsarXiv:2601.03791v1PDF
0

Posted in cs.CL · 2026-01-07 · Zhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa Tsuruoka

NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning

Neologism-aware machine translation aims to translate source sentences containing neologisms into target languages. This field remains underexplored compared with general machine translation (MT). In this paper, we propose an agentic framework, NeoAMT, for neologism-aware machine translation equipped with a Wiktionary-based search...

💬 0 commentsarXiv:2601.03790v4PDF
0

Posted in cs.CY · 2026-01-07 · Anamaria Mojica-Hanke, Thomas Goger, Svenja Wölfel, Brian Valerius, Steffen Herbold

Criminal Liability of Generative Artificial Intelligence Providers for User-Generated Child Sexual Abuse Material

The development of more powerful Generative Artificial Intelligence (GenAI) has expanded its capabilities and the variety of outputs. This has introduced significant legal challenges, including gray areas in various legal systems, such as the assessment of criminal liability for those responsible for these models. Therefore, we...

💬 0 commentsarXiv:2601.03788v1PDF
0

Posted in cs.CL · 2026-01-07 · Loris Schoenegger, Benjamin Roth

Compact Example-Based Explanations for Language Models

Training data influence estimation methods quantify the contribution of training documents to a model's output, making them a promising source of information for example-based explanations. As humans cannot interpret thousands of documents, only a small subset of the training data can be presented as an explanation. Although the...

💬 0 commentsarXiv:2601.03786v2PDF
0

Posted in cs.CL · 2026-01-07 · Dehao Tao, Guoliang Ma, Yongfeng Huang, Minghu Jiang

Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agents

Long-term human-agent dialogues are organized by topic continuity: adjacent turns often develop the same goal, plan, problem, or event, while related activities may recur across distant sessions. Yet many LLM agent memory systems first decompose histories into isolated turns or fixed-size chunks, then compensate through enrichment,...

💬 0 commentsarXiv:2601.03785v3PDF
0

Posted in cs.CV · 2026-01-07 · Steven Moonen, Rob Salaets, Kenneth Batstone, Abdellatif Bey-Temsamani, Nick Michiels

A Comparative Study of 3D Model Acquisition Methods for Synthetic Data Generation of Agricultural Products

In the manufacturing industry, computer vision systems based on artificial intelligence (AI) are widely used to reduce costs and increase production. Training these AI models requires a large amount of training data that is costly to acquire and annotate, especially in high-variance, low-volume manufacturing environments. A popular...

💬 0 commentsarXiv:2601.03784v1PDF
0

Posted in cs.CL · 2026-01-07 · Jin Wang, Liang Lin, Kaiwen Luo, Weiliu Wang, Yitian Chen, Moayad Aloqaily, Xuehai Tang, Zhenhong Zhou, Kun Wang, Li Sun, Qingsong Wen

HearSay Benchmark: Do Audio LLMs Leak What They Hear?

While Audio Large Language Models (ALLMs) have achieved remarkable progress in understanding and generation, their potential privacy implications remain largely unexplored. This paper takes the first step to investigate whether ALLMs inadvertently leak user privacy solely through acoustic voiceprints and introduces $\textit{HearSay}$,...

💬 0 commentsarXiv:2601.03783v1PDF
0

Posted in cs.RO · 2026-01-07 · Wenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu, Dieter Fox, Kaichun Mo, Li Fei-Fei

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre-trained 3D world model that unifies state and action in a shared 3D space as 3D point flows: given one or few RGB-D images and a...

💬 0 commentsarXiv:2601.03782v1PDF
0

Posted in cs.CV · 2026-01-07 · Xiaokun Sun, Zezhong Wu, Zewen Ding, Linli Xu

MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction

Reinforcement learning based post-training paradigms for Video Large Language Models (VideoLLMs) have achieved significant success by optimizing for visual-semantic tasks such as captioning or VideoQA. However, while these approaches effectively enhance perception abilities, they primarily target holistic content understanding, often...

💬 0 commentsarXiv:2601.03781v1PDF
0

Posted in cs.SE · 2026-01-07 · Md Ahasanuzzaman, Bram Adams, Emad Fallahzadeh, Gustavo A. Oliva, Ahmed E. Hassan

Assessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical Study

Large Language Models (LLMs) such as GPT-4, Claude and LLaMA have shown impressive performance in code generation, typically evaluated using benchmarks (e.g., HumanEval). However, effective code generation requires models to understand and apply a wide range of language concepts. If the concepts exercised in benchmarks are not...

💬 0 commentsarXiv:2601.03780v1PDF
0

Posted in cs.CL · 2026-01-07 · Marco Baroni, Emily Cheng, Iria de-Dios-Flores, Francesca Franzon

Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations

We explore intrinsic dimension (ID) of LLM representations as a marker of linguistic complexity. Specifically, we test whether ID differences across model layers reflect well-known complexity contrasts established in (psycho)linguistics: coordination vs. subordination, right-branching vs. center-embedding, and unambiguous vs....

💬 0 commentsarXiv:2601.03779v2PDF
0

Posted in cs.LG · 2026-01-07 · Sebastian Müller, Tobias Schneider, Ruben Kemna, Vanessa Toborek

Improving Compactness and Reducing Ambiguity of CFIRE Rule-Based Explanations

Models trained on tabular data are widely used in sensitive domains, increasing the demand for explanation methods to meet transparency needs. CFIRE is a recent algorithm in this domain that constructs compact surrogate rule models from local explanations. While effective, CFIRE may assign rules associated with different classes to...

💬 0 commentsarXiv:2601.03776v1PDF
0

Posted in cs.CL · 2026-01-07 · Pingjun Hong, Benjamin Roth

Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations

Large Language Models (LLMs) can produce verbalized self-explanations, yet prior studies suggest that such rationales may not reliably reflect the model's true decision process. We ask whether these explanations nevertheless help users predict model behavior, operationalized as counterfactual simulatability. Using StrategyQA, we...

💬 0 commentsarXiv:2601.03775v1PDF
0

Posted in cs.AI · 2026-01-07 · Zihang Li, Yuhang Wang, Yikun Zong, Wenhan Yu, Xiaokun Yuan, Runhan Jiang, Zirui Liu, Tong Yang, Arthur Jiang

EntroCoT: Enhancing Chain-of-Thought via Adaptive Entropy-Guided Segmentation

Chain-of-Thought (CoT) prompting has significantly enhanced the mathematical reasoning capabilities of Large Language Models. We find existing fine-tuning datasets frequently suffer from the "answer right but reasoning wrong" probelm, where correct final answers are derived from hallucinated, redundant, or logically invalid...

💬 0 commentsarXiv:2601.03769v3PDF
0

Posted in cs.PL · 2026-01-07 · Yichen Xu, Martin Odersky

Agentic Proof Automation: A Case Study

Proof engineering is notoriously labor-intensive: proofs that are straightforward on paper often require lengthy scripts in theorem provers. Recent advances in large language models (LLMs) create new opportunities for proof automation: modern LLMs not only generate proof scripts, but also support agentic behavior, exploring codebases...

💬 0 commentsarXiv:2601.03768v1PDF
0

Posted in cs.LG · 2026-01-07 · Noam Levi

Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model

We analyze neural scaling laws in a solvable model of last-layer fine-tuning where targets have intrinsic, instance-heterogeneous difficulty. In our Latent Instance Difficulty (LID) model, each input's target variance is governed by a latent ``precision'' drawn from a heavy-tailed distribution. While generalization loss recovers...

💬 0 commentsarXiv:2601.03764v1PDF
0

Posted in cs.NI · 2026-01-07 · Nguyen Cong Luong, Zeping Sui, Duc Van Le, Jie Cao, Bo Ma, Nguyen Duc Hai, Ruichen Zhang, Vu Van Quang, Dusit Niyato, Shaohan Feng

Incentive Mechanism Design for Resource Management in Satellite Networks: A Comprehensive Survey

Resource management is one of the challenges in satellite networks due to their high mobility, wide coverage, long propagation distances, and stringent constraints on energy, communication, and computation resources. Traditional resource allocation approaches rely only on hard and rigid system performance metrics. Meanwhile, incentive...

💬 0 commentsarXiv:2601.03757v1PDF
0

Posted in cs.LG · 2026-01-07 · Paulius Rauba, Viktor Cikojevic, Fran Bartolic, Sam Levang, Ty Dickinson, Chase Dwelle

Probabilistic Transformers for Joint Modeling of Global Weather Dynamics and Decision-Centric Variables

Weather forecasts sit upstream of high-stakes decisions in domains such as grid operations, aviation, agriculture, and emergency response. Yet forecast users often face a difficult trade-off. Many decision-relevant targets are functionals of the atmospheric state variables, such as extrema, accumulations, and threshold exceedances,...

💬 0 commentsarXiv:2601.03753v1PDF
0

Posted in cs.CL · 2026-01-07 · Dominik Macko

Evaluation of Multilingual LLMs Personalized Text Generation Capabilities Targeting Groups and Social-Media Platforms

Capabilities of large language models to generate multilingual coherent text have continuously enhanced in recent years, which opens concerns about their potential misuse. Previous research has shown that they can be misused for generation of personalized disinformation in multiple languages. It has also been observed that...

💬 0 commentsarXiv:2601.03752v1PDF
0

Posted in cs.IR · 2026-01-07 · Dario Maio, Stefano Rizzi

Bridging OLAP and RAG: A Multidimensional Approach to the Design of Corpus Partitioning

Retrieval-Augmented Generation (RAG) systems are increasingly deployed on large-scale document collections, often comprising millions of documents and tens of millions of text chunks. In industrial-scale retrieval platforms, scalability is typically addressed through horizontal sharding and a combination of Approximate...

💬 0 commentsarXiv:2601.03748v1PDF