Qwen Councils
arXiv could not process that search. Try a simpler keyword search or an arXiv field query such as all:quantum.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through September 14, 2026 — 23:16:38 EST

0

Posted in cs.CV · 2026-01-08 · Akbar Saadat

Defocus Aberration Theory Confirms Gaussian Model in Most Imaging Devices

Over the past three decades, defocus has consistently provided groundbreaking depth information in scene images. However, accurately estimating depth from 2D images continues to be a persistent and fundamental challenge in the field of 3D recovery. Heuristic approaches involve with the ill-posed problem for inferring the spatial...

💬 0 commentsarXiv:2601.04779v1PDF
0

Posted in cs.CV · 2026-01-08 · Tobia Poppi, Burak Uzkent, Amanmeet Garg, Lucas Porto, Garin Kessler, Yezhou Yang, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara, Florian Schiffers

CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models

Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal order. Existing mitigation strategies, such as textual filtering or random video perturbations, often fail to address the root cause: over-reliance on language priors rather than...

💬 0 commentsarXiv:2601.04778v1PDF
0

Posted in cs.CV · 2026-01-08 · Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang, Yufei Zhan, Ming Tang, Jinqiao Wang

GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have demonstrated impressive progress in single-image grounding and general multi-image understanding. Recently, some methods begin to address multi-image grounding. However, they are constrained by single-target localization and limited types of practical tasks, due to the lack of unified...

💬 0 commentsarXiv:2601.04777v1PDF
0

Posted in cs.CV · 2026-01-08 · Jinyu Zhang, Xu Ma, Weili Chen

Segmentation-Driven Monocular Shape from Polarization based on Physical Model

Monocular shape-from-polarization (SfP) leverages the intrinsic relationship between light polarization properties and surface geometry to recover surface normals from single-view polarized images, providing a compact and robust approach for three-dimensional (3D) reconstruction. Despite its potential, existing monocular SfP methods...

💬 0 commentsarXiv:2601.04776v2PDF
0

Posted in cs.AI · 2026-01-08 · Encheng Su, Jianyu Wu, Chen Tang, Lintao Wang, Pengze Li, Aoran Wang, Jinouwen Zhang, Yizhou Wang, Yuan Meng, Xinzhu Ma, Shixiang Tang, Houqiang Li

SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence

As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of scientific inquiry. Existing benchmarks exhibit a critical blind spot: general instruction-following metrics focus on superficial formatting, while...

💬 0 commentsarXiv:2601.04770v2PDF
0

Posted in cs.SE · 2026-01-08 · Yelena Mujibur Sheikh, Awez Akhtar Khatik, Luoxi Tang, Yuqiao Meng, Zhaohan Xi

RiskBridge: Turning CVEs into Business-Aligned Patch Priorities

Enterprises are confronted with an unprecedented escalation in cybersecurity vulnerabilities, with thousands of new CVEs disclosed each month. Conventional prioritization frameworks such as CVSS offer static severity metrics that fail to account for exploit probability, compliance urgency, and operational impact, resulting in...

💬 0 commentsarXiv:2601.06201v2PDF
0

Posted in cs.CL · 2026-01-08 · Dongjun Kim, Jeongho Yoon, Chanjun Park, Heuiseok Lim

LANGSAE EDITING: Improving Multilingual Information Retrieval via Post-hoc Language Identity Removal

Dense retrieval in multilingual settings often searches over mixed-language collections, yet multilingual embeddings encode language identity alongside semantics. This language signal can inflate similarity for same-language pairs and crowd out relevant evidence written in other languages. We propose LANGSAE EDITING, a post-hoc sparse...

💬 0 commentsarXiv:2601.04768v1PDF
0

Posted in cs.AI · 2026-01-08 · Zefang Zong, Dingwei Chen, Yang Li, Qi Yi, Bo Zhou, Chengming Li, Bo Qian, Peng Chen, Jie Jiang

AT$^2$PO: Agentic Turn-based Policy Optimization via Tree Search

LLM agents have emerged as powerful systems for tackling multi-turn tasks by interleaving internal reasoning and external tool interactions. Agentic Reinforcement Learning has recently drawn significant research attention as a critical post-training paradigm to further refine these capabilities. In this paper, we present AT$^2$PO...

💬 0 commentsarXiv:2601.04767v1PDF
0

Posted in cs.CL · 2026-01-08 · Shengyin Sun, Yiming Li, Renxi Liu, Weizhe Lin, Hui-Ling Zhen, Xianzhi Yu, Mingxuan Yuan, Chen Ma

Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence

Judge Decoding accelerates LLM inference by relaxing the strict verification of Speculative Decoding, yet it typically relies on expensive and noisy supervision. In this work, we revisit this paradigm from first principles, revealing that the ``criticality'' scores learned via costly supervision are intrinsically encoded in the...

💬 0 commentsarXiv:2601.04766v1PDF
0

Posted in cs.CL · 2026-01-08 · Santiago Acevedo, Alessandro Laio, Marco Baroni

Differential syntactic and semantic encoding in LLMs

We study how syntactic and semantic information is encoded in inner layer representations of Large Language Models (LLMs), focusing on the very large DeepSeek-V3. We find that, by averaging hidden-representation vectors of sentences sharing syntactic structure or meaning, we obtain vectors that capture a significant proportion of the...

💬 0 commentsarXiv:2601.04765v5PDF
0

Posted in cs.AI · 2026-01-08 · Zhen Chen, Weihao Xie, Peilin Chen, Shiqi Wang, Jianping Wang

Orion-RAG: Path-Aligned Hybrid Retrieval for Graphless Data

Retrieval-Augmented Generation (RAG) has proven effective for knowledge synthesis, yet it encounters significant challenges in practical scenarios where data is inherently discrete and fragmented. In most environments, information is distributed across isolated files like reports and logs that lack explicit links. Standard search...

💬 0 commentsarXiv:2601.04764v1PDF
0

Posted in cs.LG · 2026-01-08 · Rupsa Rani Mishra, D. Chandrasekhar Rao, Ajaya Kumar Tripathy

Smart IoT-Based Wearable Device for Detection and Monitoring of Common Cow Diseases Using a Novel Machine Learning Technique

Manual observation and monitoring of individual cows for disease detection present significant challenges in large-scale farming operations, as the process is labor-intensive, time-consuming, and prone to reduced accuracy. The reliance on human observation often leads to delays in identifying symptoms, as the sheer number of animals...

💬 0 commentsarXiv:2601.04761v1PDF
0

Posted in cs.CL · 2026-01-08 · Yehoon Jang, Chaewon Lee, Hyun-seok Min, Sungchul Choi

PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks

The Patent Trial and Appeal Board (PTAB) of the USPTO adjudicates thousands of ex parte appeals each year, requiring the integration of technical understanding and legal reasoning. While large language models (LLMs) are increasingly applied in patent and legal practice, their use has remained limited to lightweight tasks, with no...

💬 0 commentsarXiv:2601.04758v1PDF
0

Posted in cs.DB · 2026-01-08 · Cristian Riveros, Benjamin Scheidt, Nicole Schweikardt

Structural Indexing of Relational Databases for the Evaluation of Free-Connex Acyclic Conjunctive Queries

We present an index structure to boost the evaluation of free-connex acyclic conjunctive queries (fc-ACQs) over relational databases. The main ingredient of the index associated with a given database $D$ is an auxiliary database $D_{col}$. Our main result states that for any fc-ACQ $Q$ over $D$, we can count the number of answers of...

💬 0 commentsarXiv:2601.04757v1PDF
0

Posted in cs.DS · 2026-01-08 · Tuukka Korhonen, Sang-il Oum

Branch-width of connectivity functions is fixed-parameter tractable

A connectivity function on a finite set $V$ is a symmetric submodular function $f \colon 2^V \to \mathbb{Z}$ with $f(\emptyset)=0$. We prove that finding a branch-decomposition of width at most $k$ for a connectivity function given by an oracle is fixed-parameter tractable (FPT), by providing an algorithm of running time $2^{O(k^2)}...

💬 0 commentsarXiv:2601.04756v2PDF
0

Posted in cs.CV · 2026-01-08 · Yen-Jen Chiou, Wei-Tse Cheng, Yuan-Fu Yang

ProFuse: Efficient Cross-View Context Fusion for Open-Vocabulary 3D Gaussian Splatting

We present ProFuse, an efficient context-aware framework for open-vocabulary 3D scene understanding with 3D Gaussian Splatting (3DGS). The pipeline enhances cross-view consistency and intra-mask cohesion within a direct registration setup, adding minimal overhead and requiring no render-supervised fine-tuning. Instead of relying on a...

💬 0 commentsarXiv:2601.04754v2PDF
0

Posted in cs.CV · 2026-01-08 · Masatomo Yoshida, Haruto Namura, Nicola Adami, Masahiro Okuda

Skeletonization-Based Adversarial Perturbations on Large Vision Language Model's Mathematical Text Recognition

This work explores the visual capabilities and limitations of foundation models by introducing a novel adversarial attack method utilizing skeletonization to reduce the search space effectively. Our approach specifically targets images containing text, particularly mathematical formula images, which are more challenging due to their...

💬 0 commentsarXiv:2601.04752v1PDF
0

Posted in cs.LG · 2026-01-08 · Luca Lanzilao, Angela Meyer

Intraday spatiotemporal PV power prediction at national scale using satellite-based solar forecast models

We present a novel framework for spatiotemporal photovoltaic (PV) power forecasting and use it to evaluate the reliability, sharpness, and overall performance of seven intraday PV power nowcasting models. The model suite includes satellite-based deep learning and optical-flow approaches and physics-based numerical weather prediction...

💬 0 commentsarXiv:2601.04751v1PDF
0

Posted in cs.DC · 2026-01-08 · Krishna Chaitanya Sunkara

Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers

This work presents DCIM 3.0, a unified framework integrating semantic reasoning, predictive analytics, autonomous orchestration, and unified connectivity for next-generation AI data center management. The framework addresses critical challenges in infrastructure automation, sustainability, and digital-twin design through knowledge...

💬 0 commentsarXiv:2601.04750v1PDF
0

Posted in cs.AI · 2026-01-08 · Xiaoxiao Li

When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail

Multi-agent AI systems have proven effective for complex reasoning. These systems are compounded by specialized agents, which collaborate through explicit communication, but incur substantial computational overhead. A natural question arises: can we achieve similar modularity benefits with a single agent that selects from a library of...

💬 0 commentsarXiv:2601.04748v2PDF
0

Posted in cs.AI · 2026-01-08 · Tingyu Wu, Zhisheng Chen, Ziyan Weng, Shuhe Wang, Chenglong Li, Shuo Zhang, Sen Hu, Silin Wu, Qizhen Lan, Huacan Wang, Ronghao Chen

KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions

Existing long-horizon memory benchmarks mostly use multi-turn dialogues or synthetic user histories, which makes retrieval performance an imperfect proxy for person understanding. We present \BenchName, a publicly releasable benchmark built from long-form autobiographical narratives, where actions, context, and inner thoughts provide...

💬 0 commentsarXiv:2601.04745v2PDF
0

Posted in cs.SD · 2026-01-08 · Xingyuan Li, Mengyue Wu

Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling

Detecting medical conditions from speech acoustics is fundamentally a weakly-supervised learning problem: a single, often noisy, session-level label must be linked to nuanced patterns within a long, complex audio recording. This task is further hampered by severe data scarcity and the subjective nature of clinical annotations. While...

💬 0 commentsarXiv:2601.04744v2PDF
0

Posted in cs.CL · 2026-01-08 · Seyeon Jeong, Yeonjun Choi, JongWook Kim, Beakcheol Jang

Tool-MAD: A Multi-Agent Debate Framework for Fact Verification with Diverse Tool Augmentation and Adaptive Retrieval

Large Language Models (LLMs) suffer from hallucinations and factual inaccuracies, especially in complex reasoning and fact verification tasks. Multi-Agent Debate (MAD) systems aim to improve answer accuracy by enabling multiple LLM agents to engage in dialogue, promoting diverse reasoning and mutual verification. However, existing MAD...

💬 0 commentsarXiv:2601.04742v1PDF
0

Posted in cs.LG · 2026-01-08 · Kota Nakamura, Koki Kawabata, Yasuko Matsubara, Yasushi Sakurai

Fast Mining and Dynamic Time-to-Event Prediction over Multi-sensor Data Streams

Given real-time sensor data streams obtained from machines, how can we continuously predict when a machine failure will occur? This work aims to continuously forecast the timing of future events by analyzing multi-sensor data streams. A key characteristic of real-world data streams is their dynamic nature, where the underlying...

💬 0 commentsarXiv:2601.04741v2PDF
0

Posted in cs.CL · 2026-01-08 · Huawei Zheng, Xinqi Jiang, Sen Yang, Shouling Ji, Yingcai Wu, Dazhen Deng

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risks. Domain-specific datasets of harmful prompts remain scarce and still largely rely on manual construction; public datasets mainly focus on explicit harmful prompts, which modern LLM...

💬 0 commentsarXiv:2601.04740v3PDF