Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 12, 2026 — 04:09:00 EST

0

Posted in cs.CL · 2026-01-13 · Youwei Liu, Jian Wang, Hanlin Wang, Beichen Guo, Wenjie Li

Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models

Recent advances in world models have shown promise for modeling future dynamics of environmental states, enabling agents to reason and act without accessing real environments. Current methods mainly perform single-step or fixed-horizon rollouts, leaving their potential for complex task planning under-exploited. We propose...

💬 0 commentsarXiv:2601.08955v2PDF
0

Posted in cs.HC · 2026-01-13 · Sumin Hong, Jewoong Moon, Taeyeon Eom, Juno Hwang, Jibeom Seo

Leveraging learning analytics to enhance immersive teacher simulations: Challenges and opportunities

This chapter examines how data analytics can be leveraged to enhance immersive teacher simulations, situating this inquiry within the broader learning sciences discourse on embodied cognition, data-informed feedback, and teacher professional learning. It explores both conceptual foundations and empirical cases to illustrate how...

💬 0 commentsarXiv:2601.08954v1PDF
0

Posted in cs.RO · 2026-01-13 · Le Liu, Bangguo Yu, Nynke Vellinga, Ming Cao

Fairness risk and its privacy-enabled solution in AI-driven robotic applications

Complex decision-making by autonomous machines and algorithms could underpin the foundations of future society. Generative AI is emerging as a powerful engine for such transitions. However, we show that Generative AI-driven developments pose a critical pitfall: fairness concerns. In robotic applications, although intuitions about...

💬 0 commentsarXiv:2601.08953v1PDF
0

Posted in cs.CY · 2026-01-13 · Jing-Jing Li, Joel Mire, Eve Fleisig, Valentina Pyatkin, Anne Collins, Maarten Sap, Sydney Levine

PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduce PluriHarms, a benchmark designed...

💬 0 commentsarXiv:2601.08951v2PDF
0

Posted in cs.AI · 2026-01-13 · Mayank Sharma, Roy Pea, Hari Subramonyam

ConvoLearn: A Learning Sciences Grounded Dataset for Fine-Tuning Dialogic AI Tutors

Despite their growing adoption in education, LLMs remain misaligned with the core principle of effective tutoring: the dialogic construction of knowledge. We introduce ConvoLearn, a dataset of 2,134 semi-synthetic tutor-student dialogues operationalizing six dimensions of dialogic tutoring grounded in knowledge-building theory,...

💬 0 commentsarXiv:2601.08950v4PDF
0

Posted in cs.CR · 2026-01-13 · David Brundage

Synthetic Data for Veterinary EHR De-identification: Benefits, Limits, and Safety Trade-offs Under Fixed Compute

Veterinary electronic health records (vEHRs) contain privacy-sensitive identifiers that limit secondary use. While PetEVAL provides a benchmark for veterinary de-identification, the domain remains low-resource. This study evaluates whether large language model (LLM)-generated synthetic narratives improve de-identification safety under...

💬 0 commentsarXiv:2601.09756v1PDF
0

Posted in cs.DB · 2026-01-12 · Zehai Yang, Shimin Chen

RAIRS: Optimizing Redundant Assignment and List Layout for IVF-Based ANN Search

IVF is one of the most widely used ANNS (Approximate Nearest Neighbors Search) methods in vector databases. The idea of redundant assignment is to assign a data vector to more than one IVF lists for reducing the chance of missing true neighbors in IVF search. However, the naive strategy, which selects the second IVF list based on the...

💬 0 commentsarXiv:2601.07183v1PDF
0

Posted in cs.LG · 2026-01-12 · Ruiyi Ding, Yongxuan Lv, Xianhui Meng, Jiahe Song, Chao Wang, Chen Jiang, Yuan Cheng

PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization

Policy optimization for large language models often suffers from sparse reward signals in multi-step reasoning tasks. Critic-free methods like GRPO assign a single normalized outcome reward to all tokens, providing limited guidance for intermediate reasoning . While Process Reward Models (PRMs) offer dense feedback, they risk...

💬 0 commentsarXiv:2601.07182v3PDF
0

Posted in cs.CV · 2026-01-12 · Yichun Zhang, Xiangwu Guo, Yauhong Goh, Jessica Hu, Zhiheng Chen, Xin Wang, Difei Gao, Mike Zheng Shou

ShowUI-Aloha: Human-Taught GUI Agent

Graphical User Interfaces (GUIs) are central to human-computer interaction, yet automating complex GUI tasks remains a major challenge for autonomous agents, largely due to a lack of scalable, high-quality training data. While recordings of human demonstrations offer a rich data source, they are typically long, unstructured, and lack...

💬 0 commentsarXiv:2601.07181v1PDF
0

Posted in cs.CL · 2026-01-12 · Jinyi Han, Zixiang Di, Zishang Jiang, Ying Liao, Jiaqing Liang, Yongqi Wang, Yanghua Xiao

Structured Reasoning for Large Language Models

Large language models (LLMs) achieve strong performance by generating long chains of thought, but longer traces always introduce redundant or ineffective reasoning steps. One typical behavior is that they often perform unnecessary verification and revisions even if they have reached the correct answers. This limitation stems from the...

💬 0 commentsarXiv:2601.07180v1PDF
0

Posted in cs.CV · 2026-01-12 · Weilin Zhou, Zonghao Ying, Chunlei Meng, Jiahui Liu, Hengyang Zhou, Quanchen Zou, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang

DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection

Multimodal fake news detection is crucial for mitigating adversarial misinformation. Existing methods, relying on static fusion or LLMs, face computational redundancy and hallucination risks due to weak visual foundations. To address this, we propose DIVER (Dynamic Iterative Visual Evidence Reasoning), a framework grounded in a...

💬 0 commentsarXiv:2601.07178v1PDF
0

Posted in cs.CR · 2026-01-12 · Mingxiang Tao, Yu Tian, Wenxuan Tu, Yue Yang, Xue Yang, Xiangyan Tang

Safe-FedLLM: Delving into the Safety of Federated Large Language Models

Federated learning (FL) addresses privacy and data-silo issues in the training of large language models (LLMs). Most prior work focuses on improving the efficiency of federated learning for LLMs (FedLLM). However, security in open federated environments, particularly defenses against malicious clients, remains underexplored. To...

💬 0 commentsarXiv:2601.07177v5PDF
0

Posted in cs.ET · 2026-01-12 · Mehran Moghadam, Sercan Aygun, M. Hassan Najafi

TranSC: Hardware-Aware Design of Transcendental Functions Using Stochastic Logic

The hardware-friendly implementation of transcendental functions remains a longstanding challenge in design automation. These functions, which cannot be expressed as finite combinations of algebraic operations, pose significant complexity in digital circuit design. This study introduces a novel approach, TranSC, that utilizes...

💬 0 commentsarXiv:2601.07172v1PDF
0

Posted in cs.CL · 2026-01-12 · Xun Xu

G-MemLLM: Gated Latent Memory Augmentation for Long-Context Reasoning in Large Language Models

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding, yet they remain constrained by the finite capacity of their context windows and the inherent difficulty of maintaining long-term factual consistency during multi-hop reasoning. While existing methods utilize context compression or...

💬 0 commentsarXiv:2602.00015v1PDF
0

Posted in cs.IR · 2026-01-12 · Zihang Li, Wenjun Liu, Yikun Zong, Jiawen Tao, Siying Dai, Songcheng Ren, Zirui Liu, Yuhang Wang, Yanbing Jiang, Tong Yang

Bridge-RAG: An Abstract Bridge Tree Based Retrieval Augmented Generation Algorithm

As an important paradigm for enhancing the generation quality of Large Language Models (LLMs), retrieval-augmented generation (RAG) faces the two challenges regarding retrieval accuracy and computational efficiency. This paper presents a novel RAG framework called Bridge-RAG. To overcome the accuracy challenge, we introduce the...

💬 0 commentsarXiv:2603.26668v2PDF
0

Posted in cs.LG · 2026-01-12 · Min Wang, Xin Li, Mingzhong Wang, Hasnaa Bennis

Offline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature Overgeneralization

Offline meta-reinforcement learning (OMRL) combines the strengths of learning from diverse datasets in offline RL with the adaptability to new tasks of meta-RL, promising safe and efficient knowledge acquisition by RL agents. However, OMRL still suffers extrapolation errors due to out-of-distribution (OOD) actions, compromised by...

💬 0 commentsarXiv:2601.07164v1PDF
0

Posted in cs.CV · 2026-01-12 · Shu Shen, C. L. Philip Chen, Tong Zhang

Test-time Adaptive Hierarchical Co-enhanced Denoising Network for Reliable Multimodal Classification

Reliable learning of multimodal data (e.g., multi-omics) is a widely concerning issue, especially in safety-critical applications such as medical diagnosis. However, low-quality data induced by multimodal noise poses a major challenge in this domain, causing existing methods to suffer from two key limitations. First, they struggle to...

💬 0 commentsarXiv:2601.07163v2PDF
0

Posted in cs.AI · 2026-01-12 · Xinzi Cao, Jianyang Zhai, Pengfei Li, Zhiheng Hu, Cen Yan, Bingxu Mu, Guanghuan Fang, Bin She, Jiayu Li, Yihan Su, Dongyang Tao, Xiansong Huang, Fan Xu, Feidiao Yang, Yao Lu, Chang-Dong Wang, Yutong Lu, Weicheng Xue, Bin Zhou, Yonghong Tian

AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units

To meet the ever-increasing demand for computational efficiency, Neural Processing Units (NPUs) have become critical in modern AI infrastructure. However, unlocking their full potential requires developing high-performance compute kernels using vendor-specific Domain-Specific Languages (DSLs), a task that demands deep hardware...

💬 0 commentsarXiv:2601.07160v2PDF
0

Posted in cs.LG · 2026-01-12 · Ijun Jang, Jewon Yeom, Juan Yeo, Hyunggyu Lim, Taesup Kim

Stable On-Policy Distillation through Adaptive Target Reformulation

Knowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models; however, conventional supervised KD often suffers from a distribution mismatch between training and inference. While on-policy KD approaches attempt to mitigate this issue by learning directly from...

💬 0 commentsarXiv:2601.07155v3PDF
0

Posted in cs.CV · 2026-01-12 · Si-En Hong, James Tribble, Alexander Lake, Hao Wang, Chaoyi Zhou, Ashish Bastola, Siyu Huang, Eisa Chaudhary, Brian Canada, Ismahan Arslan-Ari, Abolfazl Razi

Motion Focus Recognition in Fast-Moving Egocentric Video

From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To bridge this gap, we propose a real-time motion focus recognition method that estimates the...

💬 0 commentsarXiv:2601.07154v3PDF
0

Posted in cs.CL · 2026-01-12 · Genta Indra Winata, David Anugraha, Patrick Amadeus Irawan, Anirban Das, Haneul Yoo, Paresh Dashore, Shreyas Kulkarni, Ruochen Zhang, Haruki Sakajo, Frederikus Hudi, Anaelia Ovalle, Syrielle Montariol, Felix Gaschi, Michael Anugraha, Rutuj Ravindra Puranik, Zawad Hayat Ahmed, Adril Putra Merin, Emmanuele Chersoni

Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?

Code-switching is a pervasive phenomenon in multilingual communication, yet the robustness of large language models (LLMs) in mixed-language settings remains insufficiently understood. In this work, we present a comprehensive evaluation of LLM capabilities in understanding, reasoning over, and generating code-switched text. We...

💬 0 commentsarXiv:2601.07153v1PDF
0

Posted in cs.MA · 2026-01-12 · Aja Khanal, Kaushik T. Ranade, Rishabh Agrawal, Kalyan S. Basu, Apurva Narayan

Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)

Generating high-quality structured data such as JSON records, remains a fundamental challenge for large language models (LLMs), particularly when semantic richness must coexist with strict schema adherence. While autoregressive LLMs offer strong structural consistency, they often struggle with semantic variation and output diversity....

💬 0 commentsarXiv:2601.07152v1PDF
0

Posted in cs.AI · 2026-01-12 · Zhaoyan Li, Hang Lei, Yujia Wang, Lanbo Liu, Hao Liu, Liang Yu

Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling

While Large Language Models (LLMs) can generate fluent text, producing high-quality creative stories remains challenging. Reinforcement Learning (RL) offers a promising solution but faces two critical obstacles: designing reliable reward signals for subjective storytelling quality and mitigating training instability. This paper...

💬 0 commentsarXiv:2601.07149v1PDF
0

Posted in cs.CL · 2026-01-12 · Zhengxiang Wang, Zeyu Dong

Measuring Iterative Temporal Reasoning with Time Puzzles

Tool use, such as web search, has become a standard capability even in freely available large language models (LLMs). However, existing benchmarks evaluate temporal reasoning mainly in static, non-tool-using settings, which poorly reflect how LLMs perform temporal reasoning in practice. We introduce Time Puzzles, a constraint-based...

💬 0 commentsarXiv:2601.07148v3PDF