Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 12, 2026 — 20:45:46 EST

0

Posted in cs.CV · 2026-01-13 · Dongting Hu, Aarush Gupta, Magzhan Gabidolla, Arpit Sahni, Huseyin Coskun, Yanyu Li, Yerlan Idelbayev, Ahsan Mahmood, Aleksei Lebedev, Dishani Lahiri, Anujraaj Goyal, Ju Hu, Mingming Gong, Sergey Tulyakov, Anil Kag

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices

Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to their high computational and memory costs. In this work, we present an efficient DiT framework tailored for mobile and edge devices that achieves transformer-level generation quality under...

💬 0 commentsarXiv:2601.08303v3PDF
0

Posted in cs.CL · 2026-01-13 · Marvin Schmitt, Anne Schwerk, Sebastian Lempert

Enhancing Sentiment Classification and Irony Detection in Large Language Models through Advanced Prompt Engineering Techniques

This study investigates the use of prompt engineering to enhance large language models (LLMs), specifically GPT-4o-mini and gemini-1.5-flash, in sentiment analysis tasks. It evaluates advanced prompting techniques like few-shot learning, chain-of-thought prompting, and self-consistency against a baseline. Key tasks include sentiment...

💬 0 commentsarXiv:2601.08302v1PDF
0

Posted in cs.CV · 2026-01-13 · Qizhen Lan, Yu-Chun Hsu, Nida Saddaf Khan, Xiaoqian Jiang

ReCo-KD: Region- and Context-Aware Knowledge Distillation for Efficient 3D Medical Image Segmentation

Accurate 3D medical image segmentation is vital for diagnosis and treatment planning, but state-of-the-art models are often too large for clinics with limited computing resources. Lightweight architectures typically suffer significant performance loss. To address these deployment and speed constraints, we propose Region- and...

💬 0 commentsarXiv:2601.08301v1PDF
0

Posted in cs.CL · 2026-01-13 · Tony Cristofano

Surgical Refusal Ablation: Disentangling Safety from Intelligence via Concept-Guided Spectral Cleaning

Safety-aligned language models systematically refuse harmful requests. While activation steering can modulate refusal, ablating the raw "refusal vector" calculated from contrastive harmful and harmless prompts often causes collateral damage and distribution drift. We argue this degradation occurs because the raw vector is...

💬 0 commentsarXiv:2601.08489v1PDF
0

Posted in cs.RO · 2026-01-13 · Chong Zhang, Victor Klemm, Fan Yang, Marco Hutter

AME-2: Agile and Generalized Legged Locomotion via Attention-Based Neural Map Encoding

Achieving agile and generalized legged locomotion across terrains requires tight integration of perception and control, especially under occlusions and sparse footholds. Existing methods have demonstrated agility on parkour courses but often rely on end-to-end sensorimotor models with limited generalization and interpretability. By...

💬 0 commentsarXiv:2601.08485v2PDF
0

Posted in cs.CV · 2026-01-13 · MD Fatin Ishraque Ayon, Sabrin Nahar, Ataur Rahman, Md. Taslim Arif, Abdul Hasib, A. S. M. Ahsanul Sarkar Akib

An IoT-Enabled Smart Aquarium System for Real-Time Water Quality Monitoring and Automated Feeding

Maintaining optimal water quality in aquariums is critical for aquatic health but remains challenging due to the need for continuous monitoring of multiple parameters. Traditional manual methods are inefficient, labor-intensive, and prone to human error, often leading to suboptimal aquatic conditions. This paper presents an IoT-based...

💬 0 commentsarXiv:2601.08484v1PDF
0

Posted in cs.LG · 2026-01-13 · Chenxu Han, Sean Bin Yang, Jilin Hu

DiffMM: Efficient Method for Accurate Noisy and Sparse Trajectory Map Matching via One Step Diffusion

Map matching for sparse trajectories is a fundamental problem for many trajectory-based applications, e.g., traffic scheduling and traffic flow analysis. Existing methods for map matching are generally based on Hidden Markov Model (HMM) or encoder-decoder framework. However, these methods continue to face significant challenges when...

💬 0 commentsarXiv:2601.08482v1PDF
0

Posted in cs.CR · 2026-01-13 · Aryan Pasikhani, Prosanta Gope, Yang Yang, Shagufta Mehnaz, Biplab Sikdar

Baiting AI: Deceptive Adversary Against AI-Protected Industrial Infrastructures

This paper explores a new cyber-attack vector targeting Industrial Control Systems (ICS), particularly focusing on water treatment facilities. Developing a new multi-agent Deep Reinforcement Learning (DRL) approach, adversaries craft stealthy, strategically timed, wear-out attacks designed to subtly degrade product quality and reduce...

💬 0 commentsarXiv:2601.08481v1PDF
0

Posted in cs.CL · 2026-01-13 · Francesco Dettori, Matteo Forasassi, Lorenzo Veronese, Livia Lestingi, Vincenzo Scotti, Matteo Giovanni Rossi

Do You Understand How I Feel?: Towards Verified Empathy in Therapy Chatbots

Conversational agents are increasingly used as support tools along mental therapeutic pathways with significant societal impacts. In particular, empathy is a key non-functional requirement in therapeutic contexts, yet current chatbot development practices provide no systematic means to specify or verify it. This paper envisions a...

💬 0 commentsarXiv:2601.08477v1PDF
0

Posted in cs.CV · 2026-01-13 · Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin

Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of...

💬 0 commentsarXiv:2601.08476v2PDF
0

Posted in cs.AI · 2026-01-13 · JungMin Yun, Juhwan Choi, Kyohoon Jin, Soojin Jang, Jinhee Jang, YoungBin Kim

SUMMPILOT: Bridging Efficiency and Customization for Interactive Summarization System

This paper incorporates the efficiency of automatic summarization and addresses the challenge of generating personalized summaries tailored to individual users' interests and requirements. To tackle this challenge, we introduce SummPilot, an interaction-based customizable summarization system. SummPilot leverages a large language...

💬 0 commentsarXiv:2601.08475v1PDF
0

Posted in cs.LO · 2026-01-13 · M. E. Coniglio, F. Esteva, J. Gispert, L. Godo

Degree-preserving Godel logics with an involution: intermediate logics and (ideal) paraconsistency

In this paper we study intermediate logics between the degree preserving companion of Godel fuzzy logic with an involution and classical propositional logic CPL, as well as the intermediate logics of their finite-valued counterparts. Although these degree-preserving Godel logics are explosive with respect to Godel negation, they are...

💬 0 commentsarXiv:2601.08474v1PDF
0

Posted in cs.CL · 2026-01-13 · Benedikt Droste, Jan Philipp Harries, Maximilian Idahl, Björn Plüster

sui-1: Grounded and Verifiable Long-Form Summarization

Large language models frequently generate plausible but unfaithful summaries that users cannot verify against source text, a critical limitation in compliance-sensitive domains such as government and legal analysis. We present sui-1, a 24B parameter model that produces abstractive summaries with inline citations, enabling users to...

💬 0 commentsarXiv:2601.08472v1PDF
0

Posted in cs.CV · 2026-01-13 · Takara Taniguchi, Kuniaki Saito, Atsushi Hashimoto

Towards Safer Mobile Agents: Scalable Generation and Evaluation of Diverse Scenarios for VLMs

Vision Language Models (VLMs) are increasingly deployed in autonomous vehicles and mobile systems, making it crucial to evaluate their ability to support safer decision-making in complex environments. However, existing benchmarks inadequately cover diverse hazardous situations, especially anomalous scenarios with spatio-temporal...

💬 0 commentsarXiv:2601.08470v1PDF
0

Posted in cs.NE · 2026-01-13 · Jakub Fil, Yulia Sandamirskaya, Hector Gonzalez, Loïc Azzalin, Stefan Glüge, Lukas Friedenstab, Friedrich Wolf, Tim Rosmeisl, Matthias Lohrmann, Mahmoud Akl, Khaleel Khan, Leonie Wolf, Kristin Richter, Holm Puder, Mazhar Ali Bari, Xuan Choo, Noha Alharthi, Michael Hopkins, Mansoor Hanif Christian Mayr, Jens Struckmeier, Steve Furber

Heterogeneous computing platform for real-time robotics

After Industry 4.0 has embraced tight integration between machinery (OT), software (IT), and the Internet, creating a web of sensors, data, and algorithms in service of efficient and reliable production, a new concept of Society 5.0 is emerging, in which infrastructure of a city will be instrumented to increase reliability,...

💬 0 commentsarXiv:2601.09755v1PDF
0

Posted in cs.CL · 2026-01-13 · Jiangshan Duo, Hanyu Li, Hailin Zhang, Yudong Wang, Sujian Li, Liang Zhao

JudgeRLVR: Judge First, Generate Second for Efficient Reasoning

Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for reasoning in Large Language Models. However, optimizing solely for final-answer correctness often drives models into aimless, verbose exploration, where they rely on exhaustive trial-and-error tactics rather than structured planning to reach...

💬 0 commentsarXiv:2601.08468v1PDF
0

Posted in cs.CV · 2026-01-13 · Takamichi Miyata, Sumiko Miyata, Andrew Morris

Zero-Shot Distracted Driver Detection via Vision Language Models with Double Decoupling

Distracted driving is a major cause of traffic collisions, calling for robust and scalable detection methods. Vision-language models (VLMs) enable strong zero-shot image classification, but existing VLM-based distracted driver detectors often underperform in real-world conditions. We identify subject-specific appearance variations...

💬 0 commentsarXiv:2601.08467v3PDF
0

Posted in cs.CV · 2026-01-13 · Evgenii Maslov, Valentin Khrulkov, Anastasia Volkova, Anton Gusarov, Andrey Kuznetsov, Ivan Oseledets

CoMa: Contextual Massing Generation with Vision-Language Models

The conceptual design phase in architecture and urban planning, particularly building massing, is complex and heavily reliant on designer intuition and manual effort. To address this, we propose an automated framework for generating building massing based on functional requirements and site context. A primary obstacle to such...

💬 0 commentsarXiv:2601.08464v1PDF
0

Posted in cs.AI · 2026-01-13 · Sixiong Xie, Zhuofan Shi, Haiyang Shen, Yun Ma, Xiang Jing

M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games

Existing benchmarks for LLM agents' social behavior typically focus on a single capability dimension and evaluate only behavioral outcomes, overlooking process signals from reasoning and communication. We present M3-BENCH, a benchmark of 24 mixed-motive games with a process-aware evaluation framework spanning three complementary...

💬 0 commentsarXiv:2601.08462v2PDF
0

Posted in cs.CV · 2026-01-13 · Chao Tian, Zikun Zhou, Chao Yang, Guoqing Zhu, Fu'an Zhong, Zhenyu He

Modality-Decoupled RGB-Thermal Object Detector via Query Fusion

The advantage of RGB-Thermal (RGB-T) detection lies in its ability to perform modality fusion and integrate cross-modality complementary information, enabling robust detection under diverse illumination and weather conditions. However, under extreme conditions where one modality exhibits poor quality and disturbs detection, modality...

💬 0 commentsarXiv:2601.08458v1PDF
0

Posted in cs.AI · 2026-01-13 · Sargam Yadav, Abhishek Kaushik, Kevin Mc Daid

An Under-Explored Application for Explainable Multimodal Misogyny Detection in code-mixed Hindi-English

Digital platforms have an ever-expanding user base, and act as a hub for communication, business, and connectivity. However, this has also allowed for the spread of hate speech and misogyny. Artificial intelligence models have emerged as an effective solution for countering online hate speech but are under explored for low resource...

💬 0 commentsarXiv:2601.08457v1PDF
0

Posted in cs.CV · 2026-01-13 · Sepideh Hatamikia, Geevarghese George, Florian Schwarzhans, Amirreza Mahbod, Marika AV Reinius, Ali Abbasian Ardakani, Mercedes Jimenez-Linan, Satish Viswanath, Mireia Crispin-Ortuzar, Lorena Escudero Sanchez, Evis Sala, James D Brenton, Ramona Woitek

Developing Predictive and Robust Radiomics Models for Chemotherapy Response in High-Grade Serous Ovarian Carcinoma

Objectives: High-grade serous ovarian carcinoma (HGSOC) is typically diagnosed at an advanced stage with extensive peritoneal metastases, making treatment challenging. Neoadjuvant chemotherapy (NACT) is often used to reduce tumor burden before surgery, but about 40% of patients show limited response. Radiomics, combined with machine...

💬 0 commentsarXiv:2601.08455v1PDF
0

Posted in cs.RO · 2026-01-13 · Alessandro Adami, Sebastian Zudaire, Ruggero Carli, Pietro Falco

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs

Constructing physically accurate simulation environments (Real2Sim) traditionally relies on manual system identification or rigid, exhaustive exploration routines. These task-agnostic pipelines often fail to leverage semantic scene context, leading to redundant physical interactions and inefficient data acquisition. In this paper, we...

💬 0 commentsarXiv:2601.08454v2PDF
0

Posted in cs.CR · 2026-01-13 · Shuiyin Liu, Amin Sakzad

On the Maximum Toroidal Distance Code for Lattice-Based Public-Key Cryptography

We propose a maximum toroidal distance (MTD) code for lattice-based public-key encryption (PKE). By formulating the encryption encoding problem as the selection of $2^\ell$ points in the discrete $\ell$-dimensional torus $\mathbb{Z}_q^\ell$, the proposed construction maximizes the minimum $L_2$-norm toroidal distance to reduce the...

💬 0 commentsarXiv:2601.08452v1PDF
0

Posted in cs.SD · 2026-01-13 · Minghui Zhao, Anton Ragni

Decoding Order Matters in Autoregressive Speech Synthesis

Autoregressive speech synthesis often adopts a left-to-right order, yet generation order is a modelling choice. We investigate decoding order through masked diffusion framework, which progressively unmasks positions and allows arbitrary decoding orders during training and inference. By interpolating between identity and random...

💬 0 commentsarXiv:2601.08450v1PDF