Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 14, 2026 — 18:00:23 EST

0

Posted in cs.CL · 2026-01-07 · Jonggeun Lee, Junseong Pyo, Gyuhyeon Seo, Yohan Jo

SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?

Large Audio-Language Models (LALMs) as judges have emerged as a prominent approach for evaluating speech generation quality, yet their ability to assess speaker consistency across multi-turn dialogues remains unexplored. We present \textbf{SpeakerSleuth}, a benchmark evaluating whether LALMs can reliably judge speaker consistency...

💬 0 commentsarXiv:2601.04029v2PDF
0

Posted in cs.CL · 2026-01-07 · Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew Lan

Simulated Students in Tutoring Dialogues: Substance or Illusion?

Advances in large language models (LLMs) enable many new innovations in education. However, evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up. Therefore, many recent works on LLM-powered tutoring solutions have used simulated students for both training and evaluation,...

💬 0 commentsarXiv:2601.04025v2PDF
0

Posted in cs.LG · 2026-01-07 · Kevin Innerebner, Stephan Bartl, Markus Reiter-Haas, Elisabeth Lex

Modeling Behavioral Patterns in News Recommendations Using Fuzzy Neural Networks

News recommender systems are increasingly driven by black-box models, offering little transparency for editorial decision-making. In this work, we introduce a transparent recommender system that uses fuzzy neural networks to learn human-readable rules from behavioral data for predicting article clicks. By extracting the rules at...

💬 0 commentsarXiv:2601.04019v1PDF
0

Posted in cs.DL · 2026-01-07 · George Macgregor, Joy Davidson

Examining persistence of European open repository infrastructure and its diffusion in the scholarly record

This article seeks to determine the extent to which the principle of persistence is observed by repositories and the organizations that operate them. We also evaluate the impact that negative repository persistence levels may be having on the scholarly record. We do this by interrogating and combining data about European repositories...

💬 0 commentsarXiv:2601.04015v3PDF
0

Posted in cs.IR · 2026-01-07 · Minglei Yin, Chuanbo Hu, Bin Liu, Neil Zhenqiang Gong, Yanfang, Ye, Xin Li

Correct and Weight: A Simple Yet Effective Loss for Implicit Feedback Recommendation

Learning from implicit feedback has become the standard paradigm for modern recommender systems. However, this setting is fraught with the persistent challenge of false negatives, where unobserved user-item interactions are not necessarily indicative of negative preference. To address this issue, this paper introduces a novel and...

💬 0 commentsarXiv:2601.04291v1PDF
0

Posted in cs.IT · 2026-01-07 · Wei Shi, Wei Xu, Yongming Huang, Jiacheng Yao, Wenhao Hu, Dongming Wang

Flexible-Duplex Cell-Free Architecture for Secure Uplink Communications in Low-Altitude Wireless Networks

Low-altitude wireless networks (LAWNs) are expected to play a central role in future 6G infrastructures, yet uplink transmissions of uncrewed aerial vehicles (UAVs) remain vulnerable to eavesdropping due to their limited transmit power, constrained antenna resources, and highly exposed air-ground propagation conditions. To address...

💬 0 commentsarXiv:2601.04011v1PDF
0

Posted in cs.SE · 2026-01-07 · Yannick Landeck, Dian Balta, Martin Wimmer, Christian Knierim

An Ontology-Based Approach to Security Risk Identification of Container Deployments in OT Contexts

In operational technology (OT) contexts, containerised applications often require elevated privileges to access low-level network interfaces or perform administrative tasks such as application monitoring. These privileges reduce the default isolation provided by containers and introduce significant security risks. Security risk...

💬 0 commentsarXiv:2601.04010v1PDF
0

Posted in cs.CY · 2026-01-07 · Jonas Klingwort, Nina M. Leach, Joep Burger

Performance of models for monitoring sustainable development goals from remote sensing: A three-level meta-regression

Machine learning (ML) is a tool to exploit remote sensing data for the monitoring and implementation of the United Nations' Sustainable Development Goals (SDGs). In this paper, we report on a meta-analysis to evaluate the performance of ML applied to remote sensing data to monitor SDGs. Specifically, we aim to 1) estimate the average...

💬 0 commentsarXiv:2601.06178v1PDF
0

Posted in cs.CV · 2026-01-07 · Onur Keleş, A. Murat Tekalp

Padé Neurons for Efficient Neural Models

Neural networks commonly employ the McCulloch-Pitts neuron model, which is a linear model followed by a point-wise non-linear activation. Various researchers have already advanced inherently non-linear neuron models, such as quadratic neurons, generalized operational neurons, generative neurons, and super neurons, which offer stronger...

💬 0 commentsarXiv:2601.04005v1PDF
0

Posted in cs.CL · 2026-01-07 · José Pedro Evans, Luís Filipe Cunha, Purificação Silvano, Alípio Jorge, Nuno Guimarães, Sérgio Nunes, Ricardo Campos

VotIE: Information Extraction from Meeting Minutes

Municipal meeting minutes record key decisions in local democratic processes. Unlike parliamentary proceedings, which typically adhere to standardized formats, they encode voting outcomes in highly heterogeneous, free-form narrative text that varies widely across municipalities, posing significant challenges for automated extraction....

💬 0 commentsarXiv:2601.03997v3PDF
0

Posted in cs.CV · 2026-01-07 · Junle Liu, Peirong Zhang, Yuyi Zhang, Pengyu Yan, Hui Zhou, Xinyue Zhou, Fengjun Guo, Lianwen Jin

PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography

Commercial-grade poster design demands the seamless integration of aesthetic appeal with precise, informative content delivery. Current automated poster generation systems face significant limitations, including incomplete design workflows, poor text rendering accuracy, and insufficient flexibility for commercial applications. To...

💬 0 commentsarXiv:2601.03993v1PDF
0

Posted in cs.DC · 2026-01-07 · Qi Wu, Chao Fang, Jiayuan Chen, Ye Lin, Yueqi Zhang, Yichuan Bai, Yuan Du, Li Du

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems

Mixture-of-Experts (MoE) models facilitate edge deployment by decoupling model capacity from active computation, yet their large memory footprint drives the need for GPU systems with near-data processing (NDP) capabilities that offload experts to dedicated processing units. However, deploying MoE models on such edge-based GPU-NDP...

💬 0 commentsarXiv:2601.03992v1PDF
0

Posted in cs.SE · 2026-01-07 · Nicolas Lacroix, Mireille Blay-Fornarino, Sébastien Mosser, Frederic Precioso

Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures

Background: Extracting the stages that structure Machine Learning (ML) pipelines from source code is key for gaining a deeper understanding of data science practices. However, the diversity caused by the constant evolution of the ML ecosystem (e.g., algorithms, libraries, datasets) makes this task challenging. Existing approaches...

💬 0 commentsarXiv:2601.03988v1PDF
0

Posted in cs.CL · 2026-01-07 · Qi Qian, Chengsong Huang, Jingwen Xu, Changze Lv, Muling Wu, Wenhao Liu, Xiaohua Wang, Zhenghua Wang, Zisu Huang, Muzhao Tian, Jianhan Xu, Kun Hu, He-Da Wang, Yao Hu, Xuanjing Huang, Xiaoqing Zheng

Benchmark^2: Systematic Evaluation of LLM Benchmarks

The rapid proliferation of benchmarks for evaluating large language models (LLMs) has created an urgent need for systematic methods to assess benchmark quality itself. We propose Benchmark^2, a comprehensive framework comprising three complementary metrics: (1) Cross-Benchmark Ranking Consistency, measuring whether a benchmark...

💬 0 commentsarXiv:2601.03986v1PDF
0

Posted in cs.LG · 2026-01-07 · Anmol Guragain

Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset

We present a systematic study of multimodal emotion recognition using the EAV dataset, investigating whether complex attention mechanisms improve performance on small datasets. We implement three model categories: baseline transformers (M1), novel factorized attention mechanisms (M2), and improved CNN baselines (M3). Our experiments...

💬 0 commentsarXiv:2601.22161v2PDF
0

Posted in cs.CL · 2026-01-07 · Yuechen Jiang, Zhiwei Liu, Yupeng Cao, Yueru He, Ziyang Xu, Chen Xu, Zhiyang Deng, Prayag Tiwari, Xi Chen, Alejandro Lopez-Lira, Jimin Huang, Junichi Tsujii, Sophia Ananiadou

All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection

We introduce RFC Bench, a benchmark for evaluating large language models on financial misinformation under realistic news. RFC Bench operates at the paragraph level and captures the contextual complexity of financial news where meaning emerges from dispersed cues. The benchmark defines two complementary tasks: reference free...

💬 0 commentsarXiv:2601.04160v3PDF
0

Posted in cs.CV · 2026-01-07 · Vladimir Frants, Sos Agaian, Karen Panetta

ToTMNet: FFT-Accelerated Toeplitz Temporal Mixing Network for Lightweight Remote Photoplethysmography

Remote photoplethysmography (rPPG) estimates a blood volume pulse (BVP) waveform from facial videos captured by commodity cameras. Although recent deep models improve robustness compared to classical signal-processing approaches, many methods increase computational cost and parameter count, and attention-based temporal modeling...

💬 0 commentsarXiv:2601.04159v1PDF
0

Posted in cs.CL · 2026-01-07 · Adar Avsian, Christopher Richardson, Anirudh Sundar, Larry Heck

FLEx: Language Modeling with Few-shot Language Explanations

Language models have become effective at a wide range of tasks, from math problem solving to open-domain question answering. However, they still make mistakes, and these mistakes are often repeated across related queries. Natural language explanations can help correct these errors, but collecting them at scale may be infeasible,...

💬 0 commentsarXiv:2601.04157v2PDF
0

Posted in cs.CV · 2026-01-07 · Chenye Meng, Zejian Li, Zhongni Liu, Yize Li, Changle Xie, Kaixin Jia, Ling Yang, Huanghuang Deng, Shiying Ding, Shengyuan Zhang, Jiayi Li, Lingyun Sun

Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes

Post-training alignment of diffusion models relies on simplified signals, such as scalar rewards or binary preferences. This limits alignment with complex human expertise, which is hierarchical and fine-grained. To address this, we first construct a hierarchical, fine-grained evaluation criteria with domain experts, which decomposes...

💬 0 commentsarXiv:2601.04300v1PDF
0

Posted in cs.CV · 2026-01-07 · Yifan Wang, Yanyu Li, Gordon Guocheng Qian, Sergey Tulyakov, Yun Fu, Anil Kag

Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning

Video diffusion alignment has been heavily relied on scalar rewards. These rewards are typically derived from learned reward models in human preference datasets, requiring additional training and extensive collection. Moreover, scalar rewards provide coarse, global supervision, offering limited prompt-generation mismatch credit...

💬 0 commentsarXiv:2601.04153v2PDF
0

Posted in cs.CV · 2026-01-07 · Jun Wang, Chunyu Qiang, Yuxin Guo, Yiran Wang, Xijuan Zeng, Feng Deng

Apollo: Unified Multi-Task Audio-Video Joint Generation

Audio-video joint generation has progressed rapidly, yet substantial challenges still remain. Non-commercial approaches still suffer audio-visual asynchrony, poor lip-speech alignment, and unimodal degradation, which can be stemmed from weak audio-visual correspondence modeling, limited generalization, and scarce high-quality...

💬 0 commentsarXiv:2601.04151v2PDF
0

Posted in cs.LG · 2026-01-07 · Pir Bakhsh Khokhar, Carmine Gravino, Fabio Palomba, Sule Yildrim Yayilgan, Sarang Shaikh

Transformer-Based Multi-Modal Temporal Embeddings for Explainable Metabolic Phenotyping in Type 1 Diabetes

Type 1 diabetes (T1D) is a highly metabolically heterogeneous disease that cannot be adequately characterized by conventional biomarkers such as glycated hemoglobin (HbA1c). This study proposes an explainable deep learning framework that integrates continuous glucose monitoring (CGM) data with laboratory profiles to learn multimodal...

💬 0 commentsarXiv:2601.04299v1PDF
0

Posted in cs.CR · 2026-01-07 · M. Amin Rahimian, Benjamin Panny, James Joshi

Privacy at Scale in Networked Healthcare

Digitized, networked healthcare promises earlier detection, precision therapeutics, and continuous care; yet, it also expands the surface for privacy loss and compliance risk. We argue for a shift from siloed, application-specific protections to privacy-by-design at scale, centered on decision-theoretic differential privacy (DP)...

💬 0 commentsarXiv:2601.04298v1PDF
0

Posted in cs.RO · 2026-01-07 · Chun-Kai Fan, Xiaowei Chi, Xiaozhu Ju, Hao Li, Yong Bao, Yu-Kai Wang, Lizhang Chen, Zhiyuan Jiang, Kuangzhi Ge, Ying Li, Weishi Mi, Qingpo Wuwu, Peidong Jia, Yulin Luo, Kevin Zhang, Zhiyuan Qin, Yong Dai, Sirui Han, Yike Guo, Shanghang Zhang, Jian Tang

Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test

As world models gain momentum in Embodied AI, an increasing number of works explore using video foundation models as predictive world models for downstream embodied tasks like 3D prediction or interactive generation. However, before exploring these downstream tasks, video foundation models still have two critical questions unanswered:...

💬 0 commentsarXiv:2601.04137v1PDF
0

Posted in cs.CL · 2026-01-07 · Leonardo Bottona, Nicolò Penzo, Bruno Lepri, Marco Guerini, Sara Tonelli

LLMberjack: Guided Trimming of Debate Trees for Multi-Party Conversation Creation

We present LLMberjack, a platform for creating multi-party conversations starting from existing debates, originally structured as reply trees. The system offers an interactive interface that visualizes discussion trees and enables users to construct coherent linearized dialogue sequences while preserving participant identity and...

💬 0 commentsarXiv:2601.04135v1PDF