Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through September 11, 2026 — 03:55:32 EST

0

Posted in cs.CY · 2026-01-14 · Sonia Katyal

Lex Reformatica: Five Principles of Policy Reform for the Technological Age

Twenty-five years ago, Joel Reidenberg argued that technology itself, not just law and regulation, imposes rules on communities in the Information Society. System design choices like network architecture and configurations create regulatory norms he termed "Lex Informatica"-referencing the merchant-driven medieval "Lex Mercatoria"...

💬 0 commentsarXiv:2601.17001v1PDF
0

Posted in cs.CL · 2026-01-14 · Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu

SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages

Tonal low-resource languages are widely spoken yet remain underserved by modern speech technology. A key challenge is learning representations that are robust to nuisance variation such as gender while remaining tone-aware for different lexical meanings. To address this, we propose SITA, a lightweight adaptation recipe that enforces...

💬 0 commentsarXiv:2601.09050v1PDF
0

Posted in cs.CL · 2026-01-14 · Shikhar Shiromani, Archie Chaudhury, Sri Pranav Kunda

The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders

Large Language Models (LLMs) frequently exhibit unfaithful behavior, producing a final answer that differs significantly from their internal chain of thought (CoT) reasoning in order to appease the user they are conversing with. In order to better detect this behavior, we introduce the Hypocrisy Gap, a mechanistic metric utilizing...

💬 0 commentsarXiv:2602.02496v1PDF
0

Posted in cs.CY · 2026-01-14 · Sonia Katyal

Democracy and Distrust in an Era of Artificial Intelligence

This essay examines how judicial review should adapt to address challenges posed by artificial intelligence decision-making, particularly regarding minority rights and interests. As I argue in this essay, the rise of three trends-privatization, prediction, and automation in AI-have combined to pose similar risks to minorities. Here, I...

💬 0 commentsarXiv:2601.09757v1PDF
0

Posted in cs.CL · 2026-01-14 · Kaiyu He, Zhang Mian, Peilin Wu, Xinya Du, Zhiyu Chen

Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers

While Large Language Models (LLMs) excel at factual retrieval, they often struggle with the "curse of two-hop reasoning" in compositional tasks. Recent research suggests that parameter-sharing transformers can bridge this gap by forming a "Generalization Circuit" during a prolonged "grokking" phase. A fundamental question arises: Is a...

💬 0 commentsarXiv:2601.09049v1PDF
0

Posted in cs.HC · 2026-01-14 · Yuki Kobayashi, Koichi Toida

Immersive XR That Moves People: How XR Advertising Transforms Comprehension, Empathy, and Behavioural Intention

Extended Reality (XR) affords an enhanced sense of bodily presence that supports experiential modes of comprehension and affective engagement which exceed the possibilities of conventional information delivery. Nevertheless, the psychological processes engendered by XR, and the manner in which these processes inform subsequent...

💬 0 commentsarXiv:2601.09048v1PDF
0

Posted in cs.HC · 2026-01-14 · Hasan Tarik Akbaba, Efe Bozkir, Anna Puhl, Süleyman Özdel, Enkelejda Kasneci

Exploring Organizational Readiness and Ecosystem Coordination for Industrial XR

Extended Reality (XR) offers transformative potential for industrial support, training, and maintenance; yet, widespread adoption lags despite demonstrated occupational value and hardware maturity. Organizations successfully implement XR in isolated pilots, yet struggle to scale these into sustained operational deployment, a...

💬 0 commentsarXiv:2601.09045v2PDF
0

Posted in cs.LG · 2026-01-14 · Neelkamal Bhuyan, Debankur Mukherjee, Adam Wierman

SCaLE: Switching Cost aware Learning and Exploration

This work addresses the fundamental problem of unbounded metric movement costs in bandit online convex optimization, by considering high-dimensional dynamic quadratic hitting costs and $\ell_2$-norm switching costs in a noisy bandit feedback model. For a general class of stochastic environments, we provide the first algorithm SCaLE...

💬 0 commentsarXiv:2601.09042v1PDF
0

Posted in cs.CL · 2026-01-14 · Samhita Bollepally, Aurora Sloman-Moll, Takashi Yamauchi

Can LLMs interpret figurative language as humans do?: surface-level vs representational similarity

Large language models generate judgments that resemble those of humans. Yet the extent to which these models align with human judgments in interpreting figurative and socially grounded language remains uncertain. To investigate this, human participants and four instruction-tuned LLMs of different sizes (GPT-4, Gemma-2-9B, Llama-3.2,...

💬 0 commentsarXiv:2601.09041v1PDF
0

Posted in cs.CV · 2026-01-14 · Jonas Römer, Timo Dickscheid

Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers

End-to-end backpropagation couples all layers through a global error signal, enabling coordinated learning but requiring long-range credit assignment. Motivated by recent progress in blockwise self-supervised learning (BWSSL), we ask whether masked video transformers can be trained without end-to-end backpropagation. Applying BWSSL to...

💬 0 commentsarXiv:2601.09040v1PDF
0

Posted in cs.IT · 2026-01-14 · Mete Erdogan, Abhiram Gorle, Shubham Chandak, Mert Pilanci, Tsachy Weissman

An Information-Theoretic Perspective on LLM Tokenizers

Large language model (LLM) tokenizers act as structured compressors: by mapping text to discrete token sequences, they determine token count (and thus compute and context usage) and the statistical structure seen by downstream models. Despite their central role in LLM pipelines, the link between tokenization, compression efficiency...

💬 0 commentsarXiv:2601.09039v1PDF
0

Posted in cs.ET · 2026-01-14 · M Mahmudul Hasan Sajeeb, Kevin Callahan-Coray, Corentin Delacour, Sanjay Seshan, Tathagata Srimani, Kerem Y. Camsari

Probabilistic Computers for MIMO Detection: From Sparsification to 2D Parallel Tempering

Probabilistic computers built from p-bits offer a promising path for combinatorial optimization, but the dense connectivity required by real-world problems scales poorly in hardware. Here, we address this through graph sparsification with auxiliary copy variables and demonstrate two fully on-chip parallel tempering solvers on an FPGA....

💬 0 commentsarXiv:2601.09037v2PDF
0

Posted in cs.CL · 2026-01-14 · Sreya Vangara, Jagjit Nanda, Yan-Kai Tzeng, Eric Darve

SpectraQuery: A Hybrid Retrieval-Augmented Conversational Assistant for Battery Science

Scientific reasoning increasingly requires linking structured experimental data with the unstructured literature that explains it, yet most large language model (LLM) assistants cannot reason jointly across these modalities. We introduce SpectraQuery, a hybrid natural-language query framework that integrates a relational Raman...

💬 0 commentsarXiv:2601.09036v1PDF
0

Posted in cs.CR · 2026-01-14 · Aniesh Chawla, Udbhav Prasad

A Decompilation-Driven Framework for Malware Detection with Large Language Models

The parallel evolution of Large Language Models (LLMs) with advanced code-understanding capabilities and the increasing sophistication of malware presents a new frontier for cybersecurity research. This paper evaluates the efficacy of state-of-the-art LLMs in classifying executable code as either benign or malicious. We introduce an...

💬 0 commentsarXiv:2601.09035v1PDF
0

Posted in cs.AI · 2026-01-14 · Yiwen Tu, Xuan Liu, Lianhui Qin, Haojian Jin

PrivacyReasoner: Can LLM Emulate a Human-like Privacy Mind?

Prior work on LLM-based privacy focuses on norm judgment over synthetic vignettes, rather than how people think about a specific data practice and formulate their opinions. We address this gap by designing PrivacyReasoner, an agent architecture grounded in three key ideas: (1) LLMs can detect subtle privacy cues in natural language...

💬 0 commentsarXiv:2601.09152v2PDF
0

Posted in cs.LG · 2026-01-14 · Yang Nan, Qihao Wen, Jiahao Wang, Pengfei He, Ravi Tandon, Yong Ge, Han Xu

Interpretable Probability Estimation with LLMs via Shapley Reconstruction

Large Language Models (LLMs) demonstrate potential to estimate the probability of uncertain events, by leveraging their extensive knowledge and reasoning capabilities. This ability can be applied to support intelligent decision-making across diverse fields, such as financial forecasting and preventive healthcare. However, directly...

💬 0 commentsarXiv:2601.09151v1PDF
0

Posted in cs.HC · 2026-01-14 · Jianwen Sun, Yukang Feng, Kaining Ying, Chuanhao Li, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yifan Chang, Yu Dai, Yifei Huang, Kaipeng Zhang

World Craft: Agentic Framework to Create Visualizable Worlds via Text

Large Language Models (LLMs) motivate generative agent simulation (e.g., AI Town) to create a ``dynamic world'', holding immense value across entertainment and research. However, for non-experts, especially those without programming skills, it isn't easy to customize a visualizable environment by themselves. In this paper, we...

💬 0 commentsarXiv:2601.09150v4PDF
0

Posted in cs.CV · 2026-01-14 · Chenhao Fu, Han Fang, Xiuzheng Zheng, Wenbo Wei, Yonghua Li, Hao Sun, Xuelong Li

SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection

Zero-Shot Anomaly Detection (ZSAD) leverages Vision-Language Models (VLMs) to enable supervision-free industrial inspection. However, existing ZSAD paradigms are constrained by single visual backbones, which struggle to balance global semantic generalization with fine-grained structural discriminability. To bridge this gap, we propose...

💬 0 commentsarXiv:2601.09147v2PDF
0

Posted in cs.LG · 2026-01-14 · Xiucheng Xu, Bingbing Xu, Xueyun Tian, Zihe Huang, Rongxin Chen, Yunfan Li, Huawei Shen

Chain-of-Memory: Lightweight Memory Construction with Dynamic Evolution for LLM Agents

External memory systems are pivotal for enabling Large Language Model (LLM) agents to maintain persistent knowledge and perform long-horizon decision-making. Existing paradigms typically follow a two-stage process: computationally expensive memory construction (e.g., structuring data into graphs) followed by naive retrieval-augmented...

💬 0 commentsarXiv:2601.14287v2PDF
0

Posted in cs.DC · 2026-01-14 · Lingkang Shangguan

Transaction-Driven Dynamic Reconfiguration for Certificate-Based Payment Systems

We present a transaction-driven dynamic reconfiguration protocol in Modern payment systems based on Byzantine Consistent Broadcast which can achieve high performance by avoiding global transaction ordering. We demonstrate the fundamental paradigm of modern payment systems, which combines user nonce based transactions ordering with...

💬 0 commentsarXiv:2601.09146v1PDF
0

Posted in cs.LG · 2026-01-14 · Jinshuai Bai, Haolin Li, Zahra Sharif Khodaei, M. H. Aliabadi, YuanTong Gu, Xi-Qiao Feng

Discrete Solution Operator Learning for Geometry-Dependent PDEs

Neural operator learning accelerates PDE solution by approximating operators as mappings between continuous function spaces. Yet in many engineering settings, varying geometry induces discrete structural changes, including topological changes, abrupt changes in boundary conditions or boundary types, and changes in the computational...

💬 0 commentsarXiv:2601.09143v3PDF
0

Posted in cs.CY · 2026-01-14 · Caitlin A. Stamatis, Jonah Meyerhoff, Richard Zhang, Olivier Tieleman, Matteo Malgaroli, Thomas D. Hull

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Large language models (LLMs) are increasingly used for mental health support, yet existing safety evaluations rely primarily on small, simulation-based test sets that have an unknown relationship to the linguistic distribution of real usage. In this study, we present replications of four published safety test sets targeting suicide...

💬 0 commentsarXiv:2601.17003v1PDF
0

Posted in cs.LG · 2026-01-14 · Shijian Ma, Yan Lin, Yi Yang

EvasionBench: A Large-Scale Benchmark for Detecting Managerial Evasion in Earnings Call Q&A

We present EvasionBench, a comprehensive benchmark for detecting evasive responses in corporate earnings call question-and-answer sessions. Drawing from 22.7 million Q&A pairs extracted from S&P Capital IQ transcripts, we construct a rigorously filtered dataset and introduce a three-level evasion taxonomy: direct, intermediate, and...

💬 0 commentsarXiv:2601.09142v2PDF
0

Posted in cs.CL · 2026-01-14 · Miao Zhang, Kelly Chen, Md Mehrab Tanjim, Rumi Chunara

Identity-Robust Language Model Generation via Content Integrity Preservation

Large Language Model (LLM) outputs often vary across user sociodemographic attributes, leading to disparities in factual accuracy, utility, and safety, even for objective questions where demographic information is irrelevant. Unlike prior work on stereotypical or representational bias, this paper studies identity-dependent degradation...

💬 0 commentsarXiv:2601.09141v1PDF
0

Posted in cs.DS · 2026-01-14 · Gramoz Goranci, Monika Henzinger, Peter Kiss, Ali Momeni, Gernot Zöcklein

Dynamic Hierarchical $j$-Tree Decomposition and Its Applications

We develop a new algorithmic framework for designing approximation algorithms for cut-based optimization problems on capacitated undirected graphs that undergo edge insertions and deletions. Specifically, our framework dynamically maintains a variant of the hierarchical $j$-tree decomposition of [Madry FOCS'10], achieving a...

💬 0 commentsarXiv:2601.09139v1PDF