works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.CL2026

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

Zirui Song, Huaxing Liu, Xiang Wang +8

Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outpu…

cs.LG2026

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Zixiang Xu, Sixian Li, Huaxing Liu +4

The paper investigates how biases in large language models used as judges are reflected in their hidden activations, identifying low-dimensional subspaces that encode bias and show…

cs.CL2026

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression

Lingjie Zeng, Xiaofan Chen, Yanbo Wang +1

Long chain-of-thought (Long-CoT) reasoning models have motivated a growing body of work on compressing reasoning traces to reduce inference cost, yet existing evaluations focus alm…

cs.AI2026

M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities

Ao Li, Jinghui Zhang, Luyu Li +10

As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to improve answer quality and support complex…

cs.CL2025

Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies

Zirui Song, Yuan Huang, Junchang Liu +7

Social deduction games like Werewolf combine language, reasoning, and strategy, providing a testbed for studying natural language and social intelligence. However, most studies red…

cs.CL2025

Do LLMs "Feel"? Emotion Circuits Discovery and Control

Chenxi Wang, Yixuan Zhang, Ruiji Yu +8

As the demand for emotional intelligence in large language models (LLMs) grows, a key challenge lies in understanding the internal mechanisms that give rise to emotional expression…