papers

Publications (32)

cs.CL2026

Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding

Yanzheng Xiang, Lan Wei, Yizhen Yao +8

Parallel diffusion decoding can accelerate diffusion language model inference by unmasking multiple tokens per step, but aggressive parallelism often harms quality. Revocable decod…

cs.CL2026

Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation

Zhanghao Hu, Qinglin Zhu, Runcong Zhao +4

Standard Retrieval Augmented Generation (RAG) is poorly matched to agent memory. Unlike large heterogeneous corpora, agent memory forms a bounded and coherent interaction stream in…

cs.CL2026

AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor

Shu Yang, Jingyu Hu, Tong Li +3

We introduce AutoMonitor-Bench, the first benchmark designed to systematically evaluate the reliability of LLM-based misbehavior monitors across diverse tasks and failure modes. Au…

cs.HC2026

Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations

Hainiu Xu, Zhaoyue Sun, Hanqi Yan +3

Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cognitive appraisals, the subjecti…

cs.CL2025

CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation

Zhenyi Shen, Hanqi Yan, Linhai Zhang +3

Chain-of-Thought (CoT) reasoning enhances Large Language Models (LLMs) by encouraging step-by-step reasoning in natural language. However, leveraging a latent continuous space for…

cs.CV2026

Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models

Hanqi Yan, Xiangxiang Cui, Lu Yin +4

The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, modality gaps persist in existing align…

cs.IR2025

GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery

Italo Luis da Silva, Hanqi Yan, Lin Gui +1

Large Language Models (LLMs) show strong reasoning and text generation capabilities, prompting their use in scientific literature analysis, including novelty assessment. While eval…

cs.CL2026

GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics

Yujing Wang, Yuanbang Liang, Yukun Lai +2

Detecting whether a model's internal knowledge is sufficient to correctly answer a given question is a fundamental challenge in deploying responsible LLMs. In addition to verbalisi…

cs.CL2024

The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and Analysis

Yuxiang Zhou, Jiazheng Li, Yanzheng Xiang +3

Understanding in-context learning (ICL) capability that enables large language models (LLMs) to excel in proficiency through demonstration examples is of utmost importance. This im…

cs.LG2026

Large Language Models Hack Rewards, and Society

Wei Liu, Xinyi Mou, Hanqi Yan +2

Reinforcement learning (RL) has become a dominant post-training paradigm, enabling large language models (LLMs) to learn from rewards. We observe that societal regulations are stru…

cs.CL2024

Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models

Yanzheng Xiang, Hanqi Yan, Lin Gui +1

In-context learning has become a popular paradigm in natural language processing. However, its performance can be significantly influenced by the order of in-context demonstration…

cs.CL2025

Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration

Qinglin Zhu, Runcong Zhao, Hanqi Yan +3

Large Language Models (LLMs) struggle with complex reasoning due to limited diversity and inefficient search. We propose Soft Reasoning, an embedding-based search framework that op…

cs.CL2024

Weak Reward Model Transforms Generative Models into Robust Causal Event Extraction Systems

Italo Luis da Silva, Hanqi Yan, Lin Gui +1

The inherent ambiguity of cause and effect boundaries poses a challenge in evaluating causal event extraction tasks. Traditional metrics like Exact Match and BertScore poorly refle…

cs.CL2025

Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering

Zhanghao Hu, Hanqi Yan, Qinglin Zhu +3

Large language models have recently pushed open domain question answering (ODQA) to new frontiers. However, prevailing retriever-reader pipelines often depend on multiple rounds of…

cs.IR2023

Tracking Brand-Associated Polarity-Bearing Topics in User Reviews

Runcong Zhao, Lin Gui, Hanqi Yan +1

Monitoring online customer reviews is important for business organisations to measure customer satisfaction and better manage their reputations. In this paper, we propose a novel d…

cs.CL2025

SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers

Yanzheng Xiang, Hanqi Yan, Shuyin Ouyang +2

This study evaluates large language models (LLMs) in generating code from algorithm descriptions in recent NLP papers. The task requires two key competencies: (1) algorithm compreh…

cs.LG2026

PreAct-Bench: Benchmarking Predictive Monitoring in LLMs

Hainiu Xu, Italo Luis da Silva, Jiangnan Ye +7

Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objective. While existing safety rese…

cs.AI2025

Constrain Alignment with Sparse Autoencoders

Qingyu Yin, Chak Tou Leong, Minjun Zhu +7

The alignment of large language models (LLMs) with human preferences remains a key challenge. While post-training techniques like Reinforcement Learning from Human Feedback (RLHF)…

physics.soc-ph2026

AI adoption induces divergent net energy changes across economic sectors

Wei He, Daoping Wang, Hanqi Yan +2

Energy planning for artificial intelligence focuses on data-centre electricity, missing the induced operational energy change caused by the deployment of AI in commercial buildings…

cs.CL2023

Position Bias Mitigation: A Knowledge-Aware Graph Model for Emotion Cause Extraction

Hanqi Yan, Lin Gui, Gabriele Pergola +1

The Emotion Cause Extraction (ECE)} task aims to identify clauses which contain emotion-evoking information for a particular emotion expressed in text. We observe that a widely-use…

cs.CL2022

Hierarchical Interpretation of Neural Text Classification

Hanqi Yan, Lin Gui, Yulan He

Recent years have witnessed increasing interests in developing interpretable models in Natural Language Processing (NLP). Most existing models aim at identifying input features suc…

cs.CL2026

Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

Zhengyi Zhao, Shubo Zhang, Zezhong Wang +6

When asked what a meme or sarcastic post means, Large Vision Language Models (LVLMs) tend to describe what the image shows rather than what the author is trying to communicate. Sta…

cs.CL2024

Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective

Hanqi Yan, Yanzheng Xiang, Guangyi Chen +3

To better interpret the intrinsic mechanism of large language models (LLMs), recent studies focus on monosemanticity on its basic units. A monosemantic neuron is dedicated to a sin…

cs.CL2026

When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignment

Hanqi Yan, Hainiu Xu, Siya Qi +2

With the growing accessibility and wide adoption of large language models, concerns about their safety and alignment with human values have become paramount. In this paper, we iden…

cs.CL2025

Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score

Zhanghao Hu, Qinglin Zhu, Siya Qi +3

Large Language Models (LLMs) have shown improved generation performance through retrieval-augmented generation (RAG) following the retriever-reader paradigm, which supplements mode…

cs.CL2024

Mirror: A Multiple-perspective Self-Reflection Method for Knowledge-rich Reasoning

Hanqi Yan, Qinglin Zhu, Xinyu Wang +2

While Large language models (LLMs) have the capability to iteratively reflect on their own outputs, recent studies have observed their struggles with knowledge-rich problems withou…

cs.LG2024

Counterfactual Generation with Identifiability Guarantees

Hanqi Yan, Lingjing Kong, Lin Gui +4

Counterfactual generation lies at the core of various machine learning tasks, including image translation and controllable text generation. This generation process usually requires…

cs.LG2026

Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning

Lukas Twist, Shu Yang, Hanqi Yan +4

Large Language Models (LLMs) increasingly exhibit strong reasoning abilities, often attributed to their capacity to generate chain-of-thought-style intermediate reasoning. Recent w…

cs.CL2026

Fix the Structural Bottleneck: Context Compression via Explicit Information Transmission

Jiangnan Ye, Hanqi Yan, Zhenyi Shen +3

Long-context LLM agents often struggle with growing token, memory, and latency costs, making efficient context compression essential for practical deployment. Existing LLM-as-a-com…

cs.CL2023

Addressing Token Uniformity in Transformers via Singular Value Transformation

Hanqi Yan, Lin Gui, Wenjie Li +1

Token uniformity is commonly observed in transformer-based models, in which different tokens share a large proportion of similar information after going through stacked multiple se…

cs.IR2024

Explainable Recommender with Geometric Information Bottleneck

Hanqi Yan, Lin Gui, Menghan Wang +2

Explainable recommender systems can explain their recommendation decisions, enhancing user trust in the systems. Most explainable recommender systems either rely on human-annotated…

cs.CL2023

Distinguishability Calibration to In-Context Learning

Hongjing Li, Hanqi Yan, Yanran Li +3

Recent years have witnessed increasing interests in prompt-based learning in which models can be trained on only a few annotated instances, making them suitable in low-resource set…