collaborators

9 papers

cs.LG2026

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity

Fengyuan Liu, Yongliang Miao, Zirui He +3

Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framewo…

cs.CL2026

Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

Haiyan Zhao, Zirui He, Guanchu Wang +3

Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each model explains only its own ac…

cs.CL2026

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

Zirui He, Haiyan Zhao, Yingcong Li +2

Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memorized solutions appear as genuine…

cs.CL2026

Rep2Text: Decoding Full Text from a Single LLM Token Representation

Haiyan Zhao, Zirui He, Yiming Tang +4

Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In this work, we investigate a fundamental…

cs.IR2026

RealRoute: Dynamic Query Routing System via Retrieve-then-Verify Paradigm

Jiahe Liu, Qinkai Yu, Jingcheng Niu +5

Despite the success of Retrieval-Augmented Generation (RAG) in grounding LLMs with external knowledge, its application over heterogeneous sources (e.g., private databases, global c…

cs.CL2026

FinAnchor: Aligned Multi-Model Representations for Financial Prediction

Zirui He, Huopu Zhang, Yanguang Liu +2

Financial prediction from long documents involves significant challenges, as actionable signals are often sparse and obscured by noise, and the optimal LLM for generating embedding…