9 papers
DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity
Fengyuan Liu, Yongliang Miao, Zirui He +3
Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framewo…
Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
Haiyan Zhao, Zirui He, Guanchu Wang +3
Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each model explains only its own ac…
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
Zirui He, Haiyan Zhao, Yingcong Li +2
Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memorized solutions appear as genuine…
Rep2Text: Decoding Full Text from a Single LLM Token Representation
Haiyan Zhao, Zirui He, Yiming Tang +4
Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In this work, we investigate a fundamental…
RealRoute: Dynamic Query Routing System via Retrieve-then-Verify Paradigm
Jiahe Liu, Qinkai Yu, Jingcheng Niu +5
Despite the success of Retrieval-Augmented Generation (RAG) in grounding LLMs with external knowledge, its application over heterogeneous sources (e.g., private databases, global c…
FinAnchor: Aligned Multi-Model Representations for Financial Prediction
Zirui He, Huopu Zhang, Yanguang Liu +2
Financial prediction from long documents involves significant challenges, as actionable signals are often sparse and obscured by noise, and the optimal LLM for generating embedding…