4 citations · 4 across the 4 of their papers we have counts for
16 papers · 1 filter
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
Zirui He, Haiyan Zhao, Yingcong Li +2
Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memorized solutions appear as genuine…
Rep2Text: Decoding Full Text from a Single LLM Token Representation
Haiyan Zhao, Zirui He, Yiming Tang +4
Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In this work, we investigate a fundamental…
SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
Jiaojiao Han, Wujiang Xu, Mingyu Jin +1
Large language models (LLMs) have achieved remarkable progress, yet their internal mechanisms remain largely opaque, posing a significant challenge to their safe and reliable deplo…
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
Haiyan Zhao, Xuansheng Wu, Fan Yang +3
Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder…
Knowledge Graph Large Language Model (KG-LLM) for Link Prediction
Dong Shu, Tianle Chen, Mingyu Jin +3
The task of multi-hop link prediction within knowledge graphs (KGs) stands as a challenge in the field of knowledge graph analysis, as it requires the model to reason through and u…
What if LLMs Have Different World Views: Simulating Alien Civilizations with LLM-based Agents
Zhaoqian Xue, Beichen Wang, Suiyuan Zhu +5
This study introduces "CosmoAgent," an innovative artificial intelligence system that utilizes Large Language Models (LLMs) to simulate complex interactions between human and extra…