2 papers
cs.CL2026
Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
Jianxiang Zang, Yongda Wei, Ruxue Bai +5
Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods focus solely on preference percept…
cs.CL2025
Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion
Jianxiang Zang, Meiling Ning, Yongda Wei +7
Recently, the concept of ``compression as intelligence'' has provided a novel informatics metric perspective for language models (LMs), emphasizing that highly structured represent…