4 papers
Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
Jianxiang Zang
The reward model (RM), as the core component of reinforcement learning from human feedback (RLHF) for large language models (LLMs), responsible for providing reward signals to gene…
Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
Meiling Ning, Zhongbao Zhang, Junda Ye +2
The emergence of LM-based judging reward modeling, represented by generative reward models, has successfully made reinforcement learning from AI feedback (RLAIF) efficient and scal…
Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion
Jianxiang Zang, Meiling Ning, Yongda Wei +7
Recently, the concept of ``compression as intelligence'' has provided a novel informatics metric perspective for language models (LMs), emphasizing that highly structured represent…
S2Sent: Nested Selectivity Aware Sentence Representation Learning
Jianxiang Zang, Nijia Mo, Yonda Wei +2
The combination of Transformer-based encoders with contrastive learning represents the current mainstream paradigm for sentence representation learning. This paradigm is typically…