7 papers
Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
Jianxiang Zang, Yongda Wei, Ruxue Bai +5
Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods focus solely on preference percept…
Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
Jianxiang Zang
The reward model (RM), as the core component of reinforcement learning from human feedback (RLHF) for large language models (LLMs), responsible for providing reward signals to gene…
Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion
Jianxiang Zang, Meiling Ning, Yongda Wei +7
Recently, the concept of ``compression as intelligence'' has provided a novel informatics metric perspective for language models (LMs), emphasizing that highly structured represent…
S2Sent: Nested Selectivity Aware Sentence Representation Learning
Jianxiang Zang, Nijia Mo, Yonda Wei +2
The combination of Transformer-based encoders with contrastive learning represents the current mainstream paradigm for sentence representation learning. This paradigm is typically…
Multi-Programming Language Sandbox for LLMs
Shihan Dou, Jiazheng Zhang, Jianxiang Zang +25
We introduce MPLSandbox, an out-of-the-box multi-programming language sandbox designed to provide unified and comprehensive feedback from compiler and analysis tools for Large Lang…
Modeling Selective Feature Attention for Representation-based Siamese Text Matching
Jianxiang Zang, Hui Liu
Representation-based Siamese networks have risen to popularity in lightweight text matching due to their low deployment and inference costs. While word-level attention mechanisms h…