4 papers
Triviality Corrected Endogenous Reward
Xinda Wang, Zhengxu Hou, Yangshijie Zhang +6
Reinforcement learning for open-ended text generation is constrained by the lack of verifiable rewards, necessitating reliance on judge models that require either annotated data or…
MoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start Recommendation
Jialin Liu, Zhaorui Zhang, Ray C. C. Cheung
Graph neural networks (GNNs) have revolutionized recommender systems by effectively modeling complex user-item interactions, yet data sparsity and the item cold-start problem signi…
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
Shaobo Wang, Xuan Ouyang, Tianyi Xu +9
As high-quality public text approaches exhaustion, a phenomenon known as the Data Wall, pre-training is shifting from more tokens to better tokens. However, existing methods either…
Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent
Yangshijie Zhang, Xinda Wang, Jialin Liu +3
With social media growth, users employ stylistic fonts and font-like emoji to express individuality, creating visually appealing text that remains human-readable. However, these fo…