8 papers
MLGIB: Multi-Label Graph Information Bottleneck for Expressive and Robust Message Passing
Chaokai Wu, Haofu Shi, Ningxuan Ma +2
Graph Neural Networks (GNNs) suffer from over-squashing in deep message passing, where information from exponentially growing neighborhoods is compressed into fixed-dimensional rep…
GRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training
Junjie Li, Ziao Wang, NingXuan Ma +2
Existing reasoning data curation pipelines score whole samples, treating every intermediate step as equally valuable. In reality, steps within a trace contribute very unevenly, and…
expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling
Mingxiong Lin, Zhangquan Gong, Maowen Tang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, where Group Relative Policy Optimization (GRPO) serves as the…
fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum
Mingxiong Lin, Zhangquan Gong, Maowen Tang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Optimization (GRPO) serving as the…
Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
Junjie Li, Ziao Wang, Jianghong Ma +1
Large vision-language models (VLMs) achieve strong benchmark performance, but controlling their behavior through instruction tuning remains difficult. Reducing the budget of instru…
Diversity Recommendation via Causal Deconfounding of Co-purchase Relations and Counterfactual Exposure
Jingmao Zhang, Zhiting Zhao, Yunqi Lin +4
Beyond user-item modeling, item-to-item relationships are increasingly used to enhance recommendation. However, common methods largely rely on co-occurrence, making them prone to i…