6 papers
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
Zhan'ao Yao, Liang Yin, Zhihao Gao +9
Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model…
MatMind: A Structure-Activity Knowledge-Driven Generative Foundation Model for Materials Science
Zhan'ao Yao, Boxuan Zhang, Jingyuan Shu +10
Progress in AI-driven crystal materials science has so far been carried by narrow architectures purpose-built for individual tasks -- graph neural networks for property prediction,…
Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation
Fei Ding, Yongkang Zhang, youwei wang +1
Reinforcement learning for multi-step reasoning with large language models (LLMs) typically relies on sparse terminal rewards, which creates a poorly conditioned credit-assignment…
Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning
Feiding, Yongkang Zhang, Yuhao Liao +10
Vision--language models (VLMs) are increasingly aligned via Group Relative Policy Optimization (GRPO)-style training. However, relying solely on terminal outcome rewards yields spa…
Deep Sparse Latent Feature Models for Knowledge Graph Completion
Haotian Li, Rui Zhang, Lingzhi Wang +6
Recent advances in knowledge graph completion (KGC) have emphasized text-based approaches to navigate the inherent complexities of large-scale knowledge graphs (KGs). While these m…
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
Fei Ding, Baiqiao Wang, Zijian Zeng +1
The Group Relative Policy Optimization (GRPO) algorithm has demonstrated considerable success in enhancing the reasoning capabilities of large language models (LLMs), as evidenced…