1 citations · 1 across the 12 of their papers we have counts for
8 papers · 1 filter
Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
Pyrros Koussios, Chenhao Li, Xin Chen +1
Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, co…
Gaussian Process Bandit Optimization with Machine Learning Predictions and Application to Hypothesis Generation
Xin Jennifer Chen, Yunjin Tong
Many real-world optimization problems involve an expensive ground-truth oracle (e.g., human evaluation, physical experiments) and a cheap, low-fidelity prediction oracle (e.g., mac…
Dual-Phase LLM Reasoning: Self-Evolved Mathematical Frameworks
ShaoZhen Liu, Xinting Huang, Houwen Peng +4
In recent years, large language models (LLMs) have demonstrated significant potential in complex reasoning tasks like mathematical problem-solving. However, existing research predo…
All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training
Chi Liu, Xin Chen
Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). However, GRPO inherits PPO's token-level…
EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuning
Jing-Cheng Pang, Liu Sun, Chang Zhou +10
Domain-specific large language models (LLMs), typically developed by fine-tuning a pre-trained general-purpose LLM on specialized datasets, represent a significant advancement in a…
Learning Safety Constraints for Large Language Models
Xin Chen, Yarden As, Andreas Krause
Large language models (LLMs) have emerged as powerful tools but pose significant safety risks through harmful outputs and vulnerability to adversarial attacks. We propose SaP, shor…