activity
20242026
most citedX-GridAgent: An LLM-Powered Agentic AI System for Assisting Power Grid Analysis

1 citations · 1 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning

Pyrros Koussios, Chenhao Li, Xin Chen +1

Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, co…

cs.LG2026

Gaussian Process Bandit Optimization with Machine Learning Predictions and Application to Hypothesis Generation

Xin Jennifer Chen, Yunjin Tong

Many real-world optimization problems involve an expensive ground-truth oracle (e.g., human evaluation, physical experiments) and a cheap, low-fidelity prediction oracle (e.g., mac…

cs.LG2026

Dual-Phase LLM Reasoning: Self-Evolved Mathematical Frameworks

ShaoZhen Liu, Xinting Huang, Houwen Peng +4

In recent years, large language models (LLMs) have demonstrated significant potential in complex reasoning tasks like mathematical problem-solving. However, existing research predo…

cs.LG2026

All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training

Chi Liu, Xin Chen

Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). However, GRPO inherits PPO's token-level…

cs.LG2026

EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuning

Jing-Cheng Pang, Liu Sun, Chang Zhou +10

Domain-specific large language models (LLMs), typically developed by fine-tuning a pre-trained general-purpose LLM on specialized datasets, represent a significant advancement in a…

cs.LG2025

Learning Safety Constraints for Large Language Models

Xin Chen, Yarden As, Andreas Krause

Large language models (LLMs) have emerged as powerful tools but pose significant safety risks through harmful outputs and vulnerability to adversarial attacks. We propose SaP, shor…