4 papers
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
Alan Li, Yixin Liu, Arpan Sarkar +2
Scientific problem solving poses unique challenges for LLMs, requiring both deep domain knowledge and the ability to apply such knowledge through complex reasoning. While automated…
LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning
Xudong Chen, Yixin Liu, Hua Wei +1
Large language models (LLMs) have become a strong foundation for multi-agent systems, but their effectiveness depends heavily on orchestration design. Across different tasks, role…
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
Yixin Liu, Argyris Oikonomou, Weiqiang Zheng +2
Many alignment methods, including reinforcement learning from human feedback (RLHF), rely on the Bradley-Terry reward assumption, which is not always sufficient to capture the full…
Calibrating Long-form Generations from Large Language Models
Yukun Huang, Yixin Liu, Raghuveer Thirukovalluru +2
To enhance Large Language Models' (LLMs) reliability, calibration is essential -- the model's assessed confidence scores should align with the actual likelihood of its responses be…