2 papers
cs.LG2026
GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning
Ting Zhou, Zhenqing Ling, Yiyang Zhao +2
Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identif…
cs.CV2024
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
Ting Zhou, Daoyuan Chen, Qirui Jiao +3
Evaluating the nuanced human-centric video understanding capabilities of Multimodal Large Language Models (MLLMs) remains a great challenge, as existing benchmarks often overlook t…