2 papers
cs.RO2026
CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning
Yang Chen, Ye-Xin Xie, Lirong Che +8
Agent systems powered by multimodal large language models (MLLMs) have advanced rapidly in recent years, yet existing embodied-agent benchmarks still lack fine-grained diagnostics…
cs.CL2026
DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation
Guhan Chen, Songtao Tian, Bohan Li +3
Iterative preference optimization is essential for aligning Large Language Models on mathematical reasoning tasks, yet its efficiency is often throttled by signal scarcity: as the…