collaborators

9 papers

cs.AI2026

Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization

Zhengtao Yao, Runhao Li, Xupeng Chen +12

Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can…

cs.CV2026

Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency

Junming Liu, Yuqi Li, Yifei Sun +4

Vision-Language Models (VLMs) have made striking progress, yet their spatial reasoning remains fragile: models that answer an original input correctly can still fail under paired t…

cs.AI2026

COMPOSITE-Stem

Kyle Waters, Lucas Nuzzi, Tadhg Looram +20

AI agents hold growing promise for accelerating scientific discovery; yet, a lack of frontier evaluations hinders adoption into real workflows. Expert-written benchmarks have prove…

cs.AI2026

Hierarchical Memory Orchestration for Personalized Persistent Agents

Junming Liu, Yifei Sun, Weihua Cheng +4

While long-term memory is essential for intelligent agents to maintain consistent historical awareness, the accumulation of extensive interaction data often leads to performance bo…

cs.CL2026

Hit-RAG: Learning to Reason with Long Contexts via Preference Alignment

Junming Liu, Yuqi Li, Shiping Wen +2

Despite the promise of Retrieval-Augmented Generation in grounding Multimodal Large Language Models with external knowledge, the transition to extensive contexts often leads to sig…

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…