collaborators

41 papers

cs.CL2026

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Deyao Zhu, Xin Zhou, Shengling Qin +44

Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less unders…

cs.CL2026

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

Guibin Zhang, Xun Xu, Yanwei Yue +4

Memory has become a standard substrate for self-evolving agents, yet retaining experience is not the same as learning how to evolve through it. Existing memory agents can store tra…

cs.CL2026

SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

Jiamin Chen, Yidi Wu, Qiexiang Wang +6

Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot resolve. Rather than construct…

cs.CL2026

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

Jiamin Chen, Qianben Chen, Jiawen Zhang +5

Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinematic control, audio, and cros…

cs.AI2026

PersonaDual: Balancing Personalization and Objectivity via Adaptive Reasoning

Xiaoyou Liu, Xinyi Mou, Shengbin Yue +5

As users increasingly expect LLMs to align with their preferences, personalized information becomes valuable. However, personalized information can be a double-edged sword: it can…

cs.CL2026

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies

Xavier Hu, Jinxiang Xia, Shengze Xu +13

Long-horizon planning is widely recognized as a core capability of autonomous LLM-based agents; however, current evaluation frameworks suffer from being largely episodic, domain-sp…