collaborators

10 papers

cs.AI2026

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…

cs.RO2026

HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

Xiaoquan Sun, Ruijian Zhang, Chen Cao +12

World Action Models (WAMs) have emerged as a new powerful paradigm for embodied intelligence, learning action-relevant visual dynamics that significantly enhance generalization and…

cs.CV2026

Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling

Yihang Wu, Yihang Sun, Shaofeng Zhang +4

Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial inform…

cs.AI2026

Useful Memories Become Faulty When Continuously Updated by LLMs

Dylan Zhang, Yanshan Lin, Zhengkun Wu +4

Learning from past experience benefits from two complementary forms of memory: episodic traces -- raw trajectories of what happened -- and consolidated abstractions distilled acros…

cs.LG2026

STAGE: Tackling Semantic Drift in Multimodal Federated Graph Learning

Zekai Chen, Xun Wu, Xunkai Li +3

Federated graph learning (FGL) enables collaborative training on graph data across multiple clients. As graph data increasingly contain multimodal node attributes such as text and…

cs.AI2026

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

Yusheng Li, Tianjun Feng, Yunfeng Chen +6

LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on this substrate for safety-critic…