collaborators

14 papers

cs.LG2026

MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer

Xuefei Wang, Jialu Wang, Fengbo Zhang +6

Multi-agent systems (MAS) powered by large language models (LLMs) have emerged as a powerful paradigm for complex problem solving, where performance critically depends on the under…

cs.CV2026

URoPE: Universal Relative Position Embedding across Geometric Spaces

Yichen Xie, Depu Meng, Chensheng Peng +4

Relative position embedding has become a standard mechanism for encoding positional information in Transformers. However, existing formulations are typically limited to a fixed geo…

cs.AI2026

NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning

Ishaan Rawal, Shubh Gupta, Yihan Hu +1

Vision-Language-Action (VLA) models are advancing autonomous driving by replacing modular pipelines with unified end-to-end architectures. However, current VLAs face two expensive…

cs.LG2026

Ratio-Variance Regularized Policy Optimization

Yu Luo, Shuo Han, Yihan Hu +5

Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscriminately truncating high-return…

cs.CV2026

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

Tianshuo Xu, Yichen Xie, Depu Meng +5

Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity p…

cs.CV2026

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

Bo Jiang, Depu Meng, Yihan Hu +3

Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedie…