14 papers
MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer
Xuefei Wang, Jialu Wang, Fengbo Zhang +6
Multi-agent systems (MAS) powered by large language models (LLMs) have emerged as a powerful paradigm for complex problem solving, where performance critically depends on the under…
URoPE: Universal Relative Position Embedding across Geometric Spaces
Yichen Xie, Depu Meng, Chensheng Peng +4
Relative position embedding has become a standard mechanism for encoding positional information in Transformers. However, existing formulations are typically limited to a fixed geo…
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
Ishaan Rawal, Shubh Gupta, Yihan Hu +1
Vision-Language-Action (VLA) models are advancing autonomous driving by replacing modular pipelines with unified end-to-end architectures. However, current VLAs face two expensive…
Ratio-Variance Regularized Policy Optimization
Yu Luo, Shuo Han, Yihan Hu +5
Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscriminately truncating high-return…
Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
Tianshuo Xu, Yichen Xie, Depu Meng +5
Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity p…
LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation
Bo Jiang, Depu Meng, Yihan Hu +3
Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedie…