2 papers
cs.CV2026
Rethinking Video Generation Model for the Embodied World
Yufan Deng, Zilin Pan, Hongyu Zhang +6
Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and act…
cs.CV2026
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
Kewei Zhang, Ye Huang, Yufan Deng +5
While the Transformer architecture dominates many fields, its quadratic self-attention complexity hinders its use in large-scale applications. Linear attention offers an efficient…