activity
20232025
most citedVideoVLA: Video Generators Can Be Generalizable Robot Manipulators

1 citations · 1 across the 4 of their papers we have counts for

collaborators

8 papers

cs.RO20251 cited

VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

Yichao Shen, Fangyun Wei, Zhiying Du +5

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language…

cs.CV2025

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models

Tianci Bi, Xiaoyi Zhang, Yan Lu +1

The performance of Latent Diffusion Models (LDMs) is critically dependent on the quality of their visual tokenizers. While recent works have explored incorporating Vision Foundatio…

cs.LG2024

A General Theory for Compositional Generalization

Jingwen Fu, Zhizheng Zhang, Yan Lu +1

Compositional Generalization (CG) embodies the ability to comprehend novel combinations of familiar concepts, representing a significant cognitive leap in human intellectual advanc…

cs.CV2024

Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis

Tianci Bi, Xiaoyi Zhang, Zhizheng Zhang +4

Significant progress has been made in scene text detection models since the rise of deep learning, but scene text layout analysis, which aims to group detected text instances as pa…

cs.CV2024

Diffusion Model with Cross Attention as an Inductive Bias for Disentanglement

Tao Yang, Cuiling Lan, Yan Lu +1

Disentangled representation learning strives to extract the intrinsic factors within observed data. Factorizing these representations in an unsupervised manner is notably challengi…

cs.LG2023

Closing the Gap Between the Upper Bound and the Lower Bound of Adam's Iteration Complexity

Bohan Wang, Jingwen Fu, Huishuai Zhang +2

Recently, Arjevani et al. [1] established a lower bound of iteration complexity for the first-order optimization under an -smooth condition and a bounded noise variance assumpti…