150 citations · 309 across the 7 of their papers we have counts for
11 papers
FlowZero: Zero-Shot Text-to-Video Synthesis with LLM-Driven Dynamic Scene Syntax
Yu Lu, Linchao Zhu, Hehe Fan +1
Text-to-video (T2V) generation is a rapidly growing research area that aims to translate the scenes, objects, and actions within complex video text into a sequence of coherent visu…
Can We Solve 3D Vision Tasks Starting from A 2D Vision Transformer?
Yi Wang, Zhiwen Fan, Tianlong Chen +2
Vision Transformers (ViTs) have proven to be effective, in solving 2D image understanding tasks by training over large-scale image datasets; and meanwhile as a somehow separate tra…
SEFormer: Structure Embedding Transformer for 3D Object Detection
Xiaoyu Feng, Heming Du, Yueqi Duan +2
Effectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a key challenge to 3D object detection on point cloud. Recently, Transfo…
PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences
Hehe Fan, Xin Yu, Yuhang Ding +2
Point cloud sequences are irregular and unordered in the spatial dimension while exhibiting regularities and order in the temporal dimension. Therefore, existing grid based convolu…
PointRNN: Point Recurrent Neural Network for Moving Point Cloud Processing
Hehe Fan, Yi Yang
In this paper, we introduce a Point Recurrent Neural Network (PointRNN) for moving point cloud processing. At each time step, PointRNN takes point coordinates $\boldsymbol{P} \in \…
Attract or Distract: Exploit the Margin of Open Set
Qianyu Feng, Guoliang Kang, Hehe Fan +1
Open set domain adaptation aims to diminish the domain shift across domains, with partially shared classes. There exist unknown target samples out of the knowledge of source domain…