activity
20182026
most citedSEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

52 citations · 165 across the 41 of their papers we have counts for

collaborators
Showing 2022Show all

6 papers · 1 filter

cs.LG2022

Policy Adaptation from Foundation Model Feedback

Yuying Ge, Annabella Macaluso, Li Erran Li +2

Recent progress on vision-language foundation models have brought significant advancement to building general-purpose robots. By using the pre-trained models to encode the scene an…

cs.CV2022

Learning Transferable Spatiotemporal Representations from Natural Script Knowledge

Ziyun Zeng, Yuying Ge, Xihui Liu +4

Pre-training on large-scale video data has become a common recipe for learning transferable spatiotemporal representations in recent years. Despite some progress, existing methods…

cs.CV2022

MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval

Yuying Ge, Yixiao Ge, Xihui Liu +5

Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast gl…

cs.CV2022★ 3 cited

All in One: Exploring Unified Video-Language Pre-training

Alex Jinpeng Wang, Yixiao Ge, Rui Yan +7

Mainstream Video-Language Pre-training models \cite{actbert,clipbert,violet} consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They purs…

cs.CV2022

MetaDance: Few-shot Dancing Video Retargeting via Temporal-aware Meta-learning

Yuying Ge, Yibing Song, Ruimao Zhang +1

Dancing video retargeting aims to synthesize a video that transfers the dance movements from a source video to a target person. Previous work need collect a several-minute-long vid…

cs.CV2022★ 1 cited

Bridging Video-text Retrieval with Multiple Choice Questions

Yuying Ge, Yixiao Ge, Xihui Liu +4

Pre-training a model to learn transferable video-text representation for retrieval has attracted a lot of attention in recent years. Previous dominant works mainly adopt two separa…