activity
20212025
most citedContrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative Pretraining

31 citations · 36 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2025

Taming Teacher Forcing for Masked Autoregressive Video Generation

Deyu Zhou, Quan Sun, Yuang Peng +8

We introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation,…

cs.CL20231 cited

ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Liang Zhao, En Yu, Zheng Ge +8

Human-AI interactivity is a critical aspect that reflects the usability of multimodal large language models (MLLMs). However, existing end-to-end MLLMs only allow users to interact…

cs.CV2023

CORSD: Class-Oriented Relational Self Distillation

Muzhou Yu, Sia Huat Tan, Kailu Wu +3

Knowledge distillation conducts an effective model compression method while holding some limitations:(1) the feature based distillation methods only focus on distilling the feature…

cs.CV2023

CLIP-FO3D: Learning Free Open-world 3D Scene Representations from 2D Dense CLIP

Junbo Zhang, Runpei Dong, Kaisheng Ma

Training a 3D scene understanding model requires complicated human annotations, which are laborious to collect and result in a model only encoding close-set object semantics. In co…

cs.CV202331 cited

Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative Pretraining

Zekun Qi, Runpei Dong, Guofan Fan +4

Mainstream 3D representation learning approaches are built upon contrastive or generative modeling pretext tasks, where great improvements in performance on various downstream task…

cs.CV2022

Contrastive Deep Supervision

Linfeng Zhang, Xin Chen, Junbo Zhang +2

The success of deep learning is usually accompanied by the growth in neural network depth. However, the traditional training method only supervises the neural network at its last l…