10 citations · 17 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 3 cited
AMD: Automatic Multi-step Distillation of Large-scale Vision Models
Cheng Han, Qifan Wang, Sohail A. Dianat +6
Transformer-based architectures have become the de-facto standard models for diverse vision tasks owing to their superior performance. As the size of the models continues to scale…
cs.CV2024★ 4 cited
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
Jiamian Wang, Guohao Sun, Pichao Wang +5
The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, rel…
cs.CV2024★ 10 cited
Image Translation as Diffusion Visual Programmers
Cheng Han, James C. Liang, Qifan Wang +5
We introduce the novel Diffusion Visual Programmer (DVP), a neuro-symbolic image translation framework. Our proposed DVP seamlessly embeds a condition-flexible diffusion model with…