5 citations · 18 across the 9 of their papers we have counts for
9 papers
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
Wenqiao Zhang, Tianwei Lin, Jiang Liu +10
Recent advancements indicate that scaling up Multimodal Large Language Models (MLLMs) effectively enhances performance on downstream multimodal tasks. The prevailing MLLM paradigm,…
Unified Generative Modeling of 3D Molecules via Bayesian Flow Networks
Yuxuan Song, Jingjing Gong, Yanru Qu +4
Advanced generative model (e.g., diffusion model) derived from simplified continuity assumptions of data distribution, though showing promising progress, has been difficult to appl…
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
Xiaohu Huang, Hao Zhou, Kun Yao +1
In this paper, we introduce FROSTER, an effective framework for open-vocabulary action recognition. The CLIP model has achieved remarkable success in a range of image-based tasks,…
Revisiting the Markov Property for Machine Translation
Cunxiao Du, Hao Zhou, Zhaopeng Tu +1
In this paper, we re-examine the Markov property in the context of neural machine translation. We design a Markov Autoregressive Transformer~(MAT) and undertake a comprehensive ass…
HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception
Junkun Yuan, Xinyu Zhang, Hao Zhou +12
Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting…
Sign Language Translation with Iterative Prototype
Huijie Yao, Wengang Zhou, Hao Feng +3
This paper presents IP-SLT, a simple yet effective framework for sign language translation (SLT). Our IP-SLT adopts a recurrent structure and enhances the semantic representation (…