106 citations · 115 across the 10 of their papers we have counts for
10 papers
SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing
Rong-Cheng Tu, Wenhao Sun, Zhao Jin +3
While open-source video generation and editing models have made significant progress, individual models are typically limited to specific tasks, failing to meet the diverse needs o…
A Survey on Vision Autoregressive Model
Kai Jiang, Jiaxing Huang
Autoregressive models have demonstrated great performance in natural language processing (NLP) with impressive scalability, adaptability and generalizability. Inspired by their not…
Open-Vocabulary Object Detection via Language Hierarchy
Jiaxing Huang, Jingyi Zhang, Kai Jiang +1
Recent studies on generalizable object detection have attracted increasing attention with additional weak supervision from large-scale datasets with image-level labels. However, we…
Historical Test-time Prompt Tuning for Vision Foundation Models
Jingyi Zhang, Jiaxing Huang, Xiaoqin Zhang +2
Test-time prompt tuning, which learns prompts online with unlabelled test samples during the inference stage, has demonstrated great potential by learning effective prompts on-the-…
A Survey on Evaluation of Multimodal Large Language Models
Jiaxing Huang, Jingyi Zhang
Multimodal Large Language Models (MLLMs) mimic human perception and reasoning system by integrating powerful Large Language Models (LLMs) with various modality encoders (e.g., visi…
Learning to Prompt Segment Anything Models
Jiaxing Huang, Kai Jiang, Jingyi Zhang +4
Segment Anything Models (SAMs) like SEEM and SAM have demonstrated great potential in learning to segment anything. The core design of SAMs lies with Promptable Segmentation, which…