1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2024
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
Yu-Jhe Li, Xinyang Zhang, Kun Wan +3
We tackle the challenge of open-vocabulary segmentation, where we need to identify objects from a wide range of categories in different environments, using text prompts as our inpu…
cs.CV2024
Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
Zeliang Zhang, Phu Pham, Wentian Zhao +6
By treating visual tokens from visual encoders as text tokens, Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse visual understanding tasks,…
cs.CL2024★ 1 cited
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
Shijian Deng, Wentian Zhao, Yu-Jhe Li +4
Self-improvement in multimodal large language models (MLLMs) is crucial for enhancing their reliability and robustness. However, current methods often rely heavily on MLLMs themsel…