173 citations · 340 across the 15 of their papers we have counts for
40 papers
P-RAG: Progressive Retrieval Augmented Generation For Planning on Embodied Everyday Task
Weiye Xu, Min Wang, Wengang Zhou +1
Embodied Everyday Task is a popular task in the embodied AI community, requiring agents to make a sequence of actions based on natural language instructions and visual observations…
AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
Yonghui Wang, Wengang Zhou, Hao Feng +1
Over the past few years, the advancement of Multimodal Large Language Models (MLLMs) has captured the wide interest of researchers, leading to numerous innovations to enhance MLLMs…
LaneTCA: Enhancing Video Lane Detection with Temporal Context Aggregation
Keyi Zhou, Li Li, Wengang Zhou +3
In video lane detection, there are rich temporal contexts among successive frames, which is under-explored in existing lane detectors. In this work, we propose LaneTCA to bridge th…
Scaling up Multimodal Pre-training for Sign Language Understanding
Wengang Zhou, Weichao Zhao, Hezhen Hu +2
Sign language serves as the primary meaning of communication for the deaf-mute community. Different from spoken language, it commonly conveys information by the collaboration of ma…
SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection
Yonghui Wang, Shaokai Liu, Li Li +2
Shadow detection is a fundamental and challenging task in many computer vision applications. Intuitively, most shadows come from the occlusion of light by the object itself, result…
Progressive Multi-modal Conditional Prompt Tuning
Xiaoyu Qiu, Hao Feng, Yuechen Wang +2
Pre-trained vision-language models (VLMs) have shown remarkable generalization capabilities via prompting, which leverages VLMs as knowledge bases to extract information beneficial…