Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
LaneTCA: Enhancing Video Lane Detection with Temporal Context Aggregation
Keyi Zhou, Li Li, Wengang Zhou +3
In video lane detection, there are rich temporal contexts among successive frames, which is under-explored in existing lane detectors. In this work, we propose LaneTCA to bridge th…
cs.CV2023
Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs
Yonghui Wang, Wengang Zhou, Hao Feng +2
In the field of document understanding, significant advances have been made in the fine-tuning of Multimodal Large Language Models (MLLMs) with instruction-following data. Neverthe…