most citedGroupLane: End-to-End 3D Lane Detection with Channel-wise Grouping

1 citations · 3 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2024

OneChart: Purify the Chart Structural Extraction via One Auxiliary Token

Jinyue Chen, Lingyu Kong, Haoran Wei +6

Chart parsing poses a significant challenge due to the diversity of styles, values, texts, and so forth. Even advanced large vision-language models (LVLMs) with billions of paramet…

cs.CV20241 cited

Small Language Model Meets with Reinforced Vision Vocabulary

Haoran Wei, Lingyu Kong, Jinyue Chen +6

Playing Large Vision Language Models (LVLMs) in 2023 is trendy among the AI community. However, the relatively large number of parameters (more than 7B) of popular LVLMs makes it d…

cs.CV2023

SCSC: Spatial Cross-scale Convolution Module to Strengthen both CNNs and Transformers

Xijun Wang, Xiaojie Chu, Chunrui Han +1

This paper presents a module, Spatial Cross-scale Convolution (SCSC), which is verified to be effective in improving both CNNs and Transformers. Nowadays, CNNs and Transformers hav…

cs.CL20231 cited

ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Liang Zhao, En Yu, Zheng Ge +8

Human-AI interactivity is a critical aspect that reflects the usability of multimodal large language models (MLLMs). However, existing end-to-end MLLMs only allow users to interact…

cs.CV20231 cited

GroupLane: End-to-End 3D Lane Detection with Channel-wise Grouping

Zhuoling Li, Chunrui Han, Zheng Ge +5

Efficiency is quite important for 3D lane detection due to practical deployment demand. In this work, we propose a simple, fast, and end-to-end detector that still maintains high d…

cs.CV2023

Triplet Knowledge Distillation

Xijun Wang, Dongyang Liu, Meina Kan +3

In Knowledge Distillation, the teacher is generally much larger than the student, making the solution of the teacher likely to be difficult for the student to learn. To ease the mi…