activity
20232025
most citedYOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection

3 citations · 4 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20253 cited

YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection

Xu Lin, Jinlong Peng, Zhenye Gan +2

Existing Real-Time Object Detection (RTOD) methods commonly adopt YOLO-like architectures for their favorable trade-off between accuracy and speed. However, these models rely on st…

cs.CV2025

Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding

Thong Nguyen, Zhiyuan Hu, Xu Lin +3

Recent years have witnessed outstanding advances of large vision-language models (LVLMs). In order to tackle video understanding, most of them depend upon their implicit temporal u…

cs.CV2024

Towards Artwork Explanation in Large-scale Vision Language Models

Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…

cs.CV2023

Early Action Recognition with Action Prototypes

Guglielmo Camporese, Alessandro Bergamo, Xunyu Lin +2

Early action recognition is an important and challenging problem that enables the recognition of an action from a partially observed video stream where the activity is potentially…

cs.CV2023

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Chaoyou Fu, Peixian Chen, Yunhang Shen +11

Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on…