activity
20232026
most citedYOLOv10: Real-Time End-to-End Object Detection

1.1k citations · 1.1k across the 38 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

12 papers · 2 filters

cs.CV2024

YOLO-UniOW: Efficient Universal Open-World Object Detection

Lihao Liu, Juexiao Feng, Hui Chen +4

Traditional object detection models are constrained by the limitations of closed-set datasets, detecting only categories encountered during training. While multimodal models have e…

cs.CV2024

Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning

Hui-Yue Yang, Hui Chen, Ao Wang +7

Segment Anything Model (SAM) has made great progress in anomaly segmentation tasks due to its impressive generalization ability. However, existing methods that directly apply SAM t…

cs.CV2024

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Ao Wang, Fengyuan Sun, Hui Chen +3

Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across a wide range of vision-language tasks, garnering significant attention in the computer…

cs.CV2024

PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation

Ao Wang, Hui Chen, Jiaxin Li +6

Recently, large vision-language models (LVLMs) have rapidly gained popularity for their strong generation and reasoning capabilities given diverse multimodal inputs. However, these…

cs.CV2024

Context Enhancement with Reconstruction as Sequence for Unified Unsupervised Anomaly Detection

Hui-Yue Yang, Hui Chen, Lihao Liu +5

Unsupervised anomaly detection (AD) aims to train robust detection models using only normal samples, while can generalize well to unseen anomalies. Recent research focuses on a uni…

cs.CV2024

TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Leqi Shen, Tianxiang Hao, Tao He +5

Most text-video retrieval methods utilize the text-image pre-trained models like CLIP as a backbone. These methods process each sampled frame independently by the image encoder, re…