Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
Yuke Zhu, Chi Xie, Shuang Liang +2
Recent advances on Multi-modal Large Language Models have demonstrated that high-resolution image input is crucial for model capabilities, especially for fine-grained tasks. Howeve…
cs.CV2023
Edge Wasserstein Distance Loss for Oriented Object Detection
Yuke Zhu, Yumeng Ruan, Zihua Xiong +1
Regression loss design is an essential topic for oriented object detection. Due to the periodicity of the angle and the ambiguity of width and height definition, traditional L1-dis…
cs.CV2023
RotaTR: Detection Transformer for Dense and Rotated Object
Zhu Yuke, Ruan Yumeng, Yang Lei +1
Detecting the objects in dense and rotated scenes is a challenging task. Recent works on this topic are mostly based on Faster RCNN or Retinanet. As they are highly dependent on th…