activity
20232026
most citedSource-Free Object Detection with Detection Transformer

6 citations · 14 across the 30 of their papers we have counts for

collaborators
Showing cs.CVShow all

16 papers · 1 filter

cs.CV2026

CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension

Lihao Liu, Yan Wang, Biao Yang +10

Multimodal Large Language Models (MLLMs) have shown remarkable success in comprehension tasks such as visual description and visual question answering. However, their direct applic…

cs.CV2025

Tracking and Segmenting Anything in Any Modality

Tianlu Zhang, Qiang Zhang, Guiguang Ding +1

Tracking and segmentation play essential roles in video understanding, providing basic positional information and temporal association of objects within video sequences. Despite th…

cs.CV2025

PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning

Fengyuan Sun, Hui Chen, Xinhao Xu +5

While multi-modal large language models (MLLMs) have made significant progress in recent years, the issue of hallucinations remains a major challenge. To mitigate this phenomenon,…

cs.CV20256 cited

Source-Free Object Detection with Detection Transformer

Huizai Yao, Sicheng Zhao, Shuo Lu +7

Source-Free Object Detection (SFOD) enables knowledge transfer from a source domain to an unsupervised target domain for object detection without access to source data. Most existi…

cs.CV2025

Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations

Yiwen Liang, Hui Chen, Yizhe Xiong +7

Vision-language models (VLMs) exhibit remarkable zero-shot capabilities but struggle with distribution shifts in downstream tasks when labeled data is unavailable, which has motiva…

cs.CV2025

Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation

Yizhe Xiong, Zihan Zhou, Yiwen Liang +6

Test-Time Adaptation (TTA) has emerged as an effective solution for adapting Vision Transformers (ViT) to distribution shifts without additional training data. However, existing TT…