activity
20242026
collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2026

PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation

Jianjian Yin, Tao Chen, Yi Chen +4

Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract ima…

cs.CV2025

Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging

Shenghe Zheng, Hongzhi Wang, Chenyu Huang +5

With more open-source models available for diverse tasks, model merging has gained attention by combining models into one, reducing training, storage, and inference costs. Current…

cs.CV2025

Dynamic Base model Shift for Delta Compression

Chenyu Huang, Peng Ye, Shenghe Zheng +4

Transformer-based models with the pretrain-finetune paradigm bring about significant progress, along with the heavy storage and deployment costs of finetuned models on multiple tas…

cs.CV2025

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

Chongjun Tu, Lin Zhang, Pengtao Chen +5

Multimodal Large Language Models (MLLMs) have shown remarkable capabilities in video content understanding but still struggle with fine-grained motion comprehension. To comprehensi…

cs.CV2025

TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Xudong Tan, Peng Ye, Chongjun Tu +5

Multimodal Large Language Models (MLLMs) are becoming increasingly popular, while the high computational cost associated with multimodal data input, particularly from visual tokens…

cs.CV2025

Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs

Chongjun Tu, Peng Ye, Dongzhan Zhou +4

Multi-Modal Large Language Models (MLLMs) stand out in various tasks but still struggle with hallucinations. While recent training-free mitigation methods mostly introduce addition…