activity
20212026
most citedGeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation

6 citations · 9 across the 12 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

5 papers · 2 filters

cs.CV2024★ 1 cited

Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Ting Liu, Liangtao Shi, Richang Hong +3

The vision tokens in multimodal large language models usually exhibit significant spatial and temporal redundancy and take up most of the input tokens, which harms their inference…

cs.CV2024

MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension

Ting Liu, Zunnan Xu, Yue Hu +3

Referring Expression Comprehension (REC), which aims to ground a local visual region via natural language, is a task that heavily relies on multimodal alignment. Most existing meth…

cs.CV2024

M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension

Xuyang Liu, Ting Liu, Siteng Huang +6

Referring expression comprehension (REC) is a vision-language task to locate a target object in an image based on a language expression. Fully fine-tuning general-purpose pre-train…

cs.CV2024

DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding

Ting Liu, Xuyang Liu, Siteng Huang +5

Visual grounding (VG) is a challenging task to localize an object in an image based on a textual description. Recent surge in the scale of VG models has substantially improved perf…

cs.CV2024

Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference

Ting Liu, Xuyang Liu, Liangtao Shi +6

Parameter-efficient fine-tuning (PEFT) has emerged as a popular solution for adapting pre-trained Vision Transformer (ViT) models to downstream applications by updating only a smal…