activity
20142024
most citedDeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection

13 citations · 17 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

DAM: Dynamic Adapter Merging for Continual Video QA Learning

Feng Cheng, Ziyang Wang, Yi-Lin Sung +3

We present a parameter-efficient method for continual video question-answering (VidQA) learning. Our method, named DAM, uses the proposed Dynamic Adapter Merging to (i) mitigate ca…

cs.CV2024

Siamese Vision Transformers are Scalable Audio-visual Learners

Yan-Bo Lin, Gedas Bertasius

Traditional audio-visual methods rely on independent audio and visual backbones, which is costly and not scalable. In this work, we investigate using an audio-visual siamese networ…

cs.CV20241 cited

Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Xiyao Wang, Yuhang Zhou, Xiaoyu Liu +9

Multimodal Large Language Models (MLLMs) have demonstrated proficiency in handling a variety of visual-language tasks. However, current MLLM benchmarks are predominantly designed t…

cs.CV20232 cited

Unified Coarse-to-Fine Alignment for Video-Text Retrieval

Ziyang Wang, Yi-Lin Sung, Feng Cheng +2

The canonical approach to video-text retrieval leverages a coarse-grained or fine-grained alignment between visual and textual information. However, retrieving the correct video ac…

cs.CV201413 cited

DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection

Gedas Bertasius, Jianbo Shi, Lorenzo Torresani

Contour detection has been a fundamental component in many image segmentation and object detection systems. Most previous work utilizes low-level features such as texture or salien…