activity
20192026
most citedVideo Moment Retrieval from Text Queries via Single Frame Annotation

41 citations · 230 across the 96 of their papers we have counts for

collaborators
Showing 2023 · cs.CVShow all

6 papers · 2 filters

cs.CV2023

Open-Vocabulary Video Relation Extraction

Wentao Tian, Zheng Wang, Yuqian Fu +2

A comprehensive understanding of videos is inseparable from describing the action with its contextual action-object interactions. However, many current video understanding tasks pr…

cs.CV2023★ 1 cited

Instance-aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting Learning

Yang Jiao, Zequn Jie, Shaoxiang Chen +4

Camera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction…

cs.CV2023★ 6 cited

FoodLMM: A Versatile Food Assistant using Large Multi-modal Model

Yuehao Yin, Huiyan Qi, Bin Zhu +3

Large Multi-modal Models (LMMs) have made impressive progress in many vision-language tasks. Nevertheless, the performance of general LMMs in specific domains is still far from sat…

cs.CV2023★ 1 cited

ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks

Zejun Li, Ye Wang, Mengfei Du +9

Recent years have witnessed remarkable progress in the development of large vision-language models (LVLMs). Benefiting from the strong language backbones and efficient cross-modal…

cs.CV2023

On the Importance of Spatial Relations for Few-shot Action Recognition

Yilun Zhang, Yuqian Fu, Xingjun Ma +4

Deep learning has achieved great success in video recognition, yet still struggles to recognize novel actions when faced with only a few examples. To tackle this challenge, few-sho…

cs.CV2023★ 4 cited

NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

Tianwen Qian, Jingjing Chen, Linhai Zhuo +2

We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to…