activity
20192025
most citedDiffuVolume: Diffusion Model for Volume based Stereo Matching

6 citations · 10 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2025

ProEdit: Inversion-based Editing From Prompts Done Right

Zhi Ouyang, Dian Zheng, Xiao-Ming Wu +4

Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image in…

cs.CV2025

Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation

Badi Li, Ren-jie Lu, Yu Zhou +2

The Object Goal Navigation (ObjectNav) task challenges agents to locate a specified object in an unseen environment by imagining unobserved regions of the scene. Prior approaches r…

cs.CV2025

ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation

Ling-An Zeng, Guohong Huang, Yi-Lin Wei +4

We propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unli…

cs.CV20252 cited

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Shenghao Fu, Qize Yang, Qijie Mo +5

Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a…

cs.CV20241 cited

Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization

Jia-Run Du, Kun-Yu Lin, Jingke Meng +1

To address the zero-shot temporal action localization (ZSTAL) task, existing works develop models that are generalizable to detect and classify actions from unseen categories. They…

cs.CV2024

Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation

Huilin Tian, Jingke Meng, Wei-Shi Zheng +3

Vision and Language Navigation (VLN) is a challenging task that requires agents to understand instructions and navigate to the destination in a visual environment.One of the key ch…