59 citations · 130 across the 28 of their papers we have counts for
11 papers · 2 filters
StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video Sequences
Shangkun Sun, Jiaming Liu, Thomas H. Li +3
Occlusions between consecutive frames have long posed a significant challenge in optical flow estimation. The inherent ambiguity introduced by occlusions directly violates the brig…
Mug-STAN: Adapting Image-Language Pretrained Models for General Video Understanding
Ruyang Liu, Jingjia Huang, Wei Gao +2
Large-scale image-language pretrained models, e.g., CLIP, have demonstrated remarkable proficiency in acquiring general multi-modal knowledge through web-scale image-text data. Des…
Efficient Test-Time Adaptation for Super-Resolution with Second-Order Degradation and Reconstruction
Zeshuai Deng, Zhuokun Chen, Shuaicheng Niu +3
Image super-resolution (SR) aims to learn a mapping from low-resolution (LR) to high-resolution (HR) using paired HR-LR training images. Conventional SR methods typically gather th…
FGPrompt: Fine-grained Goal Prompting for Image-goal Navigation
Xinyu Sun, Peihao Chen, Jugang Fan +3
Learning to navigate to an image-specified goal is an important but challenging task for autonomous systems. The agent is required to reason the goal location from where a picture…
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
Ruyang Liu, Chen Li, Yixiao Ge +3
The recent progress in Large Language Models (LLM) has spurred various advancements in image-language conversation agents, while how to build a proficient video-based dialogue syst…
Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models
Peihao Chen, Xinyu Sun, Hongyan Zhi +5
We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by langu…