5 papers
Rethinking Text-Based Image Retrieval in Specific Domain
Jingyang Tan, Sheng Yang, Yuanpeng Chen +4
Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, existing benchmarks are predominan…
EdgeFM: Efficient Edge Inference for Vision-Language Models
Mengling Deng, Yuanpeng Chen, Sheng Yang +12
Vision-language models (VLMs) have demonstrated strong applicability in edge industrial applications, yet their deployment remains severely constrained by requirements for determin…
EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation
Yue Ma, Xu Ye, Qinghe Wang +9
Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of spatial textures and temporal…
Fast-BEV++: Fast by Algorithm, Deployable by Design
Yuanpeng Chen, Hui Song, Sheng Yang +5
The advancement of vision-only Bird's-Eye-View (BEV) perception, a core paradigm for cost-effective autonomous driving, is hindered by the long-standing fundamental trade-off betwe…
Precise Drive with VLM: First Prize Solution for PRCV 2024 Drive LM challenge
Bin Huang, Siyu Wang, Yuanpeng Chen +8
This technical report outlines the methodologies we applied for the PRCV Challenge, focusing on cognition and decision-making in driving scenarios. We employed InternVL-2.0, a pion…