activity
20162025
most citedTemporal Segment Networks: Towards Good Practices for Deep Action Recognition

289 citations · 529 across the 20 of their papers we have counts for

collaborators
Showing cs.CVShow all

19 papers · 1 filter

cs.CV2025

Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM

Xinyu Fang, Zhijian Chen, Kai Lan +10

Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts. While Large Language Models (LLMs) ha…

cs.CV202416 cited

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Zhe Chen, Weiyun Wang, Hao Tian +32

In this report, we introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models…

cs.CV20235 cited

RenderMe-360: A Large Digital Asset Library and Benchmarks Towards High-fidelity Head Avatars

Dongwei Pan, Long Zhuo, Jingtan Piao +13

Synthesizing high-fidelity head avatars is a central problem for computer vision and graphics. While head avatar synthesis algorithms have advanced rapidly, the best ones still fac…

cs.CV20234 cited

DORT: Modeling Dynamic Objects in Recurrent for Multi-Camera 3D Object Detection and Tracking

Qing Lian, Tai Wang, Dahua Lin +1

Recent multi-camera 3D object detectors usually leverage temporal information to construct multi-view stereo that alleviates the ill-posed depth estimation. However, they typically…

cs.CV2023

RIFormer: Keep Your Vision Backbone Effective While Removing Token Mixer

Jiahao Wang, Songyang Zhang, Yong Liu +6

This paper studies how to keep a vision backbone effective while removing token mixers in its basic building blocks. Token mixers, as self-attention for vision transformers (ViTs),…

cs.CV2023

Grid-guided Neural Radiance Fields for Large Urban Scenes

Linning Xu, Yuanbo Xiangli, Sida Peng +5

Purely MLP-based neural radiance fields (NeRF-based methods) often suffer from underfitting with blurred renderings on large-scale scenes due to limited model capacity. Recent appr…