42 citations · 66 across the 8 of their papers we have counts for
8 papers
Temporal and Contextual Transformer for Multi-Camera Editing of TV Shows
Anyi Rao, Xuekun Jiang, Sichen Wang +7
The ability to choose an appropriate camera view among multiple cameras plays a vital role in TV shows delivery. But it is hard to figure out the statistical pattern and apply inte…
A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language
Bing Su, Dazhao Du, Zhao Yang +6
Although artificial intelligence (AI) has made significant progress in understanding molecules in a wide range of fields, existing models generally acquire the single cognitive abi…
AutoGPart: Intermediate Supervision Search for Generalizable 3D Part Segmentation
Xueyi Liu, Xiaomeng Xu, Anyi Rao +2
Training a generalizable 3D part segmentation network is quite challenging but of great importance in real-world applications. To tackle this problem, some works design task-specif…
A Unified Framework for Shot Type Classification Based on Subject Centric Lens
Anyi Rao, Jiaze Wang, Linning Xu +4
Shots are key narrative elements of various videos, e.g. movies, TV series, and user-generated videos that are thriving over the Internet. The types of shots greatly influence how…
Online Multi-modal Person Search in Videos
Jiangyue Xia, Anyi Rao, Qingqiu Huang +3
The task of searching certain people in videos has seen increasing potential in real-world applications, such as video organization and editing. Most existing approaches are devise…
MovieNet: A Holistic Dataset for Movie Understanding
Qingqiu Huang, Yu Xiong, Anyi Rao +2
Recent years have seen remarkable advances in visual understanding. However, how to understand a story-based long video with artistic styles, e.g. movie, remains challenging. In th…