5 papers
SVAC: Scaling Is All You Need For Referring Video Object Segmentation
Li Zhang, Haoxiang Gao, Zhihao Zhang +2
Referring Video Object Segmentation (RVOS) aims to segment target objects in video sequences based on natural language descriptions. While recent advances in Multi-modal Large Lang…
Multimodal Foundation Model-Driven User Interest Modeling and Behavior Analysis on Short Video Platforms
Yushang Zhao, Yike Peng, Li Zhang +3
With the rapid expansion of user bases on short video platforms, personalized recommendation systems are playing an increasingly critical role in enhancing user experience and opti…
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
Haoxiang Gao, Li Zhang, Yu Zhao +2
Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand com…
OmniCam: Unified Multimodal Video Generation via Camera Control
Xiaoda Yang, Jiayang Xu, Kaixuan Luan +9
Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as co…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…