papers
Publications (2)
cs.CV2025
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…
cs.CV2021
2nd Place Solution for Waymo Open Dataset Challenge -- Real-time 2D Object Detection
Yueming Zhang, Xiaolin Song, Bing Bai +8
In an autonomous driving system, it is essential to recognize vehicles, pedestrians and cyclists from images. Besides the high accuracy of the prediction, the requirement of real-t…