4 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.CL2026★ 2 cited
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
cs.CV2021★ 4 cited
Multimodal Feature Fusion for Video Advertisements Tagging Via Stacking Ensemble
Qingsong Zhou, Hai Liang, Zhimin Lin +1
Automated tagging of video advertisements has been a critical yet challenging problem, and it has drawn increasing interests in last years as its applications seem to be evident in…