4 papers
VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding
Junbo Zou, Ziheng Huang, Shengjie Zhang +2
Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints and the need to capture informatio…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
LightAgent: Production-level Open-source Agentic AI Framework
Weige Cai, Tong Zhu, Jinyi Niu +6
With the rapid advancement of large language models (LLMs), Multi-agent Systems (MAS) have achieved significant progress in various application scenarios. However, substantial chal…
FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios
Yingqian Wu, Qiushi Wang, Zefei Long +7
Financial report generation tasks range from macro- to micro-economics analysis, also requiring extensive data analysis. Existing LLM models are usually fine-tuned on simple QA tas…