2 citations · 2 across the 5 of their papers we have counts for
4 papers · 1 filter
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
Zehua Zhao, Zhixian Huang, Junren Li +28
Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and mis…
When Does Multimodality Lead to Better Time Series Forecasting?
Xiyuan Zhang, Boran Han, Haoyang Fang +11
Recently, there has been growing interest in incorporating textual information into foundation models for time series forecasting. However, it remains unclear whether and under wha…
Improving Generated and Retrieved Knowledge Combination Through Zero-shot Generation
Xinkai Du, Quanjie Han, Chao Lv +5
Open-domain Question Answering (QA) has garnered substantial interest by combining the advantages of faithfully retrieved passages and relevant passages generated through Large Lan…