most citedEmu3: Next-Token Prediction is All You Need

4 citations · 4 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CL2025

Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT

Yesheng Liu, Hao Li, Haiyu Xu +9

Multiple-choice question answering (MCQA) has been a popular format for evaluating and reinforcement fine-tuning (RFT) of modern multimodal language models. Its constrained output…

cs.CL2025

FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions

Bowen Qin, Chen Yue, Fang Yin +26

We conduct a moderate-scale contamination-free (to some extent) evaluation of current large reasoning models (LRMs) with some preliminary findings. We also release ROME, our evalua…

cs.CV2025

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

Zheqi He, Yesheng Liu, Jing-shu Zheng +5

We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generati…

cs.RO2025

RoboBrain 2.0 Technical Report

BAAI RoboBrain Team, Mingyu Cao, Huajie Tan +50

We introduce RoboBrain 2.0, our latest generation of embodied vision-language foundation models, designed to unify perception, reasoning, and planning for complex embodied tasks in…

cs.CV2025

Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

Xuannan Liu, Zekun Li, Zheqi He +6

The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily…

cs.CV2025

T2VEval: Benchmark Dataset and Objective Evaluation Method for T2V-generated Videos

Zelu Qi, Ping Shi, Shuqi Wang +7

Recent advances in text-to-video (T2V) technology, as demonstrated by models such as Runway Gen-3, Pika, Sora, and Kling, have significantly broadened the applicability and popular…