benchmark 1evaluation metrics 1evolving contexts 1keyframe conditioning 1large language models 1multimodal models 1on-policy distillation 1open-ended tasks 1reverse kl 1video generation 1
From the 2 of 18 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
Yang Shi, Yuhao Dong, Yue Ding +22
The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question rem…
cs.AI2025
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
Xukai Wang, Xuanbo Liu, Mingrui Chen +16
With the advancement of powerful large-scale reasoning models, effectively evaluating the reasoning capabilities of these models has become increasingly important. However, existin…