3 papers
cs.AI2025
Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
Tingyu Li, Zheng Sun, Jingxuan Wei +4
Recent vision-language models (VLMs) achieve remarkable reasoning through reinforcement learning (RL), which provides a feasible solution for realizing continuous self-evolving lar…
cs.CV2025
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
Juanxi Tian, Siyuan Li, Conghui He +2
Current multimodal models aim to transcend the limitations of single-modality representations by unifying understanding and generation, often using text-to-image (T2I) tasks to cal…
cs.AI2025
GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
Jingxuan Wei, Caijun Jia, Xi Bai +7
The advent of Unified Multimodal Models (UMMs) signals a paradigm shift in artificial intelligence, moving from passive perception to active, cross-modal generation. Despite their…