12 citations · 46 across the 37 of their papers we have counts for
55 papers · 1 filter
HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models
Jiazi Bu, Pengyang Ling, Yujie Zhou +10
Text-Image-to-Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text-to-video (T2V) and image-to-video (I2V) generation. Given…
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games
Shengyuan Ding, Xilin Wei, Xinyu Fang +4
Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observations that are no longer visible. However, existing benchmarks e…
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
Jiazi Bu, Pengyang Ling, Yujie Zhou +8
Group Relative Policy Optimization (GRPO) has emerged as a powerful framework for preference alignment in text-to-image (T2I) flow models. However, we observe that the standard par…
Visual-ERM: Reward Modeling for Visual Equivalence
Ziyu Liu, Shengyuan Ding, Xinyu Fang +7
Vision-to-code tasks require models to reconstruct structured visual inputs, such as charts, tables, and SVGs, into executable or structured representations with high visual fideli…
Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart Parsing
Jinsong Li, Xiaoyi Dong, Yuhang Zang +3
While Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities for reasoning and self-correction at the textual level, these strengths provide minimal benefit…
Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
Yuhao Dong, Shulin Tian, Shuai Liu +6
Despite the growing video understanding capabilities of recent Multimodal Large Language Models (MLLMs), existing video benchmarks primarily assess understanding based on models' s…