16 citations · 59 across the 42 of their papers we have counts for
59 papers · 1 filter
HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models
Jiazi Bu, Pengyang Ling, Yujie Zhou +10
Text-Image-to-Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text-to-video (T2V) and image-to-video (I2V) generation. Given…
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games
Shengyuan Ding, Xilin Wei, Xinyu Fang +4
Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observations that are no longer visible. However, existing benchmarks e…
PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory
Shuai Yang, Bingjie Gao, Ziwei Liu +3
Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time a…
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning
Penghui Yang, Long Xing, Xiaoyi Dong +10
Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Curren…
AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO
Jiazi Bu, Pengyang Ling, Yujie Zhou +8
Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. However, we have identified that t…
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Haiwen Diao, Penghao Wu, Hanming Deng +55
Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fra…