activity
20232026
most citedHow Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

16 citations · 59 across the 42 of their papers we have counts for

collaborators
Showing cs.CVShow all

59 papers · 1 filter

cs.CV2026

HPSD: Hybrid-Policy Self-Distillation for Text-Image-to-Video Diffusion Models

Jiazi Bu, Pengyang Ling, Yujie Zhou +10

Text-Image-to-Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text-to-video (T2V) and image-to-video (I2V) generation. Given…

cs.CV2026

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games

Shengyuan Ding, Xilin Wei, Xinyu Fang +4

Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observations that are no longer visible. However, existing benchmarks e…

cs.CV2026

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

Shuai Yang, Bingjie Gao, Ziwei Liu +3

Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time a…

cs.CV2026

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

Penghui Yang, Long Xing, Xiaoyi Dong +10

Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Curren…

cs.CV2026

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO

Jiazi Bu, Pengyang Ling, Yujie Zhou +8

Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. However, we have identified that t…

cs.CV2026

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Haiwen Diao, Penghao Wu, Hanming Deng +55

Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fra…