5 papers
Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic
Chuou Xu, Liya Ji, Qifeng Chen
Reinforcement learning (RL) as post-training is crucial for enhancing the reasoning ability of large language models (LLMs) in coding and math. However, their capacity for visual s…
Instruction-based Image Editing with Planning, Reasoning, and Generation
Liya Ji, Chenyang Qi, Qifeng Chen
Editing images via instruction provides a natural way to generate interactive content, but it is a big challenge due to the higher requirement of scene understanding and generation…
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
Yifan Wang, Liya Ji, Zhanghan Ke +3
We propose an approach to enhancing synthetic video realism, which can re-render synthetic videos from a simulator in photorealistic fashion. Our realism enhancement approach is a…
VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module
Kam Man Wu, Zeyue Tian, Liya Ji +1
Video and audio inpainting for mixed audio-visual content has become a crucial task in multimedia editing recently. However, precisely removing an object and its corresponding audi…
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
Zhefan Rao, Liya Ji, Yazhou Xing +6
Text-to-video (T2V) generation has gained significant attention recently. However, the costs of training a T2V model from scratch remain persistently high, and there is considerabl…