4 papers
When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance
Jinjin Cao, Zhiyang Chen, Zijun Wang +3
Vision-Language Models (VLMs) have shown solid ability for multimodal understanding of both visual and language contexts. However, existing VLMs often face severe challenges of hal…
C-Evolve: Consensus-based Evolution for Prompt Groups
Tiancheng Li, Yuhang Wang, Zhiyang Chen +3
Prompt evolution algorithms offer a powerful paradigm for enhancing AI systems based on closed-source models, while few work explores whether aggregating results from multiple prom…
InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO
Xueji Fang, Liyuan Ma, Zhiyang Chen +2
Recent advances in text-to-video generation, particularly with autoregressive models, have enabled the synthesis of high-quality videos depicting individual scenes. However, extend…
Self-Guidance: Boosting Flow and Diffusion Generation on Their Own
Tiancheng Li, Weijian Luo, Zhiyang Chen +2
Proper guidance strategies are essential to achieve high-quality generation results without retraining diffusion and flow-based text-to-image models. Existing guidance either requi…