42 citations · 88 across the 34 of their papers we have counts for
9 papers · 1 filter
MLLM-Guided Semantic Correction for Text-to-Video Generation
Junhao Chen, Zheqi Lv, Keting Yin +6
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic err…
NaviCache: Test-Time Self-Calibration Caching for Video Generation
Zheqi Lv, Zhibo Zhu, Jinke Wang +6
Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from calibration data dependency, prohibitive cali…
Training-Free Semantic Correction for Autoregressive Visual Models
Junhao Chen, Chanyu Zhu, Zheqi Lv +2
Autoregressive visual models (AVMs) based on next-scale prediction have emerged as a prominent paradigm for image and video synthesis. However, decomposing the generation process i…
DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery
Jiawei Wang, Ming Lei, Yaning Yang +8
Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a fundamental challenge in biod…
UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
Keming Ye, Zhipeng Huang, Canmiao Fu +7
With the rapid advances of powerful multimodal models such as GPT-4o, Nano Banana, and Seedream 4.0 in Image Editing, the performance gap between closed-source and open-source mode…
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
Zheqi Lv, Junhao Chen, Qi Tian +3
Diffusion models have become the mainstream architecture for text-to-image generation, achieving remarkable progress in visual quality and prompt controllability. However, current…