4 papers · 1 filter
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
Yankai Yang, Yancheng Long, Bin Wen +4
Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos sha…
SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing
Yankai Yang, Yancheng Long, Wei Chen +7
Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a single whole-image reward, which…
AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation
Tongfei Chen, Shuo Yang, Yuguang Yang +7
Referring Image Segmentation (RIS) aims to segment the object in an image uniquely referred to by a natural language expression. However, RIS training often contains hard-to-align…
Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
Dian Xie, Shitong Shao, Lichen Bai +5
Classifier-free guidance (CFG) has helped diffusion models achieve great conditional generation in various fields. Recently, more diffusion guidance methods have emerged with impro…