4 papers · 1 filter
REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation
Tengjie Lin, Yutao Sun, Jingwei Ni +7
Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than genuine visual ability, thus o…
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
Yanting Miao, Yutao Sun, Dexin Wang +8
Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools or image generators. However…
Image-POSER: Reflective RL for Multi-Expert Image Generation and Editing
Hossein Mohebbi, Mohammed Abdulrahman, Yanting Miao +2
Recent advances in text-to-image generation have produced strong single-shot models, yet no individual system reliably executes the long, compositional prompts typical of creative…
Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning
Yanting Miao, William Loh, Suraj Kothawade +3
Text-to-image generative models have recently attracted considerable interest, enabling the synthesis of high-quality images from textual prompts. However, these models often lack…