2 papers
cs.CV2026
Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models
Yujun Tong, Dongliang Chang, Zijin Yin +3
The long-standing goal of multimodal AI is to build unified models in which visual understanding and visual generation mutually enhance one another. Despite recent works such as BA…
cs.CV2026
Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding
Junhan Chen, Zilu Zhou, Yujun Tong +3
Fine-grained visual understanding is shifting from static classification to knowledge-augmented reasoning, where models must justify as well as recognise. Existing approaches remai…