16 papers
RAVA: Retrieval-Augmented Viewpoint Alignment for Subject-Driven Image Generation
Qiwei Yan, Zhiqiang Yuan, Chongyang Li +4
Reference-driven image generation has made rapid progress on identity preservation, but reliable viewpoint control across different subjects remains poorly understood. The difficul…
PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums
Qiwei Yan, Zhiqiang Yuan, Zexi Jia +4
Longitudinal personal albums are weak-schema multimodal databases: noisy perceptual records whose key facts require joins across faces, text, timestamps, locations, and repeated ev…
Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues
Hanbo Bi, Zhiqiang Yuan, Chongyang Li +7
With the widespread adoption of multi-modal communication platforms, long-form dialogues interleaving text and images have become increasingly common. Users often need to retrieve…
PhotoCraft: Agentic Reasoning with Hierarchical Self-Evolving Memory for Deep Image Search
Kailin Lyu, Zhiqiang Yuan, Jianwei He +9
Deep Image Search requires multi-step reasoning over rich contextual cues, such as time, location, and event relations. However, most existing LLM-based agents are stateless and re…
CoDA: Color Distribution Probing for Efficient and Generalizable AI-Generated Image Detection
Zexi Jia, Zhiqiang Yuan, Xiaoyue Duan +3
AI-generated image detection faces a persistent trade-off between generalization and efficiency: lightweight artifact-based methods often degrade on unseen generators or domains, w…
Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
Chongyang Li, Zhiqiang Yuan, Hanbo Bi +2
Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs) to develop effective walking a…