collaborators

10 papers

cs.CL2026

Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering

Jieyuan Liu, Jianyang Gu, Shijie Chen +2

Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by conditioning a reader on passages retriev…

cs.CV2026

Leveraging Latent Visual Reasoning in Silence

Dongyao Zhu, Zhen Wang, Xi Xiao +7

Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of th…

cs.CV2026

TaxaAdapter: Vision Taxonomy Models are Key to Fine-grained Image Generation over the Tree of Life

Mridul Khurana, Amin Karimi Monsefi, Justin Lee +9

Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the rema…

cs.CV2026

Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time

Sooyoung Jeon, Hongjie Tian, Lemeng Wang +7

Camera traps are vital for large-scale biodiversity monitoring, yet accurate automated analysis remains challenging due to diverse deployment environments. While the computer visio…

cs.CV2026

BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models

Ziheng Zhang, Xinyue Ma, Arpita Chowdhury +9

This work investigates descriptive captions as an additional source of supervision for biological multimodal foundation models. Images and captions can be viewed as complementary s…

cs.CV2025

Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation

Connor Kilrain, David Carlyn, Julia Chae +3

The rise of personalized generative models raises a central question: how should we evaluate identity preservation? Given a reference image (e.g., one's pet), we expect the generat…