3 papers
cs.CV2025
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
Zeyi Huang, Utkarsh Ojha, Yuyang Ji +2
When a human undertakes a test, their responses likely follow a pattern: if they answered an easy question incorrectly, they would likely answer a more difficult one…
cs.CV2025
Aligned Datasets Improve Detection of Latent Diffusion-Generated Images
Anirudh Sundara Rajan, Utkarsh Ojha, Jedidiah Schloesser +1
As latent diffusion models (LDMs) democratize image generation capabilities, there is a growing need to detect fake images. A good detector should focus on the generative models fi…
cs.CV2024
Yo'LLaVA: Your Personalized Language and Vision Assistant
Thao Nguyen, Haotian Liu, Yuheng Li +3
Large Multimodal Models (LMMs) have shown remarkable capabilities across a variety of tasks (e.g., image captioning, visual question answering). While broad, their knowledge remain…