Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou +6
AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimoda…
cs.CV2025
Open Multimodal Retrieval-Augmented Factual Image Generation
Yang Tian, Fan Liu, Jingyuan Zhang +3
Large Multimodal Models (LMMs) have achieved remarkable progress in generating photorealistic and prompt-aligned images, but they often produce outputs that contradict verifiable k…