Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning
Jiapeng Li, Ping Wei, Wenjuan Han +2
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark mat…
cs.CV2026
Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
Zizhao Chen, Ping Wei, Ziyang Ren +2
As multimodal misinformation becomes more sophisticated, its detection and grounding are crucial. However, current multimodal verification methods, relying on passive holistic fusi…