3 papers
cs.CV2025
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
Yihao Wang, Raphael Memmesheimer, Sven Behnke
The availability of large language models and open-vocabulary object perception methods enables more flexibility for domestic service robots. The large variability of domestic task…
cs.CV2025
FMNV: A Dataset of Media-Published News Videos for Fake News Detection
Yihao Wang, Zhong Qian, Peifeng Li
News media, particularly video-based platforms, have become deeply embed-ded in daily life, concurrently amplifying the risks of misinformation dissem-ination. Consequently, multim…
cs.CV2024
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
Yihao Wang, Lizhi Chen, Zhong Qian +1
News media, especially video news media, have penetrated into every aspect of daily life, which also brings the risk of fake news. Therefore, multimodal fake news detection has rec…