1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.CV2026★ 1 cited
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
Zhiyuan Feng, Zhaolu Kang, Qijie Wang +16
Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent V…
cs.CL2025
SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media
Xilai Xu, Zilin Zhao, Chengye Song +4
With the increasing prevalence of multimodal content on social media, sentiment analysis faces significant challenges in effectively processing heterogeneous data and recognizing m…