9 papers
Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
Ziheng Qin, Yuheng Ji, Renshuai Tao +4
The pursuit of a universal AI-generated image (AIGI) detector often relies on aggregating data from numerous generators to improve generalization. However, this paper identifies a…
PReD: An LLM-based Foundation Multimodal Model for Electromagnetic Perception, Recognition, and Decision
Zehua Han, Jing Xiao, Yiqi Duan +13
Multimodal Large Language Models have demonstrated powerful cross-modal understanding and reasoning capabilities in general domains. However, in the electromagnetic (EM) domain, th…
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
Shuanghao Bai, Wenxuan Song, Jiayi Chen +14
Recent advances in vision, language, and multimodal learning have significantly accelerated progress in robotic foundation models, with robotic manipulation remaining one of the mo…
Towards Cross-View Point Correspondence in Vision-Language Models
Yipu Wang, Yuheng Ji, Yuyang Liu +10
Cross-view correspondence is a fundamental capability for spatial understanding and embodied AI. However, it is still far from being realized in Vision-Language Models (VLMs), espe…
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
Shuanghao Bai, Wenxuan Song, Jiayi Chen +15
Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer vision, natural language processing, and the rise of large-scale multimodal…
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
Yuheng Ji, Huajie Tan, Cheng Chi +8
We introduce \textsc{MathSticks}, a benchmark for Visual Symbolic Compositional Reasoning (VSCR), which unifies visual perception, symbolic manipulation, and arithmetic consistency…