1 citations · 2 across the 6 of their papers we have counts for
7 papers
An AI-driven robotic system for two-dimensional hetero-assemblies
Xiaoxi Li, Jinkun He, Haojie Liu +27
Nanomaterials stacked on-demand, such as rotationally assembled two-dimensional (2D) van der Waals (vdW) layered compounds, provides a versatile platform for quantum simulation and…
FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation
Jiayong Zhu, Jiangtong Li, Jinru Ding +3
While Large Multimodal Models (LMMs) excel in general visual tasks, their deployment in specialized financial contexts remains insufficient. Existing benchmarks prioritize isolated…
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
Run Xu, Lu Li, Rongzhao Zhang +1
Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over explicit cultural context,…
SInViG: A Self-Evolving Interactive Visual Agent for Human-Robot Interaction
Jie Xu, Hanbo Zhang, Xinghang Li +3
Linguistic ambiguity is ubiquitous in our daily lives. Previous works adopted interaction between robots and humans for language disambiguation. Nevertheless, when interactive robo…
Towards Unified Interactive Visual Grounding in The Wild
Jie Xu, Hanbo Zhang, Qingyi Si +3
Interactive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate…
Vision-Language Foundation Models as Effective Robot Imitators
Xinghang Li, Minghuan Liu, Hanbo Zhang +9
Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipul…