From the 1 of 11 linked papers with an AI index.
11 papers
OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
Jiahao Huang, Zheng Lian, Jingyi Zhang +3
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific sp…
MobileSAM2: Lightweight Segment Anything for Spatial Intelligence
Kai Jiang, Jiaxing Huang, Jingyi Zhang +5
The paper introduces MobileSAM2, a lightweight version of the SAM2 segmentation model designed for mobile devices, using hypergraph-based knowledge distillation to transfer tempora…
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models
Tianyi Lin, Chuanyu Sun, Jingyi Zhang +6
Equipping Large Language Model (LLM) agents with effective skills is crucial for solving complex tasks in real-world systems like OpenClaw. In this work, we aim to develop a framew…
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
Jingyi Zhang, Tianyi Lin, Huanjin Yao +3
In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. T…
KLDrive: Fine-Grained 3D Scene Reasoning for Autonomous Driving based on Knowledge Graph
Ye Tian, Jingyi Zhang, Zihao Wang +4
Autonomous driving requires reliable reasoning over fine-grained 3D scene facts. Fine-grained question answering over multi-modal driving observations provides a natural way to eva…
MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline
Huanjin Yao, Qixiang Yin, Min Yang +5
We aim to develop a multimodal research agent capable of explicit reasoning and planning, multi-tool invocation, and cross-modal information synthesis, enabling it to conduct deep…