works on

From the 1 of 11 linked papers with an AI index.

activity
20242026
collaborators

11 papers

cs.HC2026

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

Jiahao Huang, Zheng Lian, Jingyi Zhang +3

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific sp…

cs.CV2026

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

Kai Jiang, Jiaxing Huang, Jingyi Zhang +5

The paper introduces MobileSAM2, a lightweight version of the SAM2 segmentation model designed for mobile devices, using hypergraph-based knowledge distillation to transfer tempora…

cs.AI2026

OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models

Tianyi Lin, Chuanyu Sun, Jingyi Zhang +6

Equipping Large Language Model (LLM) agents with effective skills is crucial for solving complex tasks in real-world systems like OpenClaw. In this work, we aim to develop a framew…

cs.LG2026

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?

Jingyi Zhang, Tianyi Lin, Huanjin Yao +3

In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. T…

cs.AI2026

KLDrive: Fine-Grained 3D Scene Reasoning for Autonomous Driving based on Knowledge Graph

Ye Tian, Jingyi Zhang, Zihao Wang +4

Autonomous driving requires reliable reasoning over fine-grained 3D scene facts. Fine-grained question answering over multi-modal driving observations provides a natural way to eva…

cs.CV2026

MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline

Huanjin Yao, Qixiang Yin, Min Yang +5

We aim to develop a multimodal research agent capable of explicit reasoning and planning, multi-tool invocation, and cross-modal information synthesis, enabling it to conduct deep…