collaborators

5 papers

cs.AI2026

ProAssist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks

Lilin Xu, Bufang Yang, Siyang Jiang +6

Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support…

cs.LG2025

TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models

Kaiyuan Hou, Minghui Zhao, Lilin Xu +2

Top-down images play an important role in safety-critical settings such as autonomous navigation and aerial surveillance, where they provide holistic spatial information that front…

cs.CV2025

Exploring the Capabilities of LLMs for IMU-based Fine-grained Human Activity Understanding

Lilin Xu, Kaiyuan Hou, Xiaofan Jiang

Human activity recognition (HAR) using inertial measurement units (IMUs) increasingly leverages large language models (LLMs), yet existing approaches focus on coarse activities lik…

cs.HC2025

Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation

Yunqi Guo, Kaiyuan Hou, Heming Fu +4

Understanding sensor data can be difficult for non-experts because of the complexity and different semantic meanings of sensor modalities. This leads to a need for intuitive and ef…

cs.RO2025

FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems

Minghui Zhao, Junxi Xia, Kaiyuan Hou +3

Foundation models (FM) have shown immense human-like capabilities for generating digital media. However, foundation models that can freely sense, interact, and actuate the physical…