5 papers
ProAssist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks
Lilin Xu, Bufang Yang, Siyang Jiang +6
Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support…
TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models
Kaiyuan Hou, Minghui Zhao, Lilin Xu +2
Top-down images play an important role in safety-critical settings such as autonomous navigation and aerial surveillance, where they provide holistic spatial information that front…
Exploring the Capabilities of LLMs for IMU-based Fine-grained Human Activity Understanding
Lilin Xu, Kaiyuan Hou, Xiaofan Jiang
Human activity recognition (HAR) using inertial measurement units (IMUs) increasingly leverages large language models (LLMs), yet existing approaches focus on coarse activities lik…
Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation
Yunqi Guo, Kaiyuan Hou, Heming Fu +4
Understanding sensor data can be difficult for non-experts because of the complexity and different semantic meanings of sensor modalities. This leads to a need for intuitive and ef…
FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems
Minghui Zhao, Junxi Xia, Kaiyuan Hou +3
Foundation models (FM) have shown immense human-like capabilities for generating digital media. However, foundation models that can freely sense, interact, and actuate the physical…