6 papers · 1 filter
Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots
Shiquan Zhang, Tianyi Zhang, Le Fang +3
With the rapid advancement of large language models (LLMs), mobile agents have emerged as promising tools for phone automation, simulating human interactions on screens to accompli…
Real-Time Detection of Robot Failures Using Gaze Dynamics in Collaborative Tasks
Ramtin Tabatabaei, Vassilis Kostakos, Wafa Johal
Detecting robot failures during collaborative tasks is crucial for maintaining trust in human-robot interactions. This study investigates user gaze behaviour as an indicator of rob…
AutoJournaling: A Context-Aware Journaling System Leveraging MLLMs on Smartphone Screenshots
Tianyi Zhang, Shiquan Zhang, Le Fang +3
Journaling offers significant benefits, including fostering self-reflection, enhancing writing skills, and aiding in mood monitoring. However, many people abandon the practice beca…
ScreenTK: Seamless Detection of Time-Killing Moments Using Continuous Mobile Screen Text and On-Device LLMs
Le Fang, Shiquan Zhang, Hong Jia +2
Smartphones have become essential to people's digital lives, providing a continuous stream of information and connectivity. However, this constant flow can lead to moments where us…
Predicting Affective States from Screen Text Sentiment
Songyan Teng, Tianyi Zhang, Simon D'Alfonso +1
The proliferation of mobile sensing technologies has enabled the study of various physiological and behavioural phenomena through unobtrusive data collection from smartphone sensor…
Enabling On-Device LLMs Personalization with Smartphone Sensing
Shiquan Zhang, Ying Ma, Le Fang +3
This demo presents a novel end-to-end framework that combines on-device large language models (LLMs) with smartphone sensing technologies to achieve context-aware and personalized…