3 papers
cs.CL2025
ScreenLLM: Stateful Screen Schema for Efficient Action Understanding and Prediction
Yiqiao Jin, Stefano Petrangeli, Yu Shen +1
Graphical User Interface (GUI) agents are autonomous systems that interpret and generate actions, enabling intelligent user assistance and automation. Effective training of these a…
cs.CV2024
Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention
Uttaran Bhattacharya, Gang Wu, Stefano Petrangeli +2
We propose a method to detect individualized highlights for users on given target videos based on their preferred highlight clips marked on previous videos they have watched. Our m…
cs.CV2024
HighlightMe: Detecting Highlights from Human-Centric Videos
Uttaran Bhattacharya, Gang Wu, Stefano Petrangeli +2
We present a domain- and user-preference-agnostic approach to detect highlightable excerpts from human-centric videos. Our method works on the graph-based representation of multipl…