2 papers
cs.MM2025
Mano Technical Report
Tianyu Fu, Anyang Su, Chenxu Zhao +20
Graphical user interfaces (GUIs) are the primary medium for human-computer interaction, yet automating GUI interactions remains challenging due to the complexity of visual elements…
cs.CV2025
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
Xinxin Dong, Baoyun Peng, Haokai Ma +4
Video Question Answering (VideoQA) requires identifying sparse critical moments in long videos and reasoning about their causal relationships to answer semantically complex questio…