5 papers
Any2Poster: Any-Source Poster Generation Across Modalities and Domains
Amogh Vinaykumar, Aiden Li, Suozhi Huang +1
Visual posters are a compact medium for communicating dense information, yet progress on automatic poster generation remains difficult to measure because existing evaluations are o…
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
Aiden Yiliu Li, Nels Numan, Anthony Steed
Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps it searchable over long horiz…
Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding
Wee Joe Tan, Zi Rui Lucas Lim, Shashank Durgad +2
Evaluating web usability typically requires time-consuming user studies and expert reviews, which often limits iteration speed during product development, especially for small team…
Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts
Aiden Yiliu Li, Xinyue Hao, Shilong Liu +1
Despite advances in multimodal large language models, autonomous web agents still struggle to reliably execute long-horizon tasks on complex and dynamic web interfaces. Existing ag…
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
Aiden Yiliu Li, Bizhi Yu, Daoan Lei +2
GUI grounding aims to align natural language instructions with precise regions in complex user interfaces. Advanced multimodal large language models show strong ability in visual G…