1 citations · 1 across the 6 of their papers we have counts for
6 papers
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
Yunqiao Yang, Wenbo Li, Houxing Ren +6
The rapid evolution of Large Language Models (LLMs) has fostered diverse paradigms for automated slide generation, ranging from code-driven layouts to image-centric synthesis. Howe…
MOOM: Maintenance, Organization and Optimization of Memory in Ultra-Long Role-Playing Dialogues
Weishu Chen, Jinyi Tang, Zhouhui Hou +7
Memory extraction is crucial for maintaining coherent ultra-long dialogues in human-robot role-playing scenarios. However, existing methods often exhibit uncontrolled memory growth…
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Ziming Cheng, Binrui Xu, Lisheng Gong +14
With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously…
SpiritSight Agent: Advanced GUI Agent with One Look
Zhiyuan Huang, Ziming Cheng, Junting Pan +2
Graphical User Interface (GUI) agents show amazing abilities in assisting human-computer interaction, automating human user's navigation on digital devices. An ideal GUI agent is e…
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
Ziming Cheng, Zhiyuan Huang, Junting Pan +2
Graphical user interfaces (GUI) automation agents are emerging as powerful tools, enabling humans to accomplish increasingly complex tasks on smart devices. However, users often in…
TeacherLM: Teaching to Fish Rather Than Giving the Fish, Language Modeling Likewise
Nan He, Hanyu Lai, Chenyang Zhao +12
Large Language Models (LLMs) exhibit impressive reasoning and data augmentation capabilities in various NLP tasks. However, what about small models? In this work, we propose Teache…