9 papers
MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
Katsuya Ogata, Zongshang Pang, Mayu Otani +1
Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a…
Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems
Dianxing Shi, Bowen Wang, Junqi He +2
Self-evolving agents improve through continual self-play and self-generated learning signals, but autonomous evolution can also cause capability degradation and safety drift. Altho…
PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
Jiahao Zhang, Bowen Wang, Hong Liu +2
Visual In-Context Learning (VICL) uses input-output image pairs, referred to as in-context pairs (or examples), as prompts alongside query images to guide models in performing dive…
Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown
Bowen Wang, Zhouqiang Jiang, Yasuaki Susumu +3
The real value of knowledge lies not just in its accumulation, but in its potential to be harnessed effectively to conquer the unknown. Although recent multimodal large language mo…
DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models
Bowen Wang, Jiuyang Chang, Yiming Qian +6
Large language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GP…
E-InMeMo: Enhanced Prompting for Visual In-Context Learning
Jiahao Zhang, Bowen Wang, Hong Liu +3
Large-scale models trained on extensive datasets have become the standard due to their strong generalizability across diverse tasks. In-context learning (ICL), widely used in natur…