3 citations · 3 across the 4 of their papers we have counts for
8 papers
Dual-Grained Agent Memory and Shapley Context Attribution for Multimodal Agentic Learner
Jieke Wang, Tiancheng Shen, Yibo Yang +1
Frontier multimodal large language models (MLLMs) deliver impressive perception yet still falter on scientific and mathematical reasoning. Parameter-level adaptation is unavailable…
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment
Han-Jun Ko, Jr-Jen Chen, Haobo Yuan +4
Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes are prominent: hallu…
H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions
Shiping Zhu, Yibo Yang, Zhengyang Wang +3
Large language model agents are increasingly deployed in human-human interaction settings, such as meeting assistants and clinical documentation systems, where they must observe co…
High-Quality Entity Segmentation and Grounding
Lu Qi, Yi-Wen Chen, Tao Zhang +4
In this work, we propose ESG, a pipeline for high-quality entity segmentation and grounding supported by a new dataset EntitySeg. At first, the proposed dataset naming EntitySeg co…
ParaCook: On Time-Efficient Planning for Multi-Agent Systems
Shiqi Zhang, Xinbei Ma, Yunqing Xu +7
Large Language Models (LLMs) exhibit strong reasoning abilities for planning long-horizon, real-world tasks, yet existing agent benchmarks focus on task completion while neglecting…
Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence
Yibo Yang, Sihao Liu, Chuan Rao +5
Conventional low-rank adaptation methods build adapters without considering data context, leading to sub-optimal fine-tuning performance and severe forgetting of inherent world kno…