20 papers
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Kangning Zhang, Yixing Li, Shuai Shao +9
The paper proposes Visual Attribution Distillation (VAD), a counterfactual method that isolates the visual component of teacher corrections in multimodal on‑policy distillation and…
MMSkills: Towards Multimodal Skills for General Visual Agents
Kangning Zhang, Shuai Shao, Qingyao Li +8
Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable co…
DynaTree: Dynamic Agentic Retrieval Tree for Time-Sensitive News Retrieval
Siyuan Qi, Xinyuan Wang, Yingxuan Yang +5
Agentic Retrieval-Augmented Generation improves retrieval by integrating planning, tool use, and iterative reasoning, but existing agentic RAG methods often couple semantic expansi…
SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System
Shuai Pan, Yixiang Liu, Jiaye Gao +7
Large language model (LLM) agent systems are increasingly expected to improve after deployment, but existing work often decouples two adaptation targets: skill evolution and multi-…
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
Jie Gao, Yongan Yu, Junzhu Su +3
Adaptive learning refers to educational technologies that track learners' learning progress and adapt the instructional process based on individual learners' learning performance.…
Contexting as Recommendation: Evolutionary Collaborative Filtering for Context Engineering
Jiachen Zhu, Zhuoying Ou, Congmin Zheng +9
Large Language Models (LLMs) are highly sensitive to their input contexts, motivating the development of automated context engineering. However, existing methods predominantly trea…