From the 1 of 46 linked papers with an AI index.
46 papers
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Kangning Zhang, Yixing Li, Shuai Shao +9
The paper proposes Visual Attribution Distillation (VAD), a counterfactual method that isolates the visual component of teacher corrections in multimodal on‑policy distillation and…
DynaKRAG: A Unified Framework for Learnable Evidence Control in Multi-Hop Retrieval-Augmented Generation
Yaqi Wu, Xiaolei Guo, Chenyu Zhou +7
Multi-hop retrieval-augmented generation (RAG) acquires evidence sequentially, with each new document potentially revealing missing facts, bridge entities, query defects, or suffic…
DiffCold: A Diffusion-based Generative Model for Cold-Start Item Recommendation
Kangning Zhang, Yingjie Qin, Weinan Zhang +2
Cold-start item recommendation remains a persistent challenge in real-world systems due to the absence of interaction histories. While prior models attempt to bridge this gap using…
SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior
Zhiyu Chen, Zihan Guo, Bo Huang +4
Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organ…
MOTOR: Learning ID-free Item Representation with Token Crossing for Embedding-based Multimodal Recommendation
Kangning Zhang, Jiarui Jin, Yingjie Qin +4
While multimodal recommendation models have effectively integrated visual and textual information, their reliance on unique ID embeddings constitutes a fundamental performance bott…
LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
Aofan Yu, Chenyu Zhou, Tianyi Xu +8
Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substantial context overhead and e…