From the 2 of 28 linked papers with an AI index.
28 papers
Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?
Jiankun Wang, Yisen Gao, Ziwei Zhang +3
Visual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all available evidence should be p…
Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
Jiaxin Bai, Jiaxuan Xiong
The paper introduces Temporal-Distance JEPA, a method that learns a directed temporal cost from offline trajectories to improve latent world model predictive control, enhancing pla…
VisualPatchWorld: Code World Models as Latent Structured Representations for Planning
Jiaxin Bai, Jiaxuan Xiong
The paper presents VisualPatchWorld, a system that learns compact code programs to model world dynamics from visual observations, enabling inspection, simulation, and use in model-…
PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
Jiaxin Bai, Yue Guo, Yifei Dong +13
World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environment returns text observations after each ac…
SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding
Yueming Wang, Tianshi Zheng, Jiaxin Bai +3
Scientific discovery increasingly relies on automated systems that generate hypotheses, inspect multimodal evidence, and validate claims at scale. Yet scientific claim verification…
SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents
Qiao Xiao, Haochen Shi, Yisen Gao +9
Large language model (LLM) agents increasingly rely on agent harnesses that manage context, tools, and multi-turn execution, making tools a central interface for acting in realisti…