31 citations · 78 across the 45 of their papers we have counts for
30 papers · 1 filter
Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?
Jiankun Wang, Yisen Gao, Ziwei Zhang +3
Visual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all available evidence should be p…
Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
Jiaxin Bai, Jiaxuan Xiong
Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for late…
VisualPatchWorld: Code World Models as Latent Structured Representations for Planning
Jiaxin Bai, Jiaxuan Xiong
Different research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a form that supports perception,…
SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding
Yueming Wang, Tianshi Zheng, Jiaxin Bai +3
Scientific discovery increasingly relies on automated systems that generate hypotheses, inspect multimodal evidence, and validate claims at scale. Yet scientific claim verification…
SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents
Qiao Xiao, Haochen Shi, Yisen Gao +9
Large language model (LLM) agents increasingly rely on agent harnesses that manage context, tools, and multi-turn execution, making tools a central interface for acting in realisti…
PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
Jiaxin Bai, Yue Guo, Yifei Dong +13
World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environment returns text observations after each ac…