activity
20192026
most citedsPortfolio: Stratified Visual Analysis of Stock Portfolios

31 citations · 78 across the 45 of their papers we have counts for

collaborators
Showing cs.CLShow all

30 papers · 1 filter

cs.CL2026

Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?

Jiankun Wang, Yisen Gao, Ziwei Zhang +3

Visual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all available evidence should be p…

cs.CL2026

Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control

Jiaxin Bai, Jiaxuan Xiong

Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for late…

cs.CL2026

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

Jiaxin Bai, Jiaxuan Xiong

Different research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a form that supports perception,…

cs.CL2026

SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding

Yueming Wang, Tianshi Zheng, Jiaxin Bai +3

Scientific discovery increasingly relies on automated systems that generate hypotheses, inspect multimodal evidence, and validate claims at scale. Yet scientific claim verification…

cs.CL2026

SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents

Qiao Xiao, Haochen Shi, Yisen Gao +9

Large language model (LLM) agents increasingly rely on agent harnesses that manage context, tools, and multi-turn execution, making tools a central interface for acting in realisti…

cs.CL2026

PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments

Jiaxin Bai, Yue Guo, Yifei Dong +13

World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environment returns text observations after each ac…