collaborators

6 papers

cs.MA2026

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems

Yifan Yu, Moyan Li, Shaoyuan Xu +4

Multi-agent systems (MAS) are increasingly capable of tackling complex real-world tasks, yet their reliance on inter-agent coordination, tool use, and long-horizon reasoning makes…

cs.CL2026

Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context

Yilun Zhu, Yuan Zhuang, Nikhita Vedula +6

Many applications of LLM-based text regression require predicting a full conditional distribution rather than a single point value. We study distributional regression under empiric…

cs.CV2026

CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization

Xinhai Hou, Shaoyuan Xu, Manan Biyani +4

Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful…

cs.LG2025

AlignFlow: Improving Flow-based Generative Models with Semi-Discrete Optimal Transport

Lingkai Kong, Molei Tao, Yang Liu +4

Flow-based Generative Models (FGMs) effectively transform noise into complex data distributions. Incorporating Optimal Transport (OT) to couple noise and data during FGM training h…

cs.CV2025

QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding

Binh M. Le, Shaoyuan Xu, Jinmiao Fu +6

In Visual Document Understanding (VDU) tasks, fine-tuning a pre-trained Vision-Language Model (VLM) with new datasets often falls short in optimizing the vision encoder to identify…

cs.CV2025

Temporal-Consistent Video Restoration with Pre-trained Diffusion Models

Hengkang Wang, Yang Liu, Huidong Liu +5

Video restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they…