6 papers
CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
Yifan Yu, Moyan Li, Shaoyuan Xu +4
Multi-agent systems (MAS) are increasingly capable of tackling complex real-world tasks, yet their reliance on inter-agent coordination, tool use, and long-horizon reasoning makes…
Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context
Yilun Zhu, Yuan Zhuang, Nikhita Vedula +6
Many applications of LLM-based text regression require predicting a full conditional distribution rather than a single point value. We study distributional regression under empiric…
CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
Xinhai Hou, Shaoyuan Xu, Manan Biyani +4
Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful…
AlignFlow: Improving Flow-based Generative Models with Semi-Discrete Optimal Transport
Lingkai Kong, Molei Tao, Yang Liu +4
Flow-based Generative Models (FGMs) effectively transform noise into complex data distributions. Incorporating Optimal Transport (OT) to couple noise and data during FGM training h…
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
Binh M. Le, Shaoyuan Xu, Jinmiao Fu +6
In Visual Document Understanding (VDU) tasks, fine-tuning a pre-trained Vision-Language Model (VLM) with new datasets often falls short in optimizing the vision encoder to identify…
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
Hengkang Wang, Yang Liu, Huidong Liu +5
Video restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they…