3 papers
cs.CL2026
Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context
Yilun Zhu, Yuan Zhuang, Nikhita Vedula +6
Many applications of LLM-based text regression require predicting a full conditional distribution rather than a single point value. We study distributional regression under empiric…
cs.CV2025
CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
Xinhai Hou, Shaoyuan Xu, Manan Biyani +4
Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful…
cs.CV2025
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
Binh M. Le, Shaoyuan Xu, Jinmiao Fu +6
In Visual Document Understanding (VDU) tasks, fine-tuning a pre-trained Vision-Language Model (VLM) with new datasets often falls short in optimizing the vision encoder to identify…