4 papers
Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation
Zeteng Lin, Xingxing Li, Wen You +4
Existing Vision Language Models (VLMs) often struggle to preserve logic, entity identity, and artistic style during extended, interleaved image-text interactions. We identify this…
Understanding GUI Agent Localization Biases through Logit Sharpness
Xingjian Tao, Yiwei Wang, Yujun Cai +2
Multimodal large language models (MLLMs) have enabled GUI agents to interact with operating systems by grounding language into spatial actions. Despite their promising performance,…
Learning Few-Step Diffusion Models by Trajectory Distribution Matching
Yihong Luo, Tianyang Hu, Jiacheng Sun +2
Accelerating diffusion model sampling is crucial for efficient AIGC deployment. While diffusion distillation methods -- based on distribution matching and trajectory matching -- re…
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
Xingjian Tao, Yiwei Wang, Yujun Cai +2
Large language models (LLMs) have shown promise as parametric knowledge bases, but often underperform on question answering (QA) tasks due to hallucinations and uncertainty. While…