From the 1 of 13 linked papers with an AI index.
13 papers
SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion
Wenbin Duan, Yan Shu, Zhuoyuan Fu +4
Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, achieving fast and accurate inv…
Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models
Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang +5
The paper presents a large-scale structured reasoning dataset created via slice‑wise synthesis that encodes chain‑of‑thought explanations for 3D medical images, and uses it to inst…
Bridging Auxiliary Constraints to Resolve Instruction Following in Large Reasoning Models
Zhengyi Zhao, Shubo Zhang, Huimin Wang +7
Large Reasoning Models (LRMs) have demonstrated impressive capabilities in many tasks, yet they struggle with reliably following multiple instructions, either by failing to satisfy…
Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding
Zhengyi Zhao, Shubo Zhang, Zezhong Wang +6
When asked what a meme or sarcastic post means, Large Vision Language Models (LVLMs) tend to describe what the image shows rather than what the author is trying to communicate. Sta…
Guaranteeing Knowledge Integration with Joint Decoding for Retrieval-Augmented Generation
Zhengyi Zhao, Shubo Zhang, Zezhong Wang +7
Retrieval-Augmented Generation (RAG) significantly enhances Large Language Models (LLMs) by providing access to external knowledge. However, current research primarily focuses on r…
EventWeave: A Dynamic Framework for Capturing Core and Supporting Events in Dialogue Systems
Zhengyi Zhao, Shubo Zhang, Yiming Du +5
Large language models have improved dialogue systems, but often process conversational turns in isolation, overlooking the event structures that guide natural interactions. Hence w…