2 papers
cs.CV2026
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
Xinpeng Dong, Min Zhang, Kairong Han +3
In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual informat…
cs.CL2024
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
Zekun Li, Zhiyu Zoey Chen, Mike Ross +7
Large language models (LLMs) are increasingly prevalent in conversational systems due to their advanced understanding and generative capabilities in general contexts. However, thei…