5 papers
HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models
Feng He, Zhenting Wang, Qifan Wang +4
Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous visual evidence. Prior work mainly focuses o…
Q-Bridge: Code Translation for Quantum Machine Learning via LLMs
Runjia Zeng, Priyabrata Senapati, Ruixiang Tang +2
Large language models have recently shown potential in bridging the gap between classical machine learning and quantum machine learning. However, the lack of standardized, high-qua…
Shifting Uncertainty to Critical Moments: Towards Reliable Uncertainty Quantification for VLA Model
Yanchuan Tang, Taowen Wang, Yuefei Chen +3
Vision-Language-Action (VLA) models enable general-purpose robotic policies by mapping visual observations and language instructions to low-level actions, but they often lack relia…
TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching
Runjia Zeng, Qifan Wang, Qiang Guan +6
Fine tuning has been regarded as a de facto approach for adapting large language models (LLMs) to downstream tasks, but the high training memory consumption inherited from LLMs mak…
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
Xinyan Cai, Shiguang Wu, Dafeng Chi +4
In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imag…