3 papers
cs.RO2026
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
Suhwan Choi, Jaeyoon Jung, Sungkyung Kim +2
Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action…
cs.CL2024
Towards Efficient Visual-Language Alignment of the Q-Former for Visual Reasoning Tasks
Sungkyung Kim, Adam Lee, Junyoung Park +3
Recent advancements in large language models have demonstrated enhanced capabilities in visual reasoning tasks by employing additional encoders for aligning different modalities. W…
cs.CV2024
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
Minchan Kim, Minyeong Kim, Junik Bae +3
Hallucinations in vision-language models pose a significant challenge to their reliability, particularly in the generation of long captions. Current methods fall short of accuratel…