3 papers
cs.CV2025
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
Junha Song, Yongsik Jo, So Yeon Min +4
Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal la…
cs.LG2025
Self-Regulation and Requesting Interventions
So Yeon Min, Yue Wu, Jimin Sun +4
Human intelligence involves metacognitive abilities like self-regulation, recognizing limitations, and seeking assistance only when needed. While LLM Agents excel in many domains,…
cs.RO2025
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation
Quanting Xie, So Yeon Min, Pengliang Ji +7
There is no limit to how much a robot might explore and learn, but all of that knowledge needs to be searchable and actionable. Within language research, retrieval augmented genera…