6 papers
PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft
Yuchen Guo, Junli Gong, Weicheng Wang +3
We present PEAM, a Parametric Embodied Agent Memory framework in Minecraft that transforms agent memory from inference-time retrieval into parameter-resident skills internalized th…
Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought
Yuchen Guo, Junli Gong, Hongmin Cai +2
Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instructions remain target-referentia…
Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment
Yuchen Guo, Junli Gong, Yao Lu +3
Infrared-Visible image fusion (IVIF) aims to integrate thermal information and detailed spatial structures into a single fused image to enhance perception. However, existing evalua…
Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models
Yuchen Guo, Junli Gong, Wenjun Dong +2
Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared im…
LumiVideo: An Intelligent Agentic System for Video Color Grading
Yuchen Guo, Junli Gong, Hongmin Cai +2
Video color grading is a critical post-production process that transforms flat, log-encoded raw footage into emotionally resonant cinematic visuals. Existing automated methods act…
Fuse4Seg: Image Fusion for Multi-Modal Medical Segmentation via Bi-level Optimization
Yuchen Guo, Junli Gong, Hongmin Cai +2
Multi-modal medical image fusion is traditionally optimized for human visual perception, aiming to maximize generic contrast and structural fidelity. However, when these visually p…