2 papers
cs.CV2026
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
Jonas Herzog, Yue Wang
Recent research suggested that the embeddings produced by CLIP-like contrastive language-image training are suboptimal for image-only tasks. The main theory is that the inter-modal…
cs.RO2025
Domain-Conditioned Scene Graphs for State-Grounded Task Planning
Jonas Herzog, Jiangpin Liu, Yue Wang
Recent robotic task planning frameworks have integrated large multimodal models (LMMs) such as GPT-4o. To address grounding issues of such models, it has been suggested to split th…