2 papers
cs.CV2025
Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation
Zeteng Lin, Xingxing Li, Wen You +4
Existing Vision Language Models (VLMs) often struggle to preserve logic, entity identity, and artistic style during extended, interleaved image-text interactions. We identify this…
cs.CV2025
IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants
Vivek Chavan, Yasmina Imgrund, Tung Dao +5
We introduce IndEgo, a multimodal egocentric and exocentric dataset addressing common industrial tasks, including assembly/disassembly, logistics and organisation, inspection and r…