14 citations · 14 across the 1 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026★ 14 cited
Image Captioning via Compact Bidirectional Architecture
Zijie Song, Yuanen Zhou, Zhenzhen Hu +4
Most current image captioning models typically generate captions from left-to-right. This unidirectional property makes them can only leverage past context but not future context.…
cs.CV2024
ESceme: Vision-and-Language Navigation with Episodic Scene Memory
Qi Zheng, Daqing Liu, Chaoyue Wang +3
Vision-and-language navigation (VLN) simulates a visual agent that follows natural-language navigation instructions in real-world scenes. Existing approaches have made enormous pro…