Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MemoGen: Can Past Experience Improve Future Text-to-Image Generation?
Wenshuo Chen, Kuimou Yu, Bowen Tian +10
Modern text-to-image models have achieved strong visual synthesis, yet remain unreliable when prompts require implicit visual constraints, relational reasoning, or external knowled…
cs.CV2024
Zero-shot Prompt-based Video Encoder for Surgical Gesture Recognition
Mingxing Rao, Yinhong Qin, Soheil Kolouri +2
Purpose: In order to produce a surgical gesture recognition system that can support a wide variety of procedures, either a very large annotated dataset must be acquired, or fitted…