Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
WorldGen: From Text to Traversable and Interactive 3D Worlds
Dilin Wang, Hyunyoung Jung, Tom Monnier +22
We introduce WorldGen, a system that enables the automatic creation of large-scale, interactive 3D worlds directly from text prompts. Our approach transforms natural language descr…
cs.CV2025
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
Geon Park, Seon Bin Kim, Gunho Jung +1
With recent advancements in text-to-image (T2I) models, effectively generating multiple instances within a single image prompt has become a crucial challenge. Existing methods, whi…
cs.CV2024
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
Kangsan Kim, Geon Park, Youngwan Lee +2
Recent advancements in video large multimodal models (LMMs) have significantly improved their video understanding and reasoning capabilities. However, their performance drops on ou…