4 papers
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
Xiao Wang, Jianlong Wu, Zijia Lin +3
Recently, video-language understanding has achieved great success through large-scale pre-training. However, data scarcity remains a prevailing challenge. This study quantitatively…
Polarization entanglement enabled by orthogonally stacked van der Waals NbOCl2 crystals
Qiangbing Guo, Yun-Kun Wu, Di Zhang +5
Polarization entanglement holds significant importance for photonic quantum technologies. Recently emerging subwavelength nonlinear quantum light sources, e.g., GaP and LiNbO3 thin…
EVLM: An Efficient Vision-Language Model for Visual Understanding
Kaibing Chen, Dong Shen, Hanwen Zhong +14
In the field of multi-modal language models, the majority of methods are built on an architecture similar to LLaVA. These models use a single-layer ViT feature as a visual prompt,…
PlacidDreamer: Advancing Harmony in Text-to-3D Generation
Shuo Huang, Shikun Sun, Zixuan Wang +6
Recently, text-to-3D generation has attracted significant attention, resulting in notable performance enhancements. Previous methods utilize end-to-end 3D generation models to init…