3 papers
cs.LG2025
Relation-Aware Graph Foundation Model
Jianxiang Yu, Jiapeng Zhu, Hao Qian +3
In recent years, large language models (LLMs) have demonstrated remarkable generalization capabilities across various natural language processing (NLP) tasks. Similarly, graph foun…
cs.CV2024
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
Tiehan Fan, Kepan Nan, Rui Xie +6
Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in e…
cs.CV2024
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
Yupeng Zheng, Xiang Li, Pengfei Li +6
Monocular Semantic Occupancy Prediction aims to infer the complete 3D geometry and semantic information of scenes from only 2D images. It has garnered significant attention, partic…