22 citations · 45 across the 17 of their papers we have counts for
17 papers
PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification
Hao Yang, Qianyu Zhou, Haijia Sun +4
Domain Generalization (DG) has been recently explored to enhance the generalizability of Point Cloud Classification (PCC) models toward unseen domains. Prior works are based on con…
Dense360: Dense Understanding from Omnidirectional Panoramas
Yikang Zhou, Tao Zhang, Dizhe Zhang +3
Multimodal Large Language Models (MLLMs) require comprehensive visual inputs to achieve dense understanding of the physical world. While existing MLLMs demonstrate impressive world…
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
Zhucun Xue, Jiangning Zhang, Teng Hu +8
The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for vide…
On Path to Multimodal Generalist: General-Level and General-Bench
Hao Fei, Yuan Zhou, Juncheng Li +29
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolv…
Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
Muyi Bao, Shuchang Lyu, Zhaoyang Xu +5
Deep learning has profoundly transformed remote sensing, yet prevailing architectures like Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) remain constrained by…
An Empirical Study of GPT-4o Image Generation Capabilities
Sixiang Chen, Jinbin Bai, Zhuoran Zhao +16
The landscape of image generation has rapidly evolved, from early GAN-based approaches to diffusion models and, most recently, to unified generative architectures that seek to brid…