activity
20242026
collaborators

6 papers

cs.CV2026

M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection

Chao Wang, Wei Lu, Xiang Li +2

Single-source remote sensing object detection using optical or SAR images struggles in complex environments. Optical images offer rich textural details but are often affected by lo…

cs.CV2025

MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs

Yipeng Du, Tiehan Fan, Kepan Nan +6

Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-fra…

cs.CV2025

Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion

Zeren Xiong, Zikun Chen, Zedong Zhang +4

In this paper, we tackle a new task of 3D object synthesis, where a 3D model is composited with another object category to create a novel 3D model. However, most existing text/imag…

cs.CV2024

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

Tiehan Fan, Kepan Nan, Rui Xie +6

Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in e…

cs.CV2024

Novel Object Synthesis via Adaptive Text-Image Harmony

Zeren Xiong, Zedong Zhang, Zikun Chen +5

In this paper, we study an object synthesis task that combines an object text with an object image to create a new object image. However, most diffusion models struggle with this t…

cs.CV2024

DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine Domain

Kun Wang, Zhiqiang Yan, Junkai Fan +4

In this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial…