15 papers · 1 filter
ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE
Yunkai Yang, Yudong Zhang, Xinying Chen +7
Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrating this capability into uni…
Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering
Fan Wei, Siru Zhong, Runmin Dong +3
Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods…
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
Zhaoyang Luo, Runmin Dong, Miao Yang +4
Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Existing acceleration methods usually remov…
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution
Bin Luo, Runmin Dong, Zhaoyang Luo +4
Diffusion-based methods demonstrate significant potential for remote sensing image super-resolution at large scaling factors, particularly in reference-based super-resolution (RefS…
Continual Test-Time Adaptation for Object Detection with Adaptive Monitoring and Randomized Restoration
Shilei Cao, Juepeng Zheng, Yan Liu +5
Real-world application models are commonly deployed in dynamic environments, where the target domain distribution undergoes temporal changes. Continual Test-Time Adaptation (CTTA)…
Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping
Jienan Lyu, Miao Yang, Jinchen Cai +4
Fine-grained high-resolution remote sensing mapping typically relies on localized visual features, which restricts cross-domain generalizability and often leads to fragmented predi…