12 papers
ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE
Yunkai Yang, Yudong Zhang, Xinying Chen +7
Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrating this capability into uni…
Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering
Fan Wei, Siru Zhong, Runmin Dong +3
Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods…
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
Zhaoyang Luo, Runmin Dong, Miao Yang +4
Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Existing acceleration methods usually remov…
Exascale Hybrid Numerical-AI Ensembles for Operational Flood-Season Forecasting in East Asia: 15-km Decadal Hindcasts and 1-km High-Resolution Capability
Mengxuan Chen, Yunpu Xu, Qiuyan Sun +19
Seasonal forecasting of summer rainfall in East Asia remains a grand challenge, as predictability at 3 to 6 month lead times is constrained by the spring predictability barrier, we…
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution
Bin Luo, Runmin Dong, Zhaoyang Luo +4
Diffusion-based methods demonstrate significant potential for remote sensing image super-resolution at large scaling factors, particularly in reference-based super-resolution (RefS…
Continual Test-Time Adaptation for Object Detection with Adaptive Monitoring and Randomized Restoration
Shilei Cao, Juepeng Zheng, Yan Liu +5
Real-world application models are commonly deployed in dynamic environments, where the target domain distribution undergoes temporal changes. Continual Test-Time Adaptation (CTTA)…