activity
20242026
most citedA Survey on Long Video Generation: Challenges, Methods, and Prospects

2 citations · 9 across the 21 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Earth-o1: A Grid-free Observation-native Atmospheric World Model

Junchao Gong, Kaiyi Xu, Wangxu Wei +22

Despite the unprecedented volume of multimodal data provided by modern Earth observation systems, our ability to model atmospheric dynamics remains constrained. Traditional modelin…

cs.CV2025

A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition

Peiqin Zhuang, Lei Bai, Yichao Wu +4

Recently, action recognition has been dominated by transformer-based methods, thanks to their spatiotemporal contextual aggregation capacities. However, despite the significant pro…

cs.CV2025

Native-Resolution Image Synthesis

Zidong Wang, Lei Bai, Xiangyu Yue +2

We introduce native-resolution image synthesis, a novel generative modeling paradigm that enables the synthesis of images at arbitrary resolutions and aspect ratios. This approach…

cs.CV2025

UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines

Chen Tang, Xinzhu Ma, Encheng Su +6

Traditional spatiotemporal models generally rely on task-specific architectures, which limit their generalizability and scalability across diverse tasks due to domain-specific desi…

cs.CV2025

Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs

Chongjun Tu, Peng Ye, Dongzhan Zhou +4

Multi-Modal Large Language Models (MLLMs) stand out in various tasks but still struggle with hallucinations. While recent training-free mitigation methods mostly introduce addition…

cs.CV2025

Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform

Chenyu Huang, Peng Ye, Xiaohui Wang +5

With transformer-based models and the pretrain-finetune paradigm becoming mainstream, the high storage and deployment costs of individual finetuned models on multiple tasks pose cr…