activity
20242026
collaborators

11 papers

cs.CV2026

Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing

Lin Liu, Zhihan Xiao, Haohang Xu +4

Video editing has recently achieved remarkable progress with diffusion-based generative models, enabling diverse object-level manipulations from natural language instructions. Howe…

cs.CV2026

FineEdit: Fine-Grained Image Edit with Bounding Box Guidance

Haohang Xu, Lin Liu, Zhibo Zhang +3

Diffusion-based image editing models have achieved significant progress in real world applications. However, conventional models typically rely on natural language prompts, which o…

cs.CV2026

FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions

Peisen Zhao, Xiaopeng Zhang, Mingxing Xu +10

While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encode…

physics.ao-ph2026

AI Decodes Historical Chinese Archives to Reveal Lost Climate History

Sida He, Lingxi Xie, Xiaopeng Zhang +1

Historical archives contain qualitative descriptions of climate events, yet converting these into quantitative records has remained a fundamental challenge. Here we introduce a par…

cs.CV2025

UniLat3D: Geometry-Appearance Unified Latents for Single-Stage 3D Generation

Guanjun Wu, Jiemin Fang, Chen Yang +11

High-fidelity 3D asset generation is crucial for various industries. While recent 3D pretrained models show strong capability in producing realistic content, most are built upon di…

cs.CV2025

O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing

Yuqing Chen, Junjie Wang, Lin Liu +4

Diffusion models have recently advanced video editing, yet controllable editing remains challenging due to the need for precise manipulation of diverse object properties. Current m…