Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
Yicheng Xiao, Wenhu Zhang, Lin Song +10
Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained…
cs.CV2025
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Yicheng Xiao, Lin Song, Rui Yang +6
With the advancement of language models, unified multimodal understanding and generation have made significant strides, with model architectures evolving from separated components…