1 citations · 2 across the 12 of their papers we have counts for
13 papers
Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
Hengyuan Xu, Wei Cheng, Yumeng Ji +4
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an imag…
FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining
Jinghong Lan, Wei Cheng, Yunuo Chen +10
Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting the style of a separate style r…
MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling
Shubo Lin, Xuanyang Zhang, Wei Cheng +3
Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address…
GEditBench v2: A Human-Aligned Benchmark for General Image Editing
Zhangqi Jiang, Zheng Sun, Xianfang Zeng +7
Recent advances in image editing have enabled models to handle complex instructions with impressive realism. However, existing evaluation frameworks lag behind: current benchmarks…
Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splatting
Shuangkang Fang, I-Chao Shen, Xuanyang Zhang +5
Recent 3D Gaussian Splatting (3DGS) Dropout methods address overfitting under sparse-view conditions by randomly nullifying Gaussian opacities. However, we identify a neighbor comp…
Native 3D Editing with Full Attention
Weiwei Cai, Shuangkang Fang, Weicai Ye +7
Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimiza…