activity
20242026
collaborators

18 papers

cs.CV2026

Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off

Davide Lobba, Fulvio Sanguigni, Bin Ren +3

Recent advances in Virtual Try-On (VTON) and Virtual Try-Off (VTOFF) have greatly improved photo-realistic fashion synthesis and garment reconstruction. However, existing datasets…

cs.CV2026

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

Yan Shu, Bin Ren, Zhitong Xiong +4

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level vis…

cs.CV2026

Panoramic Affordance Prediction

Zixin Zhang, Chenfei Liao, Hongfei Zhang +10

Affordance prediction serves as a critical bridge between perception and action in embodied AI. However, existing research is confined to pinhole camera models, which suffer from n…

cs.CV2026

DVD: Deterministic Video Depth Estimation with Generative Priors

Hongfei Zhang, Harold Haodong Chen, Chenfei Liao +12

Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand…

cs.CV2026

Any Image Restoration via Efficient Spatial-Frequency Degradation Adaptation

Bin Ren, Eduard Zamfir, Zongwei Wu +7

Restoring multiple degradations efficiently via just one model has become increasingly significant and impactful, especially with the proliferation of mobile devices. Traditional s…

cs.CV2026

SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting

Mengjiao Ma, Qi Ma, Yue Li +10

3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven…