8 papers
A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology
Dichang Zhang, Jiaqi Deng, Yixuan Shao +7
Existing astronomy foundation models provide strong galaxy representations, but adapting them to new survey conditions and survey-specific morphology recognition tasks still requir…
ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport
Zhiqiang Lao
One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-ener…
TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion
Nurislam Tursynbek, Zhiqiang Lao, Heather Yu +2
Recent text-to-video diffusion transformers generate visually compelling frames, yet still struggle with temporal coherence, often producing flickering, drifting, or unstable motio…
Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision
Jiali Cui, Zhiqiang Lao, Heather Yu
Energy-based models (EBMs) are a flexible class of deep generative models and are well-suited to capture complex dependencies in multimodal data. However, learning multimodal EBM b…
SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model
Yu Guo, Zhiqiang Lao, Xiyun Song +2
Glass surfaces create complex interactions of reflected and transmitted light, making single-image reflection removal (SIRR) challenging. Existing datasets suffer from limited phys…
Neural Geometry Image-Based Representations with Optimal Transport (OT)
Xiang Gao, Yuanpeng Liu, Xinmu Wang +7
Neural representations for 3D meshes are emerging as an effective solution for compact storage and efficient processing. Existing methods often rely on neural overfitting, where a…