13 citations · 30 across the 15 of their papers we have counts for
19 papers · 1 filter
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
Ruiheng Liu, Haihong Hao, Mingfei Han +4
Advancing towards artificial superintelligence requires rich and intelligent perceptual capabilities. A critical frontier in this pursuit is overcoming the limited spatial understa…
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
Xiaoxiao Sun, Mingyang Li, Kun Yuan +7
Large Vision-Language Models (VLMs) often answer classic visual illusions "correctly" on original images, yet persist with the same responses when illusion factors are inverted, ev…
Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
Changlin Li, Jiawei Zhang, Shuhao Liu +4
Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training th…
Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models
Changlin Li, Jiawei Zhang, Zeyi Shi +3
Large-scale vision generative models, including diffusion and flow models, have demonstrated remarkable performance in visual generation tasks. However, transferring these pre-trai…
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
Haihong Hao, Mingfei Han, Changlin Li +2
Embodied navigation demands comprehensive scene understanding and precise spatial reasoning. While image-text models excel at interpreting pixel-level color and lighting cues, 3D-t…
Efficient Training of Large Vision Models via Advanced Automated Progressive Learning
Changlin Li, Jiawei Zhang, Sihao Lin +4
The rapid advancements in Large Vision Models (LVMs), such as Vision Transformers (ViTs) and diffusion models, have led to an increasing demand for computational resources, resulti…