Publications (53)
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
Songsong Yu, Yuxin Chen, Hao Ju +15
Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in V…
Measurements of Dy() at energies relevant for astrophysical process
Hao Cheng, Bao-Hua Sun, Li-Hua Zhu +20
Rare information on photodisintegration reactions of nuclei with mass numbers at astrophysical conditions impedes our understanding of the origin of -nuclei. Exp…
From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
Cheng Cheng, Lin Song, Di An +4
Autoregressive (AR) image generators offer a language-model-friendly approach to image generation by predicting discrete image tokens in a causal sequence. However, unlike diffusio…
GrootVL: Tree Topology is All You Need in State Space Model
Yicheng Xiao, Lin Song, Shaoli Huang +5
The state space models, employing recursively propagated features, demonstrate strong representation capabilities comparable to Transformer models and superior efficiency. However,…
Rethinking Learnable Tree Filter for Generic Feature Transform
Lin Song, Yanwei Li, Zhengkai Jiang +5
The Learnable Tree Filter presents a remarkable approach to model structure-preserving relations for semantic segmentation. Nevertheless, the intrinsic geometric constraint forces…
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
Yicheng Xiao, Lin Song, Rui Yang +4
Recent advances have highlighted the benefits of scaling language models to enhance performance across a wide range of NLP tasks. However, these approaches still face limitations i…