5 papers
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
Yixu Feng, Zinan Zhao, Yanxiang Ma +4
Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deployment. Existing token pruning m…
Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization
Zijian Wang, Yanxiang Ma, Chang Xu
Chain-of-Thought (CoT) reasoning is a critical capability for large language models (LLMs), enabling them to tackle com- plex multi-step tasks. While base LLMs, pre-trained on gene…
FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis
Wenyan Xu, Dawei Xiang, Yue Liu +6
Pure time series forecasting tasks typically focus exclusively on numerical features; however, real-world financial decision-making demands the comparison and analysis of heterogen…
Rethinking Causal Mask Attention for Vision-Language Inference
Xiaohuan Pei, Tao Huang, YanXiang Ma +1
Causal attention has become a foundational mechanism in autoregressive vision-language models (VLMs), unifying textual and visual inputs under a single generative framework. Howeve…
Learning Mask Invariant Mutual Information for Masked Image Modeling
Tao Huang, Yanxiang Ma, Shan You +1
Masked autoencoders (MAEs) represent a prominent self-supervised learning paradigm in computer vision. Despite their empirical success, the underlying mechanisms of MAEs remain ins…