activity
20242026
most citedHarvest Video Foundation Models via Efficient Post-Pretraining

2 citations · 2 across the 2 of their papers we have counts for

collaborators
Showing 2025Show all

8 papers · 1 filter

cs.AI2025

Truly Assessing Fluid Intelligence of Large Language Models through Dynamic Reasoning Evaluation

Yue Yang, MingKang Chen, Qihua Liu +9

Recent advances in large language models (LLMs) have demonstrated impressive reasoning capacities that mirror human-like thinking. However, whether LLMs possess genuine fluid intel…

cs.CV2025

Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation

Hao Zhang, Yongqiang Ma, Wenqi Shao +3

Capturing long-range dependencies while preserving high-resolution visual representations is crucial for dense prediction tasks such as human pose estimation. Vision Transformers (…

cs.CV2025

GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Quanfeng Lu, Wenqi Shao, Zitao Liu +7

Autonomous Graphical User Interface (GUI) navigation agents can enhance user experience in communication, entertainment, and productivity by streamlining workflows and reducing man…

cs.CV2025

SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations

Xiangchao Yan, Runjian Chen, Bo Zhang +11

Annotating 3D LiDAR point clouds for perception tasks is fundamental for many applications e.g., autonomous driving, yet it still remains notoriously labor-intensive. Pretraining-f…

cs.LG2025

EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Mengzhao Chen, Wenqi Shao, Peng Xu +4

Large language models (LLMs) are crucial in modern natural language processing and artificial intelligence. However, they face challenges in managing their significant memory requi…

cs.CV2025

MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Fanqing Meng, Lingxiao Du, Zongkai Liu +12

DeepSeek R1, and o1 have demonstrated powerful reasoning capabilities in the text domain through stable large-scale reinforcement learning. To enable broader applications, some wor…