papers

Publications (78)

cs.CV2023

Adaptive Frequency Filters As Efficient Global Token Mixers

Zhipeng Huang, Zhizheng Zhang, Cuiling Lan +3

Recent vision transformers, large-kernel CNNs and MLPs have attained remarkable successes in broad vision tasks thanks to their effective information fusion in the global scope. Ho…

cs.RO2026

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

Zhiyuan Feng, Qixiu Li, Huizhi Liang +12

Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large…

cs.LG2025

Optimizing Large Language Model Training Using FP4 Quantization

Ruizhe Wang, Yeyun Gong, Xiao Liu +5

The growing computational demands of training large language models (LLMs) necessitate more efficient methods. Quantized training presents a promising solution by enabling low-bit…

cs.CL2025

InfoAgent: Advancing Autonomous Information-Seeking Agents

Gongrui Zhang, Jialiang Zhu, Ruiqi Yang +15

Building Large Language Model agents that expand their capabilities by interacting with external tools represents a new frontier in AI research and applications. In this paper, we…

cs.CV2021

Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

Ze Liu, Yutong Lin, Yue Cao +5

This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer fro…

cs.LG2024

Simplified Diffusion Schrödinger Bridge

Zhicong Tang, Tiankai Hang, Shuyang Gu +2

This paper introduces a novel theoretical simplification of the Diffusion Schrödinger Bridge (DSB) that facilitates its unification with Score-based Generative Models (SGMs), addr…