12 papers
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
Ziyao Wang, Bingying Wang, Hanrong Zhang +7
Despite remarkable progress in Vision--Language--Action (VLA) models, a central bottleneck remains underexamined: the data infrastructure that underlies embodied learning. In this…
Cognibit: From Digital Exhaustion to Real-World Connection Through Gamified Territory Control and LLM-Powered Twin Networking
Wanghao Ye, Sihan Chen, Yiting Wang +20
We present an LLM-powered social discovery platform that uses digital twins to autonomously evaluate interpersonal compatibility through behavioral simulation. The platform unifies…
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
Ziyao Wang, Chen Chen, Jingtao Li +4
Unified models aim to support both understanding and generation by encoding images into discrete tokens and processing them alongside text within a single autoregressive framework.…
FedMOA: Federated GRPO for Personalized Reasoning LLMs under Heterogeneous Rewards
Ziyao Wang, Daeun Jung, Yexiao He +4
Group Relative Policy Optimization (GRPO) has recently emerged as an effective approach for improving the reasoning capabilities of large language models through online multi-objec…
Towards Building Non-Fine-Tunable Foundation Models
Ziyao Wang, Nizhang Li, Pingzhi Li +3
Open-sourcing foundation models (FMs) enables broad reuse but also exposes model trainers to economic and safety risks from unrestricted downstream fine-tuning. We address this pro…
MindCraft: How Concept Trees Take Shape In Deep Models
Bowei Tian, Yexiao He, Wanghao Ye +3
Large-scale foundation models demonstrate strong performance across language, vision, and reasoning tasks. However, how they internally structure and stabilize concepts remains elu…