7 papers
Patch Policy: Efficient Embodied Control via Dense Visual Representations
Gaoyue Zhou, Zichen Jeff Cui, Ada Langford +3
Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation…
Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
Zichen Jeff Cui, Omar Rayyan, Haritheja Etukuru +16
The prevalent paradigm in robot learning attempts to generalize across environments, embodiments, and tasks with language prompts at runtime. A fundamental tension limits this appr…
LLM-Empowered Cooperative Content Caching in Vehicular Fog Caching-Assisted Platoon Networks
Bowen Tan, Qiong Wu, Pingyi Fan +3
This letter proposes a novel three-tier content caching architecture for Vehicular Fog Caching (VFC)-assisted platoon, where the VFC is formed by the vehicles driving near the plat…
K2-V2: A 360-Open, Reasoning-Enhanced LLM
K2 Team, Zhengzhong Liu, Liping Tang +36
We introduce K2-V2, a 360-open LLM built from scratch as a superior base for reasoning adaptation, in addition to functions such as conversation and knowledge retrieval from genera…
M3PO: Multimodal-Model-Guided Preference Optimization for Visual Instruction Following
Ruirui Gao, Emily Johnson, Bowen Tan +1
Large Vision-Language Models (LVLMs) hold immense potential for complex multimodal instruction following, yet their development is often hindered by the high cost and inconsistency…
ARIC: An Activity Recognition Dataset in Classroom Surveillance Images
Linfeng Xu, Fanman Meng, Qingbo Wu +16
The application of activity recognition in the ``AI + Education" field is gaining increasing attention. However, current work mainly focuses on the recognition of activities in man…