most citedLoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models

1 citations · 1 across the 5 of their papers we have counts for

collaborators

7 papers

cs.RO2025

From Power to Precision: Learning Fine-grained Dexterity for Multi-fingered Robotic Hands

Jianglong Ye, Lai Wei, Guangqi Jiang +3

Human grasps can be roughly categorized into two types: power grasps and precision grasps. Precision grasping enables tool use and is believed to have influenced human evolution. T…

cs.RO2025

GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation

Guangqi Jiang, Haoran Chang, Ri-Zhao Qiu +6

This paper presents GSWorld, a robust, photo-realistic simulator for robotics manipulation that combines 3D Gaussian Splatting with physics engines. Our framework advocates "closin…

cs.AI2025

Real Deep Research for AI, Robotics and Beyond

Xueyan Zou, Jianglong Ye, Hao Zhang +7

With the rapid growth of research in AI and robotics now producing over 10,000 papers annually it has become increasingly difficult for researchers to stay up to date. Fast evolvin…

cs.CV2025

Visual Acoustic Fields

Yuelei Li, Hyunjin Kim, Fangneng Zhan +7

Objects produce different sounds when hit, and humans can intuitively infer how an object might sound based on its appearance and material properties. Inspired by this intuition, w…

cs.CV2025

M3: 3D-Spatial MultiModal Memory

Xueyan Zou, Yuchen Song, Ri-Zhao Qiu +4

We present 3D Spatial MultiModal Memory (M3), a multimodal memory system designed to retain information about medium-sized static scenes through video sources for visual perception…

cs.CV20251 cited

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models

Yuto Kojima, Jiarui Xu, Xueyan Zou +1

The rapid advancements in vision-language models (VLMs), such as CLIP, have intensified the need to address distribution shifts between training and testing datasets. Although prio…