From the 1 of 15 linked papers with an AI index.
15 papers
GeoBridge: Decoupled Semantic Conditioning for Generative Image Geolocalization
Zhiyang Dou, Xumeng Han, Fengde Peng +4
Multimodal large language models (MLLMs) have advanced image geolocalization mainly by improving how they reason about geographic cues. How that reasoning isdecoded into coordinate…
Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning
Shilin Shan, Chuhao Zhou, Ruize Wang +30
Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capability is particularly critical…
NeuralActuator: Neural Actuation Modeling for Robot Dynamics and External Force Perception
Zhiyang Dou, John U. Onyemelukwe, Hangxing Zhang +9
NeuralActuator is a neural network model that predicts actuator effort, external contact forces, and motor condition for low‑cost robot arms, using differentiable simulation and a…
Towards Interactive Global Geolocation Assistant
Zhiyang Dou, Zipeng Wang, Xumeng Han +3
Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision.…
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
Juncheng Ma, Jianxin Bi, Yufan Deng +19
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remai…
RigidFormer: Learning Rigid Dynamics using Transformers
Zhiyang Dou, Minghao Guo, Haixu Wu +3
Learning-based simulation of multi-object rigid-body dynamics remains difficult because contact is discontinuous and errors compound over long horizons. Most existing methods remai…