Publications (8)
wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2
James Tanner, Morgan Sonderegger, Jane Stuart-Smith +2
While automatic tools for speech annotation are now commonplace within phonetic research pipelines, many tasks require substantial manual correction or training sets to perform acc…
Automatic classification of stop realisation with wav2vec2.0
James Tanner, Morgan Sonderegger, Jane Stuart-Smith +2
Modern phonetic research regularly makes use of automatic tools for the annotation of speech data, however few tools exist for the annotation of many variable phonetic phenomena. A…
Exploring the anatomy of articulation rate in spontaneous English speech: relationships between utterance length effects and social factors
James Tanner, Morgan Sonderegger, Jane Stuart-Smith +4
Speech rate has been shown to vary across social categories such as gender, age, and dialect, while also being conditioned by properties of speech planning. The effect of utterance…
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Lucy Xiaoyang Shi, Brian Ichter, Michael Equi +12
Generalist robots that can perform a range of different tasks in open-world settings must be able to not only reason about the steps needed to accomplish their goals, but also proc…
: a Vision-Language-Action Model with Open-World Generalization
Physical Intelligence, Kevin Black, Noah Brown +33
In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated im…
: a VLA That Learns From Experience
Physical Intelligence, Ali Amin, Raichelle Aniceto +53
We study how vision-language-action (VLA) models can improve through real-world deployments via reinforcement learning (RL). We present a general-purpose method, RL with Experience…
: A Vision-Language-Action Flow Model for General Robot Control
Kevin Black, Noah Brown, Danny Driess +21
Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artif…
Reconciling saliency and object center-bias hypotheses in explaining free-viewing fixations
Ali Borji, James Tanner
Predicting where people look in natural scenes has attracted a lot of interest in computer vision and computational neuroscience over the past two decades. Two seemingly contrastin…