activity
20202025
most citedNatural Language as Policies: Reasoning for Coordinate-Level Embodied Control with LLMs

1 citations · 2 across the 5 of their papers we have counts for

collaborators

6 papers

cs.LG2025

Search-Based Adversarial Estimates for Improving Sample Efficiency in Off-Policy Reinforcement Learning

Federico Malato, Ville Hautamaki

Sample inefficiency is a long-lasting challenge in deep reinforcement learning (DRL). Despite dramatic improvements have been made, the problem is far from being solved and is espe…

eess.AS2024

ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR

Vishwanath Pratap Singh, Federico Malato, Ville Hautamaki +2

While automatic speech recognition (ASR) greatly benefits from data augmentation, the augmentation recipes themselves tend to be heuristic. In this paper, we address one of the heu…

cs.RO2024★ 1 cited

Natural Language as Policies: Reasoning for Coordinate-Level Embodied Control with LLMs

Yusuke Mikami, Andrew Melnik, Jun Miura +1

We demonstrate experimental results with LLMs that address robotics task planning problems. Recently, LLMs have been applied in robotics task planning, particularly using a code ge…

cs.AI2023★ 1 cited

Towards Solving Fuzzy Tasks with Human Feedback: A Retrospective of the MineRL BASALT 2022 Competition

Stephanie Milani, Anssi Kanervisto, Karolis Ramanauskas +27

To facilitate research in the direction of fine-tuning foundation models from human feedback, we held the MineRL BASALT Competition on Fine-Tuning from Human Feedback at NeurIPS 20…

cs.LG2022

The Transitive Information Theory and its Application to Deep Generative Models

Trung Ngo, Najwa Laabid, Ville Hautamäki +1

Paradoxically, a Variational Autoencoder (VAE) could be pushed in two opposite directions, utilizing powerful decoder model for generating realistic images but collapsing the learn…

cs.AI2020

General Characterization of Agents by States they Visit

Anssi Kanervisto, Tomi Kinnunen, Ville Hautamäki

Behavioural characterizations (BCs) of decision-making agents, or their policies, are used to study outcomes of training algorithms and as part of the algorithms themselves to enco…