49 citations · 97 across the 5 of their papers we have counts for
5 papers
CoV: Chain-of-View Prompting for Spatial Reasoning
Haoyu Zhao, Akide Liu, Zeyu Zhang +5
Embodied question answering (EQA) in 3D environments often requires collecting context that is distributed across multiple viewpoints and partially occluded. However, most recent v…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
Xi Cheng, Ruiyan Zhu, Ke Liu +8
We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language instructions. In this paper, we introd…
Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback
Josh Abramson, Arun Ahuja, Federico Carnevale +16
An important goal in artificial intelligence is to create agents that can both interact naturally with humans and learn from their feedback. Here we demonstrate how to use reinforc…
Imitating Interactive Intelligence
Josh Abramson, Arun Ahuja, Iain Barr +26
A common vision from science fiction is that robots will one day inhabit our physical spaces, sense the world as we do, assist our physical labours, and communicate with us through…