1 citations · 1 across the 5 of their papers we have counts for
7 papers
The Era of Real-World Human Interaction: RL from User Conversations
Chuanyang Jin, Jing Xu, Bo Liu +6
We posit that to achieve continual model improvement and multifaceted alignment, future models must learn from natural human interaction. Current conversational models are aligned…
ProToM: Promoting Prosocial Behaviour via Theory of Mind-Informed Feedback
Matteo Bortoletto, Yichao Zhou, Lance Ying +2
While humans are inherently social creatures, the challenge of identifying when and how to assist and collaborate with others - particularly when pursuing independent goals - can h…
Augmented Vision-Language Models: A Systematic Review
Anthony C Davis, Burhan Sadiq, Tianmin Shu +1
Recent advances in visual-language machine learning models have demonstrated exceptional ability to use natural language and understand visual scenes by training on large, unstruct…
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
Qiyue Gao, Xinyu Pi, Kevin Liu +21
Internal world models (WMs) enable agents to understand the world's state and predict transitions, serving as the basis for advanced deliberative reasoning. Recent large Vision-Lan…
Position: Foundation Models Need Digital Twin Representations
Yiqing Shen, Hao Ding, Lalithkumar Seenivasan +2
Current foundation models (FMs) rely on token representations that directly fragment continuous real-world multimodal data into discrete tokens. They limit FMs to learning real-wor…
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
Suyu Ye, Haojun Shi, Darren Shih +3
To achieve successful assistance with long-horizon web-based tasks, AI agents must be able to sequentially follow real-world user instructions over a long period. Unlike existing w…