activity
20212026
most citedIs Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

132 citations · 157 across the 52 of their papers we have counts for

collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI2026

Figures as Programs: Recursive Generation of Editable Scientific Figures

Yepeng Liu, Dasen Dai, Chengzhi Liu +7

Scientific methodology figures are essential for communicating complex methods clearly, yet creating them remains labor-intensive and typically requires multiple rounds of refineme…

cs.AI2026

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

Neel Mokaria, Rishie Raj, Dheeraj Baiju +13

Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestr…

cs.AI2026

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents

Binjie Zhang, Mike Zheng Shou

Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Existing works have two common g…

cs.AI2025

AUTO-Explorer: Automated Data Collection for GUI Agent

Xiangwu Guo, Difei Gao, Mike Zheng Shou

Recent advancements in GUI agents have significantly expanded their ability to interpret natural language commands to manage software interfaces. However, acquiring GUI data remain…

cs.AI2025

macOSWorld: A Multilingual Interactive Benchmark for GUI Agents

Pei Yang, Hai Ci, Mike Zheng Shou

Graphical User Interface (GUI) agents show promising capabilities for automating computer-use tasks and facilitating accessibility, but existing interactive benchmarks are mostly E…

cs.AI2025

Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models

Jiaqi Wang, Kevin Qinghong Lin, James Cheng +1

Reinforcement Learning (RL) has proven to be an effective post-training strategy for enhancing reasoning in vision-language models (VLMs). Group Relative Policy Optimization (GRPO)…