activity
20242026
collaborators

7 papers

cs.CV2026

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding

Yuxiang Duan, Huining Li, Ao Li +6

Video anomaly understanding (VAU) focuses on comprehensively interpreting abnormal events in videos, requiring models to identify anomalous occurrences, discover their supporting e…

cs.RO2025

RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI

Cong Tai, Zhaoyu Zheng, Haixu Long +13

The emerging field of Vision-Language-Action (VLA) for humanoid robots faces several fundamental challenges, including the high cost of data acquisition, the lack of a standardized…

cs.AI2025

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise

Shuai Feng, Wei-Chuang Chan, Srishti Chouhan +4

The integration of large language models (LLMs) into global applications necessitates effective cultural alignment for meaningful and culturally-sensitive interactions. Current LLM…

cs.CV2024

SeMOPO: Learning High-quality Model and Policy from Low-quality Offline Visual Datasets

Shenghua Wan, Ziyuan Chen, Le Gan +2

Model-based offline reinforcement Learning (RL) is a promising approach that leverages existing data effectively in many real-world applications, especially those involving high-di…

cs.LG2024

AD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual Distractors

Yucen Wang, Shenghua Wan, Le Gan +2

Model-based methods have significantly contributed to distinguishing task-irrelevant distractors for visual control. However, prior research has primarily focused on heterogeneous…

cs.RO2024

SENSOR: Imitate Third-Person Expert's Behaviors via Active Sensoring

Kaichen Huang, Minghao Shao, Shenghua Wan +4

In many real-world visual Imitation Learning (IL) scenarios, there is a misalignment between the agent's and the expert's perspectives, which might lead to the failure of imitation…