From the 1 of 34 linked papers with an AI index.
34 papers
EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent
Junlong Li, Junxi Li, Yuxiang Yang +3
The paper introduces EgoProceVQA, a new egocentric video question answering benchmark focused on procedural reasoning, and proposes EgoProceAgent, a self-exploring agent that learn…
TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts
Boyuan Chen, Zichen Dang, Chuang Yang +2
In real-world deployments, scene text detectors inevitably face distribution shifts beyond the training distribution. Prior work often depends on large-scale scene-text pretraining…
GUI-C: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
Junlong Li, Chao Hao, Lap-Pui Chau +1
Existing agentic reinforcement learning methods for GUI grounding have limitations at two levels. At the data level, current approaches typically treat all training samples equally…
SP-MoMamba: Superpixel-driven Mixture of State Space Experts for Efficient Image Super-Resolution
Wenbin Zou, Yawen Cui, Yi Wang +5
State space models (SSMs) have emerged as a powerful paradigm for efficient single-image super-resolution (SR) due to their linear complexity and long-range modeling capabilities.…
RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation
Sixu Lin, Junliang Chen, Huaiyuan Xu +8
Planning and acting in 3D environments is a fundamental capability for robotic manipulation in the real world. Although prior work has explored predictive flow planners to guide 3D…
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
Yuejiao Su, Xinshen Zhang, Zhen Ye +3
Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing multimodal large language mod…