From the 1 of 6 linked papers with an AI index.
6 papers
Plover: Steering GUI Agents through Plan-Centric Interaction
Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo +3
Plover is a vision‑based GUI automation system that externalizes task plans, letting users inspect, edit, and correct plans to make GUI agents more transparent, controllable, and a…
RLPR: Radar-to-LiDAR Place Recognition via Two-Stage Asymmetric Cross-Modal Alignment for Autonomous Driving
Zhangshuo Qi, Jingyi Xu, Luqi Cheng +2
All-weather autonomy is critical for autonomous driving, which necessitates reliable localization across diverse scenarios. While LiDAR place recognition is widely deployed for thi…
VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents
Sam Yu-Te Lee, Chenyang Ji, Shicheng Wen +3
Text analytics has traditionally required specialized knowledge in Natural Language Processing (NLP) or text analysis, which presents a barrier for entry-level analysts. Recent adv…
Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths
Xuezhe Ma, Shicheng Wen, Linghao Jin +11
Designing a unified neural network to efficiently and inherently process sequential data with arbitrary lengths is a central and challenging problem in sequence modeling. The desig…
UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations
Zhangshuo Qi, Jingyi Xu, Luqi Cheng +3
Place recognition is a critical component of autonomous vehicles and robotics, enabling global localization in GPS-denied environments. Recent advances have spurred significant int…
MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models
Ming-Chang Chiu, Shicheng Wen, Pin-Yu Chen +1
In vision-language models (VLMs), the ability to perceive and interpret color and physical environment is crucial for achieving contextually accurate understanding and interaction.…