6 papers
Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces
Zijing Shi, Meng Fang, Ling Chen
As autonomous web agents are increasingly deployed to perform real-world tasks, ensuring their safety has become a critical concern. In this work, we study web agent behavior under…
Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
Shunfeng Zheng, Yudi Zhang, Meng Fang +4
Retrieval-augmented generation (RAG) with foundation models has achieved strong performance across diverse tasks, but their capacity for expert-level reasoning-such as solving Olym…
Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach
Biao Wu, Meng Fang, Ling Chen +3
Recent advances in vision-language models have opened up new possibilities for reasoning-driven image geolocalization. However, existing approaches often rely on synthetic reasonin…
Spiral of Silence in Large Language Model Agents
Mingze Zhong, Meng Fang, Zijing Shi +5
The Spiral of Silence (SoS) theory holds that individuals with minority views often refrain from speaking out for fear of social isolation, enabling majority positions to dominate…
Monte Carlo Planning with Large Language Model for Text-Based Game Agents
Zijing Shi, Meng Fang, Ling Chen
Text-based games provide valuable environments for language-based autonomous agents. However, planning-then-learning paradigms, such as those combining Monte Carlo Tree Search (MCT…
MMAC-Copilot: Multi-modal Agent Collaboration Operating Copilot
Zirui Song, Yaohang Li, Meng Fang +6
Large language model agents that interact with PC applications often face limitations due to their singular mode of interaction with real-world environments, leading to restricted…