8 papers
DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing
Sowjanya Puligadda, Mengdie Zhang, Ali Zamani +3
As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cross-platform scalability. This p…
Scaling Mobile Chaos Testing with AI-Driven Test Execution
Juan Marcano, Ashish Samant, Kai Song +9
Mobile applications in large-scale distributed systems are susceptible to backend service failures, yet traditional chaos engineering approaches cannot scale mobile testing due to…
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Yaozu Wu, Wei-Chieh Huang, Jizhou Guo +11
Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework…
PINA: Prompt Injection Attack against Navigation Agents
Jiani Liu, Yixin He, Lanlan Fan +5
Navigation agents powered by large language models (LLMs) convert natural language instructions into executable plans and actions. Compared to text-based applications, their securi…
ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications
Changwen Xing, SamZaak Wong, Xinlai Wan +9
While Large Language Models (LLMs) demonstrate immense potential for automating integrated circuit (IC) development, their practical deployment is fundamentally limited by restrict…
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
Yaozu Wu, Jizhou Guo, Dongyuan Li +9
Effective guardrails are essential for safely deploying LLM-based agents in critical applications. Despite recent advances, existing guardrails suffer from two fundamental limitati…