Publications (4)
IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks
Spencer Mateega, Jeff Yang, Tiana Costello +3
IDE-Bench is a comprehensive framework for evaluating AI IDE agents on real-world software engineering tasks through an IDE-native tool interface. We present a Dockerized test harn…
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling
Jeff Yang, Duy-Khanh Vu, Minh-Tien Nguyen +3
This paper introduces layout-aware graph modeling for multimodal RAG. Different from traditional RAG methods that mostly deal with flat text chunks, the proposed method takes into…
Automatic Prompt Selection for Large Language Models
Viet-Tung Do, Van-Khanh Hoang, Duy-Hung Nguyen +5
Large Language Models (LLMs) can perform various natural language processing tasks with suitable instruction prompts. However, designing effective prompts manually is challenging a…
When Giant Language Brains Just Aren't Enough! Domain Pizzazz with Knowledge Sparkle Dust
Minh-Tien Nguyen, Duy-Hung Nguyen, Shahab Sabahi +3
Large language models (LLMs) have significantly advanced the field of natural language processing, with GPT models at the forefront. While their remarkable performance spans a rang…