3 papers
cs.SE2026
Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair
Mehdi Bahrami, Kosaku Kimura, Satoshi Munakata +24
Industrial software-engineering teams increasingly need LLM agents that turn bug reports into correct patches, yet benchmark-scale operation adds long horizons, tool-use discipline…
cs.CL2026
Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework
Kenichirou Narita, Siqi Peng, Taku Fukui +3
Performance evaluation of Retrieval-Augmented Generation (RAG) systems within enterprise environments is governed by multi-dimensional and composite factors extending far beyond si…
cs.CL2024
A Multiple-Fill-in-the-Blank Exam Approach for Enhancing Zero-Resource Hallucination Detection in Large Language Models
Satoshi Munakata, Taku Fukui, Takao Mohri
Large language models (LLMs) often fabricate a hallucinatory text. Several methods have been developed to detect such text by semantically comparing it with the multiple versions p…