From the 1 of 5 linked papers with an AI index.
5 papers
Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
Jiwon Jang, Kisu Yang, Heuiseok Lim +1
The paper evaluates 4-bit post‑training quantization of multi‑turn, tool‑calling LLM agents and finds that while standard scores remain unchanged, quantization substantially increa…
Llamion Technical Report
Kisu Yang, Yoonna Jang, Hyeonseok Moon +6
We release Llamion, a family of 14B-parameter open-weight language models obtained by transforming Orion-14B into the standardized Llama-family architecture. The transformation is…
Reliable Evaluation Protocol for Low-Precision Retrieval
Kisu Yang, Yoonna Jang, Hwanseok Jang +3
Lowering the numerical precision of model parameters and computations is widely adopted to improve the efficiency of retrieval systems. However, when computing relevance scores bet…
Dynamic Context Adaptation for Consistent Role-Playing Agents with Retrieval-Augmented Generations
Jeiyoon Park, Yongshin Han, Minseop Kim +1
Building role-playing agents (RPAs) that faithfully emulate specific characters remains challenging because collecting character-specific utterances and continually updating model…
Expanding Computation Spaces of LLMs at Inference Time
Yoonna Jang, Kisu Yang, Isabelle Augenstein
Chain-of-thought (CoT) rationale enables language models to use additional task-related text for problem-solving, benefiting not only from detailed reasoning steps but also from th…