From the 1 of 5 linked papers with an AI index.
5 papers
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs
Woongkyu Lee, Jungwook Choi
The paper empirically studies how different inference-time scaling strategies affect the performance and failure modes of locally deployed autonomous computer-use agents under hard…
ReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training
Janghwan Lee, Sihwa Lee, Jinseok Kim +4
Large Reasoning Models (LRMs) achieve strong problem-solving through long chain-of-thought, but their deployment is constrained by the high cost of full-precision inference and gro…
ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling
Sihwa Lee, Janghwan Lee, Donghoon Yoo +4
Large reasoning models (LRMs) improve complex problem-solving by generating long intermediate reasoning traces, but this substantially increases inference costs. NVFP4 inference of…
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
Jung Hyun Lee, June Yong Yang, Jungwook Choi +1
As large language models continue to scale, low-bit weight-only post-training quantization (PTQ) offers a practical solution to their memory-efficient deployment. Although block-wi…
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
Woongkyu Lee, Junhee Cho, Jungwook Choi
Large language models (LLMs) have advanced code generation from single-function tasks to competitive-programming problems, but existing multi-agent solutions either rely on costly…