3 papers
cs.LG2026
Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning
Hyunjin Kim, Youngeun Nam, Jaemin Han +2
Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation. Naive large-batch fine-tuning repeatedly inc…
cs.IR2026
TimelyRAG: Semantic-Temporal Hybrid Retrieval for Time-Critical Question Answering in Overlapping-Evolving Documents
Youngeun Nam, Joeun Kim, Hwanjun Song +3
Although large language models (LLMs) and retrieval-augmented generation (RAG) have advanced open-domain question answering (QA), they remain unreliable when documents evolve throu…
cs.LG2024
Universal Time-Series Representation Learning: A Survey
Patara Trirat, Yooju Shin, Junhyeok Kang +6
Time-series data exists in every corner of real-world systems and services, ranging from satellites in the sky to wearable devices on human bodies. Learning representations by extr…