11 papers
A Data-Centric Framework for Detecting and Correcting Corrupted Labels
Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La +3
The performance of machine learning and deep learning models largely depends on the quality of the training data. However, the quality of the real-world datasets is often compromis…
Noise-Aware Framework for Correcting Corrupted Labels
Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La +4
High-quality labeled data is essential for training reliable ML/DL models. However, real-world datasets often contain a considerable proportion of corrupted labels, which can sever…
iML: Executable, Problem-Grounded, and Broadly Exploratory Code-Driven AutoML
Dat Le, Duc-Cuong Le, Anh-Son Nguyen +4
Automated Machine Learning (AutoML) has improved access to machine learning, yet existing techniques often remain limited in flexibility, transparency, and execution reliability. C…
Structured Exploration and Exploitation of Label Functions for Automated Data Annotation
Phong Lam, Ha-Linh Nguyen, Thu-Trang Nguyen +2
High-quality labeled data is critical for training reliable machine learning and deep learning models, yet manual annotation remains costly and error-prone. Programmatic labeling a…
Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
Thanh Trong Vu, Tuan-Dung Bui, Thu-Trang Nguyen +2
Large Language Models (LLMs) have demonstrated impressive capabilities in code generation and are increasingly integrated into the software development process. However, ensuring t…
CABENCH: Benchmarking Composable AI for Solving Complex Tasks through Composing Ready-to-Use Models
Tung-Thuy Pham, Duy-Quan Luong, Minh-Quan Duong +4
Composable AI offers a scalable and effective paradigm for tackling complex AI tasks by decomposing them into sub-tasks and solving each sub-task using ready-to-use well-trained mo…