8 papers
What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective
Guangzeng Han, Xiaolei Huang
Instruction-tuning datasets often contain substantial redundancy and low-quality samples, necessitating effective data selection methods. We propose an instruction data selection f…
Knowledge-driven Augmentation and Retrieval for Integrative Temporal Adaptation
Weisi Liu, Guangzeng Han, Xiaolei Huang
Time introduces fundamental challenges in model development and deployment: models are usually trained on historical data while deployed on future data where semantic distributions…
Model-Agnostic Meta Learning for Class Imbalance Adaptation
Hanshu Rao, Guangzeng Han, Xiaolei Huang
Class imbalance is a widespread challenge in NLP tasks, significantly hindering robust performance across diverse domains and applications. We introduce Hardness-Aware Meta-Resampl…
From UAV Imagery to Agronomic Reasoning: A Multimodal LLM Benchmark for Plant Phenotyping
Yu Wu, Guangzeng Han, Ibra Niang Niang +6
To improve crop genetics, high-throughput, effective and comprehensive phenotyping is a critical prerequisite. While such tasks were traditionally performed manually, recent advanc…
Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
Guangzeng Han, Weisi Liu, Xiaolei Huang
Large Language Models (LLMs) excel at generating synthetic data, but ensuring its quality and diversity remains challenging. We propose Genetic Prompt, a novel framework that combi…
Cultivating Multidisciplinary AI Workforce Development on iTiger GPU Cluster: Practices and Challenges
Mayira Sharif, Guangzeng Han, Weisi Liu +1
To support rapid AI advances and broaden access to large-scale computing resources for under-resourced institutions at the Mid-South, we established the first regional mid-scale GP…