From the 1 of 6 linked papers with an AI index.
6 papers
Baikal: Structured Search for Deep Research over Data Lakes
Dhruv Agarwal, Rishitha Guttapalle Mohan, Aarti Kumari +5
Baikal is a framework that clusters heterogeneous tables and passages into semantic regions and uses adaptive, budgeted search policies to guide an LLM agent in generating subquest…
Balancing Multi-modal Sensor Learning via Multi-objective Optimization
Heshan Fernando, Quan Xiao, Parikshit Ram +4
Learning-enabled control systems increasingly rely on multiple sensing modalities (e.g., vision, audio, language, etc.) for perception and decision support. A key challenge is that…
Language Model Representations for Efficient Few-Shot Tabular Classification
Inwon Kang, Parikshit Ram, Yi Zhou +2
The Web is a rich source of structured data in the form of tables, from product catalogs and knowledge bases to scientific datasets. However, the heterogeneity of the structure and…
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
Heshan Fernando, Han Shen, Parikshit Ram +4
The post-training of LLMs, which typically consists of the supervised fine-tuning (SFT) stage and the preference learning stage (RLHF or DPO), is crucial to effective and safe LLM…
On the Utility of Domain-Adjacent Fine-Tuned Model Ensembles for Few-shot Problems
Md Ibrahim Ibne Alam, Parikshit Ram, Soham Dan +2
Large Language Models (LLMs) have been observed to perform well on a wide range of downstream tasks when fine-tuned on domain-specific data. However, such data may not be readily a…
On Learning Representations for Tabular Data Distillation
Inwon Kang, Parikshit Ram, Yi Zhou +2
Dataset distillation generates a small set of information-rich instances from a large dataset, resulting in reduced storage requirements, privacy or copyright risks, and computatio…