4 papers
Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance
Kaustav Kundu, Ritvik Shrivastava, Maxim Arap +13
We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \…
CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
Jiaqi Wang, Xiao Yang, Kai Sun +38
Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-…
Improving Code Reviewer Recommendation: Accuracy, Latency, Workload, and Bystanders
Peter C. Rigby, Seth Rogers, Sadruddin Saleem +5
The code review team at Meta is continuously improving the code review process. To evaluate the new recommenders, we conduct three A/B tests which are a type of randomized controll…
Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs
Yung-Chieh Chan, George Pu, Apaar Shanker +4
As large language models (LLMs) are applied to more use cases, creating high quality, task-specific datasets for fine-tuning becomes a bottleneck for model improvement. Using high…