activity
20152026
most citedSafe Exploration in Continuous Action Spaces

275 citations · 368 across the 45 of their papers we have counts for

collaborators
Showing 2024Show all

11 papers · 1 filter

cs.CL2024★ 5 cited

LitLLMs, LLMs for Literature Review: Are we there yet?

Shubham Agarwal, Gaurav Sahu, Abhay Puri +5

Literature reviews are an essential component of scientific research, but they remain time-intensive and challenging to write, especially due to the recent influx of research paper…

cs.LG2024★ 1 cited

BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks

Juan Rodriguez, Xiangru Jian, Siba Smarak Panigrahi +40

Multimodal AI has the potential to significantly enhance document-understanding tasks, such as processing receipts, understanding workflows, extracting data from documents, and sum…

cs.LG2024

Achieving the Tightest Relaxation of Sigmoids for Formal Verification

Samuel Chevalier, Duncan Starkenburg, Krishnamurthy Dvijotham

In the field of formal verification, Neural Networks (NNs) are typically reformulated into equivalent mathematical programs which are optimized over. To overcome the inherent non-c…

cs.LG2024

Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation

Katherine M. Collins, Najoung Kim, Yonatan Bitton +15

Human feedback plays a critical role in learning and refining reward models for text-to-image generation, but the optimal form the feedback should take for learning an accurate rew…

cs.DS2024★ 1 cited

Efficient and Near-Optimal Noise Generation for Streaming Differential Privacy

Krishnamurthy Dvijotham, H. Brendan McMahan, Krishna Pillutla +2

In the task of differentially private (DP) continual counting, we receive a stream of increments and our goal is to output an approximate running total of these increments, without…

cs.LG2024

Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models

Kyuyoung Kim, Jongheon Jeong, Minyong An +4

Fine-tuning text-to-image models with reward functions trained on human feedback data has proven effective for aligning model behavior with human intent. However, excessive optimiz…