Publications (13)
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
Abhay Gupta, Michael Lu, Kevin Zhu +2
Current large language models (LLMs) struggle to answer questions that span tens of thousands of tokens, especially when multi-hop reasoning is involved. While prior benchmarks exp…
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
Michael Lu, Max Qiushi Lin, Mo Chen +1
We study policy optimization for infinite-horizon, discounted constrained Markov decision processes (CMDPs). While existing theoretical guarantees typically hold for the mixture po…
Fast Confidence-Aware Human Prediction via Hardware-accelerated Bayesian Inference for Safe Robot Navigation
Michael Lu, Minh Bui, Xubo Lyu +1
As robots increasingly integrate into everyday environments, ensuring their safe navigation around humans becomes imperative. Efficient and safe motion planning requires robots to…
Empirical Evaluation of the Segment Anything Model (SAM) for Brain Tumor Segmentation
Mohammad Peivandi, Jason Zhang, Michael Lu +2
Brain tumor segmentation presents a formidable challenge in the field of Medical Image Segmentation. While deep-learning models have been useful, human expert segmentation remains…
OptimizedDP: An Efficient, User-friendly Library For Optimal Control and Dynamic Programming
Minh Bui, Hanyang Hu, Chong He +4
This paper introduces OptimizedDP, a high-performance software library for several common grid-based dynamic programming (DP) algorithms used in control theory and robotics. Specif…
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
Shaun Baek, Shaun Esua-Mensah, Cyrus Tsui +6
Large Language Models (LLMs) are primarily trained on high-resource natural languages, limiting their effectiveness in low-resource settings and in tasks requiring deep logical rea…
Real-Time Formal Verification of Autonomous Systems With An FPGA
Minh Bui, Michael Lu, Reza Hojabr +2
Hamilton-Jacobi reachability analysis is a powerful technique used to verify the safety of autonomous systems. This method is very good at handling non-linear system dynamics with…
Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting
Chloe Ho, Ishneet Sukhvinder Singh, Diya Sharma +4
Search algorithms and user query relevance have given LLMs the ability to return relevant information, but the effect of content phrasing on ad visibility remains underexplored. We…
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
Daniel Csizmadia, Andrei Codreanu, Victor Sim +5
We present Distill CLIP (DCLIP), a fine-tuned variant of the CLIP model that enhances multimodal image-text retrieval while preserving the original model's strong zero-shot classif…
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
Michael Lu, Matin Aghaei, Anant Raj +1
We consider (stochastic) softmax policy gradient (PG) methods for bandits and tabular Markov decision processes (MDPs). While the PG objective is non-concave, recent research has u…
NewsInterview: a Dataset and a Playground to Evaluate LLMs' Ground Gap via Informational Interviews
Alexander Spangher, Michael Lu, Sriya Jeslyn Kalyan +3
Large Language Models (LLMs) have demonstrated impressive capabilities in generating coherent text but often struggle with grounding language and strategic dialogue. To address thi…
InfraredTags: Embedding Invisible AR Markers and Barcodes Using Low-Cost, Infrared-Based 3D Printing and Imaging Tools
Mustafa Doga Dogan, Ahmad Taka, Michael Lu +4
Existing approaches for embedding unobtrusive tags inside 3D objects require either complex fabrication or high-cost imaging equipment. We present InfraredTags, which are 2D marker…
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
Max Qiushi Lin, Jincheng Mei, Matin Aghaei +6
Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typica…