activity
20242026
collaborators

6 papers

cs.CR2026

Model Card for OpenAI Privacy Filter

Charles de Bourcy, Sahra Ghalebikesabi, Avi Schwarzschild +22

OpenAI Privacy Filter is a compact, bidirectional token-classification model for detecting and redacting personally identifiable information (PII) and secrets in unstructured text.…

cs.CL2025

gpt-oss-120b & gpt-oss-20b Model Card

OpenAI, :, Sandhini Agarwal +124

We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert trans…

cs.CL2025

Considering Length Diversity in Retrieval-Augmented Summarization

Juseon-Do, Jaesung Hwang, Jingun Kwon +2

This study investigates retrieval-augmented summarization by specifically examining the impact of exemplar summary lengths under length constraints, not covered by previous work. W…

cs.RO2025

Look Before You Leap: Using Serialized State Machine for Language Conditioned Robotic Manipulation

Tong Mu, Yihao Liu, Mehran Armand

Imitation learning frameworks for robotic manipulation have drawn attention in the recent development of language model grounded robotics. However, the success of the frameworks la…

cs.AI2024

Rule Based Rewards for Language Model Safety

Tong Mu, Alec Helyar, Johannes Heidecke +7

Reinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior. However, in cas…

cs.CL2024

Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions

Angana Borah, Rada Mihalcea

As Large Language Models (LLMs) continue to evolve, they are increasingly being employed in numerous studies to simulate societies and execute diverse social tasks. However, LLMs a…