1 citations · 1 across the 10 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.AI2024
Constrain Alignment with Sparse Autoencoders
Qingyu Yin, Chak Tou Leong, Minjun Zhu +7
The alignment of large language models (LLMs) with human preferences remains a key challenge. While post-training techniques like Reinforcement Learning from Human Feedback (RLHF)…
cs.CL2024
E2CL: Exploration-based Error Correction Learning for Embodied Agents
Hanlin Wang, Chak Tou Leong, Jian Wang +1
Language models are exhibiting increasing capability in knowledge utilization and reasoning. However, when applied as agents in embodied environments, they often suffer from misali…