3 papers
cs.LG2026
Feedback Attribution and Representation Geometry: Metrics for Comparing Individual and Shared Rewards in MARL
Tasha Pais, Richard Higgins
Cooperative multi-agent RL systems routinely use team-averaged rewards, a feedback-attribution choice that gives each agent the team outcome regardless of its individual contributi…
cs.LG2026
Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling
Erik Nordby, Tasha Pais, Aviel Parrack
Linear probes can detect when language models produce outputs they "know" are wrong, a capability relevant to both deception and reward hacking. However, single-layer probes are fr…
cs.CV2025
On Extending Semantic Abstraction for Efficient Search of Hidden Objects
Tasha Pais, Nikhilesh Belulkar
Semantic Abstraction's key observation is that 2D VLMs' relevancy activations roughly correspond to their confidence of whether and where an object is in the scene. Thus, relevancy…