2 papers
cs.SE2025
UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
Horia Cristescu, Charles Park, Trong Canh Nguyen +3
While current Computer Use Agent (CUA) benchmarks measure task completion effectively, they provide limited assessment of enterprise deployment readiness, emphasizing functional co…
cs.LG2025
When Embedding Models Meet: Procrustes Bounds and Applications
Lucas Maystre, Alvaro Ortega Gonzalez, Charles Park +4
Embedding models trained separately on similar data often produce representations that encode stable information but are not directly interchangeable. This lack of interoperability…