4 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.SE2026
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
cs.CV2024★ 3 cited
Face-GPS: A Comprehensive Technique for Quantifying Facial Muscle Dynamics in Videos
Juni Kim, Zhikang Dong, Pawel Polak
We introduce a novel method that combines differential geometry, kernels smoothing, and spectral analysis to quantify facial muscle activity from widely accessible video recordings…
cs.CV2023★ 4 cited
Tackling Data Bias in MUSIC-AVQA: Crafting a Balanced Dataset for Unbiased Question-Answering
Xiulong Liu, Zhikang Dong, Peng Zhang
In recent years, there has been a growing emphasis on the intersection of audio, vision, and text modalities, driving forward the advancements in multimodal research. However, stro…