5 papers
What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations
Samuel Pagon, Yixuan Shen, Vishal Asnani +1
As deepfake generators approach perceptual indistinguishability, reliable detection becomes critical. Yet, detectors that score well on benchmarks routinely fail in the wild. A con…
Non-Colliding Biometric Identities for Digital Entities: Geometry, Capacity, and Million-Scale Virtual Identity Provisioning
Yuyang Ji, Yixuan Shen, Anil Jain +2
Digital entities such as AI agents and humanoid robots increasingly operate alongside real humans, yet their identity infrastructure is based on credentials rather than embodied bi…
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
Yuyang Ji, Yixuan Shen, Shengjie Zhu +2
We present BioCoach, a biomechanics-grounded vision--language framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, thro…
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
Yixuan Shen, Peng He, Honglu Liu +6
K-12 science classrooms are rich sites of inquiry where students coordinate phenomena, evidence, and explanatory models through discourse; yet, the multimodal complexity of these i…
IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition
Yuyang Ji, Yixuan Shen, Kien Nguyen +2
Video-based person recognition achieves robust identification by integrating face, body, and gait. However, current systems waste computational resources by processing all modaliti…