4 papers
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
Jiaqi Li, Shuntian Zheng, Yixian Shen +4
Video Temporal Grounding (VTG) localizes the temporal boundaries of query-relevant moments in long, untrimmed videos, making video-language-model prohibitively expensive. While rec…
Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action Localization
Jiaqi Li, Guangming Wang, Shuntian Zheng +4
Temporal Action Localization (TAL) requires identifying both the boundaries and categories of actions in untrimmed videos. While vision-language models (VLMs) offer rich semantics…
Why Learn What Physics Already Knows? Realizing Agile mmWave-based Human Pose Estimation via Physics-Guided Preprocessing
Shuntian Zheng, Jiaqi Li, Minzhe Ni +2
We revisit millimeter-wave (mmWave) human pose estimation (HPE) from a signal preprocessing perspective. A single mmWave frame provides structured dimensions that map directly to h…
Person Parametric Physics-informed Representation for mmWave-based Human Pose Estimation
Shuntian Zheng, Jiaqi Li, Guangming Wang +4
Millimeter-wave (mmWave) radar enables privacy-preserving, illumination-invariant Human Pose Estimation (HPE). However, current mmWave-based HPE systems face a signal-noise dilemma…