4 papers
Visually-Guided Spatial Audio Generation for In-the-Wild Speech Scenes
Qingyu Luo, Peng Zhang, Wenwu Wang +1
Spatial audio is a key component of immersive media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided Fir…
Grammar-Guided Hierarchical Parsing for Long-form Audio Activity Recognition
Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1
Long-form audio exhibits an inherent hierarchy: fine-grained events form sub-activities, which in turn constitute higher-level activities. Prior work often models these levels sepa…
Exploring Multiple Converged States of Network Configurations
Shunyu Yang, Dan Wang, Peng Zhang
Due to the policy-rich BGP, multiple stable forwarding states might exist for the same network topology and configuration, rendering the network convergence non-deterministic. This…
Hierarchical Activity Recognition and Captioning from Long-Form Audio
Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1
Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge…