5 papers
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
Qirui Wang, Jingyi He, Yining Pan +3
Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and autonomous driving. However, ex…
SimROD: A Simple Baseline for Raw Object Detection with Global and Local Enhancements
Haiyang Xie, Xi Shen, Shihua Huang +2
Most visual models are designed for sRGB images, yet RAW data offers significant advantages for object detection by preserving sensor information before ISP processing. This enable…
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
Huizhen Shu, Xuying Li, Qirui Wang +3
With the rapid proliferation of Natural Language Processing (NLP), especially Large Language Models (LLMs), generating adversarial examples to jailbreak LLMs remains a key challeng…
Spatial Speech Translation: Translating Across Space With Binaural Hearables
Tuochao Chen, Qirui Wang, Runlin He +1
Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spat…
Multi-Modal Video Feature Extraction for Popularity Prediction
Haixu Liu, Wenning Wang, Haoxiang Zheng +4
This work aims to predict the popularity of short videos using the videos themselves and their related features. Popularity is measured by four key engagement metrics: view count,…