collaborators

5 papers

cs.CV2026

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

Qirui Wang, Jingyi He, Yining Pan +3

Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and autonomous driving. However, ex…

cs.CV2025

SimROD: A Simple Baseline for Raw Object Detection with Global and Local Enhancements

Haiyang Xie, Xi Shen, Shihua Huang +2

Most visual models are designed for sRGB images, yet RAW data offers significant advantages for object detection by preserving sensor information before ISP processing. This enable…

cs.CL2025

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

Huizhen Shu, Xuying Li, Qirui Wang +3

With the rapid proliferation of Natural Language Processing (NLP), especially Large Language Models (LLMs), generating adversarial examples to jailbreak LLMs remains a key challeng…

cs.CL2025

Spatial Speech Translation: Translating Across Space With Binaural Hearables

Tuochao Chen, Qirui Wang, Runlin He +1

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spat…

cs.CV2025

Multi-Modal Video Feature Extraction for Popularity Prediction

Haixu Liu, Wenning Wang, Haoxiang Zheng +4

This work aims to predict the popularity of short videos using the videos themselves and their related features. Popularity is measured by four key engagement metrics: view count,…