activity
20242026
collaborators

6 papers

cs.MM2026

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

Jianghan Chao, Jianzhang Gao, Wenhui Tan +3

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of proc…

cs.CL2026

BFS-PO: Best-First Search for Large Reasoning Models

Fiorenzo Parascandolo, Wenhui Tan, Enver Sangineto +2

Large Reasoning Models (LRMs) such as OpenAI o1 and DeepSeek-R1 have shown excellent performance in reasoning tasks using long reasoning chains. However, this has also led to a sig…

cs.CL2026

Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains

Wenhui Tan, Jiaze Li, Jianzhong Ju +3

Large Language Models (LLMs) achieve superior performance through Chain-of-Thought (CoT) reasoning, but these token-level reasoning chains are computationally expensive and ineffic…

cs.HC2025

Think-Then-React: Towards Unconstrained Human Action-to-Reaction Generation

Wenhui Tan, Boyuan Li, Chuhao Jin +3

Modeling human-like action-to-reaction generation has significant real-world applications, like human-robot interaction and games. Despite recent advancements in single-person moti…

cs.RO2025

Transferring Foundation Models for Generalizable Robotic Manipulation

Jiange Yang, Wenhui Tan, Chuhao Jin +6

Improving the generalization capabilities of general-purpose robotic manipulation agents in the real world has long been a significant challenge. Existing approaches often rely on…

cs.RO2024

RoLD: Robot Latent Diffusion for Multi-task Policy Modeling

Wenhui Tan, Bei Liu, Junbo Zhang +2

Modeling generalized robot control policies poses ongoing challenges for language-guided robot manipulation tasks. Existing methods often struggle to efficiently utilize cross-data…