papers

Publications (18)

cs.LG2026

Adapting Feature Attenuation to NLP

Tianshuo Yang, Ryan Rabinowitz, Terrance E. Boult +1

Transformer classifiers such as BERT deliver impressive closed-set accuracy, yet they remain brittle when confronted with inputs from unseen categories--a common scenario for deplo…

cs.RO2025

AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory

Zhiqian Lan, Yuxuan Jiang, Ruiqi Wang +12

Vision-language-action (VLA) models have shown promise as generalist robotic policies by jointly leveraging visual, linguistic, and proprioceptive modalities to generate action tra…

math.NT2026

Matrix Kloosterman Sums, Random Matrix Statistics, and Cryptography

Tianshuo Yang

This paper presents a comprehensive study of matrix Kloosterman sums, including their computational aspects, distributional behavior, and applications in cryptographic analysis. Bu…

cs.CV2024

Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Peng Gao, Le Zhuo, Dongyang Liu +17

Sora unveils the potential of scaling Diffusion Transformer for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lac…

cs.CV2025

InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models

Nianchen Deng, Lixin Gu, Shenglong Ye +17

Recent benchmarks and datasets have been proposed to improve spatial reasoning in vision-language models (VLMs), yet existing open resources remain limited in scale, visual diversi…

cs.CV2026

AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model

Yutian Chen, Shi Guo, Renbiao Jin +7

Sparse-view 3D reconstruction is essential for modeling scenes from casual captures, but remain challenging for non-generative reconstruction. Existing diffusion-based approaches m…

cs.LG2025

LemmaHead: RAG Assisted Proof Generation Using Large Language Models

Tianbo Yang, Mingqi Yan, Hongyi Zhao +1

Developing the logic necessary to solve mathematical problems or write mathematical proofs is one of the more difficult objectives for large language models (LLMS). Currently, the…

cs.AI2026

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Haotian Liang, Mingkang Chen, Yufei Huang +27

The paper introduces RxBrain, a foundation model that jointly reasons over language and visual inputs to create embodied plans, using a multimodal Mixture-of-Transformers architect…

#embodied cognition#multimodal reasoning#language-visual planning#foundation models
cs.LG2023

Classified as unknown: A novel Bayesian neural network

Tianbo Yang, Tianshuo Yang

We establish estimations for the parameters of the output distribution for the softmax activation function using the probit function. As an application, we develop a new efficient…

cs.CV2026

4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture

Yutian Chen, Shi Guo, Tianshuo Yang +4

Reconstructing fast-dynamic scenes from multi-view videos is crucial for high-speed motion analysis and realistic 4D reconstruction. However, the majority of 4D capture systems are…

cs.IT2023

Finding the minimum distance and decoding linear codes with Gaussian Elimination Method

Tianshuo Yang

We propose an algorithm using the Gaussian elimination method to find the minimal Hamming distance and decode received messages of linear codes. This algorithm is easy to implement…

cs.CV2026

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models

Dehui Wang, Rong Wei, Yue Shi +9

The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. However, existing approaches stru…

cs.CV2026

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Zhixuan Liang, Yizhuo Li, Tianshuo Yang +9

Vision-Language-Action (VLA) models adapt large vision-language backbones to map images and instructions into robot actions. However, prevailing VLAs either generate actions autore…

cs.CV2026

HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System

Tianshuo Yang, Guanyu Chen, Yutian Chen +8

While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often compromises the profound rea…

cs.RO2025

Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS Workshop

Tianxing Chen, Kaixuan Wang, Zhaohui Yang +96

Embodied Artificial Intelligence (Embodied AI) is an emerging frontier in robotics, driven by the need for autonomous systems that can perceive, reason, and act in complex physical…

cs.CV2024

PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

Fanqing Meng, Wenqi Shao, Lixin Luo +8

Text-to-image (T2I) models have made substantial progress in generating images from textual prompts. However, they frequently fail to produce images consistent with physical common…

cs.CV2024

Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model

Lirui Zhao, Tianshuo Yang, Wenqi Shao +5

This paper addresses an important problem of object addition for images with only text guidance. It is challenging because the new object must be integrated seamlessly into the ima…

cs.CV2025

MM-ACT: Learn from Multimodal Parallel Generation to Act

Haotian Liang, Xinyi Chen, Bin Wang +12

A generalist robotic policy needs both semantic understanding for task planning and the ability to interact with the environment through predictive capabilities. To tackle this, we…