7 papers
On the Design Fundamentals of Pixel Text Representation Learning
Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang +4
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretr…
Paving the Way for Point Cloud Video Representation Learning Using A PDE Model
Zhuoxu Huang, Zhenkun Fan, Jungong Han +1
Investigating spatial-temporal correlations, specifically how spatial points vary over time, is crucial for understanding point cloud videos. Traditional methods, particularly flow…
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
Zhuoxu Huang, Mengxi Jia, Hao Sun +2
Reinforcement Learning with verifiable rewards (RLVR) has emerged as a primary learning paradigm for enhancing the reasoning capabilities of multi-modal large language models (MLLM…
Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
Zhuoxu Huang, Mingqi Gao, Jungong Han
3D object segmentation with Large Language Models (LLMs) has become a prevailing paradigm due to its broad semantics, task flexibility, and strong generalization. However, this par…
On Exploring PDE Modeling for Point Cloud Video Representation Learning
Zhuoxu Huang, Zhenkun Fan, Tao Xu +1
Point cloud video representation learning is challenging due to complex structures and unordered spatial arrangement. Traditional methods struggle with frame-to-frame correlations…
Pixel Sentence Representation Learning
Chenghao Xiao, Zhuoxu Huang, Danlu Chen +7
Pretrained language models are long known to be subpar in capturing sentence and document-level semantics. Though heavily investigated, transferring perturbation-based methods from…