Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
Tongsheng Ding, Zhen Luo, Yixuan Yang +4
Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-…
cs.CV2025
SegVec3D: A Method for Vector Embedding of 3D Objects Oriented Towards Robot manipulation
Zhihan Kang, Boyu Wang
We propose SegVec3D, a novel framework for 3D point cloud instance segmentation that integrates attention mechanisms, embedding learning, and cross-modal alignment. The approach bu…
cs.CV2024
PaLM2-VAdapter: Progressively Aligned Language Model Makes a Strong Vision-language Adapter
Junfei Xiao, Zheng Xu, Alan Yuille +2
This paper demonstrates that a progressively aligned language model can effectively bridge frozen vision encoders and large language models (LLMs). While the fundamental architectu…