5 papers
Xray-Visual Models: Scaling Vision models on Industry Scale Data
Shlok Mishra, Tsung-Yu Lin, Linda Wang +24
We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 b…
The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes
Redacted by arXiv
This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…
BioDet: Boosting Industrial Object Detection with Image Preprocessing Strategies
Jiaqi Hu, Hongli Xu, Junwen Huang +3
Accurate 6D pose estimation is essential for robotic manipulation in industrial environments. Existing pipelines typically rely on off-the-shelf object detectors followed by croppi…
GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation
Weihang Li, Hongli Xu, Junwen Huang +4
A key challenge in model-free category-level pose estimation is the extraction of contextual object features that generalize across varying instances within a specific category. Re…
PRISM: Probabilistic Representation for Integrated Shape Modeling and Generation
Lei Cheng, Mahdi Saleh, Qing Cheng +4
Despite the advancements in 3D full-shape generation, accurately modeling complex geometries and semantics of shape parts remains a significant challenge, particularly for shapes w…