collaborators

5 papers

cs.CV2026

Xray-Visual Models: Scaling Vision models on Industry Scale Data

Shlok Mishra, Tsung-Yu Lin, Linda Wang +24

We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 b…

cs.SE2026

The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes

Redacted by arXiv

This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…

cs.CV2025

BioDet: Boosting Industrial Object Detection with Image Preprocessing Strategies

Jiaqi Hu, Hongli Xu, Junwen Huang +3

Accurate 6D pose estimation is essential for robotic manipulation in industrial environments. Existing pipelines typically rely on off-the-shelf object detectors followed by croppi…

cs.CV2025

GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation

Weihang Li, Hongli Xu, Junwen Huang +4

A key challenge in model-free category-level pose estimation is the extraction of contextual object features that generalize across varying instances within a specific category. Re…

cs.CV2025

PRISM: Probabilistic Representation for Integrated Shape Modeling and Generation

Lei Cheng, Mahdi Saleh, Qing Cheng +4

Despite the advancements in 3D full-shape generation, accurately modeling complex geometries and semantics of shape parts remains a significant challenge, particularly for shapes w…