2 papers
cs.CV2026
TGIF: Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
Chenchen Lin, Sanbao Su, Rachel Luo +4
Multimodal large language models (MLLMs) typically rely on a single late-layer feature from a frozen vision encoder, leaving the encoder's rich hierarchy of visual cues under-utili…
cs.CV2025
-OCC: Uncertainty-Aware Camera-based 3D Semantic Occupancy Prediction
Sanbao Su, Nuo Chen, Chenchen Lin +3
In the realm of autonomous vehicle perception, comprehending 3D scenes is paramount for tasks such as planning and mapping. Camera-based 3D Semantic Occupancy Prediction (OCC) aims…