2 papers
cs.RO2026
FAST-EQA: Efficient Embodied Question Answering with Global and Local Region Relevancy
Haochen Zhang, Nirav Savaliya, Faizan Siddiqui +1
Embodied Question Answering (EQA) combines visual scene understanding, goal-directed exploration, spatial and temporal reasoning under partial observability. A central challenge is…
cs.CV2026
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
Haochen Zhang, Animesh Sinha, Felix Juefei-Xu +8
Conversational image generation requires a model to follow user instructions across multiple rounds of interaction, grounded in interleaved text and images that accumulate as chat…