2 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2025
START: Spatial and Textual Learning for Chart Understanding
Zhuoming Liu, Xiaofeng Gao, Feiyang Niu +3
Chart understanding is crucial for deploying multimodal large language models (MLLMs) in real-world scenarios such as analyzing scientific papers and technical reports. Unlike natu…
cs.AI2025★ 2 cited
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…
cs.CV2022★ 2 cited
A Multi-level Alignment Training Scheme for Video-and-Language Grounding
Yubo Zhang, Feiyang Niu, Qing Ping +1
To solve video-and-language grounding tasks, the key is for the network to understand the connection between the two modalities. For a pair of video and language description, their…