30 citations · 39 across the 7 of their papers we have counts for
7 papers
Valley3: Scaling Omni Foundation Models for E-commerce
Zeyu Chen, Guanghao Zhou, Qixiang Yin +6
In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understanding and reasoning capabilitie…
MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline
Huanjin Yao, Qixiang Yin, Min Yang +5
We aim to develop a multimodal research agent capable of explicit reasoning and planning, multi-tool invocation, and cross-modal information synthesis, enabling it to conduct deep…
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
Zeyu Chen, Huanjin Yao, Ziwang Zhao +1
Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging paradigm across various domains. Evaluating…
LLMvsSmall Model? Large Language Model Based Text Augmentation Enhanced Personality Detection Model
Linmei Hu, Hongyu He, Duokang Wang +3
Personality detection aims to detect one's personality traits underlying in social media posts. One challenge of this task is the scarcity of ground-truth personality traits which…
Valley: Video Assistant with Large Language model Enhanced abilitY
Ruipu Luo, Ziwang Zhao, Min Yang +6
Large Language Models (LLMs), with remarkable conversational capability, have emerged as AI assistants that can handle both visual and textual modalities. However, their effectiven…
Multimodal Matching-aware Co-attention Networks with Mutual Knowledge Distillation for Fake News Detection
Linmei Hu, Ziwang Zhao, Weijian Qi +2
Fake news often involves multimedia information such as text and image to mislead readers, proliferating and expanding its influence. Most existing fake news detection methods appl…