Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
WALT: Web Agents that Learn Tools
Viraj Prabhu, Yutong Dai, Matthew Fernandez +8
Web agents promise to automate complex browser tasks, but current methods remain brittle -- relying on step-by-step UI interactions and heavy LLM reasoning that break under dynamic…
cs.CV2024
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
Can Qin, Congying Xia, Krithika Ramakrishnan +16
We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI'…