activity
20242026
most citedBenchmarking and Evaluating VLMs for Software Architecture Diagram Understanding

1 citations · 1 across the 3 of their papers we have counts for

collaborators

8 papers

cs.AI2026

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

Wenhui Chen, Shiwen Cheng, Hao Dong +39

Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumul…

cs.SE2026

Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation

Shuyin Ouyang, Zhaozhi Qian, Faroq AL-Tam +2

Reinforcement Learning (RL) is an important paradigm for aligning Diffusion Language Models (DLMs) toward functional correctness in code generation. However, these models often enc…

cs.SE20261 cited

Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding

Shuyin Ouyang, Jie M. Zhang, Jingzhi Gong +7

Software architecture diagrams are important design artifacts for communicating system structure, behavior, and data organization throughout the software development lifecycle. Alt…

cs.CV2025

ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps

Sicheng Feng, Song Wang, Shuyi Ouyang +5

Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performa…

cs.SE2025

DSCodeBench: A Realistic Benchmark for Data Science Code Generation

Shuyin Ouyang, Dong Huang, Jingwen Guo +3

We introduce DSCodeBench, a new benchmark designed to evaluate large language models (LLMs) on complicated and realistic data science code generation tasks. DSCodeBench consists of…

cs.CL2025

SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers

Yanzheng Xiang, Hanqi Yan, Shuyin Ouyang +2

This study evaluates large language models (LLMs) in generating code from algorithm descriptions in recent NLP papers. The task requires two key competencies: (1) algorithm compreh…