Score: 0

OODEval: Evaluating Large Language Models on Object-Oriented Design

Published: January 12, 2026 | arXiv ID: 2601.07602v1

By: Bingxu Xiao , Yunwei Dong , Yiqi Tang and more

Recent advances in large language models (LLMs) have driven extensive evaluations in software engineering. however, most prior work concentrates on code-level tasks, leaving software design capabilities underexplored. To fill this gap, we conduct a comprehensive empirical study evaluating 29 LLMs on object-oriented design (OOD) tasks. Owing to the lack of standardized benchmarks and metrics, we introduce OODEval, a manually constructed benchmark comprising 50 OOD tasks of varying difficulty, and OODEval-Human, the first human-rated OOD benchmark, which includes 940 undergraduate-submitted class diagrams evaluated by instructors. We further propose CLUE (Class Likeness Unified Evaluation), a unified metric set that assesses both global correctness and fine-grained design quality in class diagram generation. Using these benchmarks and metrics, we investigate five research questions: overall correctness, comparison with humans, model dimension analysis, task feature analysis, and bad case analysis. The results indicate that while LLMs achieve high syntactic accuracy, they exhibit substantial semantic deficiencies, particularly in method and relationship generation. Among the evaluated models, Qwen3-Coder-30B achieves the best overall performance, rivaling DeepSeek-R1 and GPT-4o, while Gemma3-4B-IT outperforms GPT-4o-Mini despite its smaller parameter scale. Although top-performing LLMs nearly match the average performance of undergraduates, they remain significantly below the level of the best human designers. Further analysis shows that parameter scale, code specialization, and instruction tuning strongly influence performance, whereas increased design complexity and lower requirement readability degrade it. Bad case analysis reveals common failure modes, including keyword misuse, missing classes or relationships, and omitted methods.

Holistic Evaluation of State-of-the-Art LLMs for Code Generation

Software Engineering

Makes computers write better, error-free code.

19 Dec 2025 1

89%

CodeEval: A pedagogical approach for targeted evaluation of code-trained Large Language Models

Software Engineering

Tests AI's ability to write computer code.

6 Jan 2026 0

89%

Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code

Software Engineering

Computers struggle to fix messy code on their own.

25 Nov 2025 1

View PDF Login to Bookmark

OODEval: Evaluating Large Language Models on Object-Oriented Design

Technical Abstract

Holistic Evaluation of State-of-the-Art LLMs for Code Generation

CodeEval: A pedagogical approach for targeted evaluation of code-trained Large Language Models

Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code