Score: 3

E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task

Published: October 16, 2025 | arXiv ID: 2510.14509v1

By: Jingyao Liu , Chen Huang , Zhizhao Guan and more

Potential Business Impact:

Tests computer code automatically, saving time and money.

Business Areas:

Developer Tools Software

E2EDev comprises (i) a fine-grained set of user requirements, (ii) {multiple BDD test scenarios with corresponding Python step implementations for each requirement}, and (iii) a fully automated testing pipeline built on the Behave framework. To ensure its quality while reducing the annotation effort, E2EDev leverages our proposed Human-in-the-Loop Multi-Agent Annotation Framework (HITL-MAA). {By evaluating various E2ESD frameworks and LLM backbones with E2EDev}, our analysis reveals a persistent struggle to effectively solve these tasks, underscoring the critical need for more effective and cost-efficient E2ESD solutions. Our codebase and benchmark are publicly available at https://github.com/SCUNLP/E2EDev.

E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task

Software Engineering

Tests if AI can build working computer programs.

16 Oct 2025 3

88%

Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development

Software Engineering

Helps AI build better computer programs.

6 Nov 2025 1

88%

EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents

Software Engineering

Builds software faster by connecting ideas.

4 Nov 2025 1

View PDF Login to Bookmark

Country of Origin

🇨🇳 🇸🇬 China, Singapore

Repos / Data Links

github.com github.com github.com huggingface.co

Page Count

52 pages

E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task

Tests computer code automatically, saving time and money.

Technical Abstract

E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task

Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development

EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents