Score: 0

Evaluating Large Language Models on Non-Code Software Engineering Tasks

Published: June 12, 2025 | arXiv ID: 2506.10833v1

By: Fabian C. Peña, Steffen Herbold

Potential Business Impact:

Tests computers on understanding software tasks.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

Large Language Models (LLMs) have demonstrated remarkable capabilities in code understanding and generation; however, their effectiveness on non-code Software Engineering (SE) tasks remains underexplored. We present the first comprehensive benchmark, which we name `Software Engineering Language Understanding' (SELU), for evaluating LLMs on 17 non-code tasks, spanning from identifying whether a requirement is functional or non-functional to estimating the effort and complexity of backlog items. SELU covers classification, regression, Named Entity Recognition (NER), and Masked Language Modeling (MLM) targets, with data drawn from diverse sources such as code repositories, issue tracking systems, and developer forums. We fine-tune 22 open-source LLMs, prompt two proprietary alternatives, and train two baselines. Performance is measured using metrics such as F1-macro, SMAPE, F1-micro, and accuracy, and compared via the Bayesian signed-rank test. Our results show that moderate-scale decoder-only models consistently form a top-tier, exhibiting high mean performance and low across-task variance, while domain adaptation via code-focused pre-training might yield only modest improvements. These insights guide model selection for non-code SE workflows and highlight directions for expanding SELU to generative and design-oriented scenarios.

Augmenting the Generality and Performance of Large Language Models for Software Engineering

Software Engineering

Helps computers understand and create software ideas.

13 Jun 2025 0

92%

Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks

Software Engineering

Tests how well AI helps build computer programs.

13 May 2025 1

90%

Benchmarking Large Language Models for Multi-Language Software Vulnerability Detection

Software Engineering

Finds hidden bugs in computer code.

3 Mar 2025 2

View PDF Login to Bookmark

Country of Origin

🇩🇪 Germany

Page Count

16 pages

Evaluating Large Language Models on Non-Code Software Engineering Tasks

Tests computers on understanding software tasks.

Technical Abstract

Augmenting the Generality and Performance of Large Language Models for Software Engineering

Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks

Benchmarking Large Language Models for Multi-Language Software Vulnerability Detection