Software Engineering

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
Avatar
Tanmay Sah
5 views
LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review
Avatar
Ali Hassaan Mughal
238 views
SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces
Avatar
librarian
157 views
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
Avatar
Ali Hassaan Mughal
214 views
AI-Generated Smells: An Analysis of Code and Architecture in LLM and Agent-Driven Development
Avatar
librarian
222 views
Finding Duplicates in 1.1M BDD Steps: cukereuse, a Paraphrase-Robust Static Detector for Cucumber and Gherkin
Avatar
Ali Hassaan Mughal
232 views
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
Avatar
librarian
273 views
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
Avatar
Tzafrir Rehan
301 views
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
Avatar
librarian
281 views
Rethinking Autonomy: Preventing Failures in AI-Driven Software
  Engineering
Avatar
Joydeep
641 views
Are Large Language Models Robust in Understanding Code Against
  Semantics-Preserving Mutations?
Avatar
librarian
708 views
Mutation Testing framework for Machine Learning
Avatar
rsingh80
948 views
Patched RTC: evaluating LLMs for diverse software development tasks
Avatar
Asankhaya Sharma
894 views
Patched MOA: optimizing inference for diverse software development tasks
Avatar
Asankhaya Sharma
1022 views
Pitfalls in Language Models for Code Intelligence: A Taxonomy and Survey
Avatar
Xinyu She
1057 views
Runtime Resolution of Feature Interactions through Adaptive Requirement
  Weakening
Avatar
Simon Chu
1031 views
Demystifying Compiler Unstable Feature Usage and Impacts in the Rust
  Ecosystem
Avatar
Chenghao Li
1004 views
Towards the decentralized coordination of multiple self-adaptive systems
Avatar
Paul-Andrei Dragan
1042 views
Variance of ML-based software fault predictors: are we really improving
  fault prediction?
Avatar
Domenic Bubel
1129 views
Exploring Behaviours of RESTful APIs in an Industrial Setting
Avatar
Stefan Karlsson
1004 views
Evaluating Pre-trained Language Models for Repairing API Misuses
Avatar
Ting Zhang
1000 views
Formal Runtime Error Detection During Development in the Automotive
  Industry
Avatar
Jesko Hecking-Harbusch
1107 views
Exploring Large Language Models for Code Explanation
Avatar
Paheli Bhattacharya
970 views
Leveraging Deep Learning for Abstractive Code Summarization of
  Unofficial Documentation
Avatar
AmirHossein Naghshzan
1102 views
Vision-Based Mobile App GUI Testing: A Survey
Avatar
Shengcheng Yu
1032 views
Using ChatGPT throughout the Software Development Life Cycle by Novice
  Developers
Avatar
Muhammad Waseem
1068 views
Less is More? An Empirical Study on Configuration Issues in Python PyPI
  Ecosystem
Avatar
Yun Peng
1000 views
Unleashing the Power of Clippy in Real-World Rust Projects
Avatar
Chunmiao Li
1104 views
The Effects of Computational Resources on Flaky Tests
Avatar
Denini Silva
974 views
A comprehensible analysis of the efficacy of Ensemble Models for Bug
  Prediction
Avatar
Ingrid Marc¸al
1005 views
Large Language Models for Code Analysis: Do LLMs Really Do Their Job?
Avatar
Chongzhou Fang
1179 views