Benchmarking AI ML Applications
Generative artificial intelligence (AI) and machine learning (ML) are no longer confined to coding assistants that help developers write software faster — increasingly, they are the functionality being delivered. Recommendation engines, computer vision, fraud detection, forecasting and conversational AI are becoming ordinary line items in development portfolios.
In this short report we look at what it takes to benchmark these AI/ML application projects themselves. We bring together the latest external evidence on AI project outcomes with what an ISBSG-style, size-based benchmark of these projects would need to measure. We show why the functional sizing community’s own standards are only just starting to catch up.
