Dev Team Built an AI Estimation Benchmark, Then Their Own Tool Failed It
A software development team set out to build an AI-powered IT project cost estimation tool, but first designed a rigorous public benchmark to test whether such tools actually work. Using nine real-world datasets and two input channels — structured project attributes and natural-language requirements — they tested multiple model classes against standard baselines and human expert estimates. Their pre-registered criteria required models to achieve at least 55% of estimates within 25% of actual costs, among other thresholds, across out-of-sample evaluations with strict leakage controls. No model class they tested met the benchmark, leading them to conclude the performance threshold was unreachable on the available public data. The team published the benchmark under an MIT license, with full reproducibility and a transparent git history serving as pre-registration of their methodology.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in