Day 2: The model I want is the one that's boring everywhere
Kaggle Benchmarking Challenge. Previously: Day 0, the benchmark · Day 1, most of my bugs looked like model behaviour The benchmark asks one question twice: can the model do the job, and does it know when it can't? There are 200 invented items in four everyday shapes: route, classify, judge and ground. In each shape, one item in five can only be answered with ESCALATE. Every model gets two numbers that are never merged: a task score, and a false-confidence rate (how often it answered anyway when the right reply was ESCALATE).
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in