Fine-Tuned Model Hit 100% Accuracy — Until a Better Benchmark Exposed the Truth
A developer fine-tuned Mistral 7B using LoRA on a personal laptop to detect personal data in log lines and support messages, initially achieving a perfect 100% score on a self-generated test set. The result was misleading because the test data was built from the same templates as the training data, effectively measuring memorisation rather than generalisation. When the benchmark was rebuilt using real public data, the fine-tuned model dropped to 95% accuracy while few-shot prompting collapsed from 94% to just 66%. The experiment — run entirely on an Apple Silicon Mac at zero cost — showed a genuine 29-point performance gap in favour of fine-tuning, reversing the original conclusion. The author highlights that overly easy or template-matched test sets can silently corrupt evaluation results, making benchmark design as critical as model training itself.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in