Fine-tuned AI model for Barbados outperforms base on some benchmarks, fails on others

Researchers fine-tuned a 30-billion-parameter Qwen3 model on Barbados newspaper archives as part of a Caribbean AI buildathon project called Pulse, which aims to build a public-signal intelligence system for the island. Two versions of the model were evaluated against the unmodified base using three distinct benchmarks covering factual recall, radio transcript quality, and TikTok content extraction. The fully trained Version 4 outperformed the base model by 13.3 percentage points on factual recall and nearly doubled quality scores on radio transcription, correctly identifying local proper nouns like Crop Over event names. However, on the TikTok extraction benchmark, V4 performed worse than V3 by generating far more false positives — 23 versus 8 — due to over-emission of observations and misclassification of entity types. The findings highlight how a single benchmark can be misleading, and that a fine-tuned model may genuinely improve in some domains while regressing in others depending on how outputs are structured and scored.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in