WHALE Method Matches Full Fine-Tuning at a Fraction of Compute Cost
Researchers have developed WHALE, a training recipe that alternates between updating model weights and optimizing executable harness code, achieving accuracy comparable to full fine-tuning while using significantly less GPU time and fewer training rollouts. The approach outperforms weight-only, harness-only, and the prior Fast-Slow Training method by 4.15 to 24.38 percentage points in best accuracy. WHALE was tested on tasks including search-based question answering, mathematical reasoning with Python execution, and chess puzzles, consistently matching end-to-end fine-tuning benchmarks. Earlier joint-adaptation methods treated the surrounding harness code as fixed, a limitation WHALE directly addresses by making harness search an active part of the training loop. The study notes open questions around scalability beyond the Qwen 3.5-2B and 4B model sizes and whether the compute advantage holds as harness complexity increases.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in