Training LLMs on GitHub Commit History Could Teach AI Real Engineering Reasoning
Current large language models are trained on billions of lines of finished code from GitHub, but they only learn the final output rather than the decision-making process behind it. A developer argues this is a fundamental gap, comparing it to teaching an architect solely through photos of completed buildings without explaining the reasoning behind design choices. The proposed solution is to train models on full commit histories of major open-source projects like Linux, PostgreSQL, and Rust, so they can learn how and why code evolved over time. Such training could help AI understand refactoring, error handling evolution, and architectural trade-offs rather than simply mimicking clean code patterns. The author acknowledges that most commit data is noisy but suggests focusing on high-quality, well-maintained repositories to extract the most meaningful engineering lessons.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in