Open Method Uses Public Model Files to Test 'Built From Scratch' LLM Claims
A technique called Model DNA allows independent researchers to assess whether a large language model was genuinely trained from scratch or derived from an existing open-weight base, using only publicly available files such as config.json, tokenizer data, and embedding weights. The method combines three signals — architectural configuration, tokenizer vocabulary overlap, and embedding-space similarity via Linear CKA — to place a model on a lineage spectrum. It gained prominence around mid-2026 when 'self-developed' claims by AI labs began facing public scrutiny, including a widely shared Zhihu thread that questioned several such announcements. The approach has been applied to major Korean AI developers including LG, NAVER, Kakao, SKT, and others, though its authors stress it produces a lineage label rather than an accusation of wrongdoing. Key limitations include sensitivity to similarity thresholds, a gray zone around continued pretraining, and the fact that the analysis is restricted to embedding layers alone.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in