Researcher Tests 5 AI Systems on Architecture Choices 10 Times, Finds Patterns and Pitfalls
A developer ran a controlled experiment testing five AI systems — Claude, Gemini, DeepSeek, Yandex Alice, and GPT — across 10 rounds using an abstract information-processing prompt. Each model had to independently select five architectural dimensions most fundamental to a general-purpose system, with item order and wording varied across rounds. One dimension — the set of basic computational operations available to a system — was chosen by all five models in 47 out of 49 runs, even when repositioned or rephrased. However, the author concluded this likely reflects a conceptual advantage built into the question itself rather than independent AI consensus on a universal principle. A seemingly stable four-item 'universal core' that emerged in Round 9 also collapsed in Round 10 after descriptions were rewritten, highlighting how sensitive results were to prompt framing.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in