LLM Trading Bots Beat Benchmark, But Pre-Set Rules Declare Result Inconclusive
A developer ran a two-month paper-trading experiment using LLM agents managing a simulated $100 cryptocurrency portfolio across two seasons, with no real money ever at stake. In Season 2, all four AI agents — using Claude Opus and GPT via Codex in autonomous and structured configurations — outperformed a passive 50/50 WETH/USDC benchmark by 3 to 10 percentage points at the 30-day mark. However, pre-committed decision rules required consistent outperformance across two 30-day windows with a minimum 2 percentage point net edge, a threshold the agents failed to meet within the frozen evaluation window. Inside that designated window, both primary agents posted negative returns and underperformed cash and the passive portfolio, triggering a formal 'STOP' verdict under the pre-registered rules. The author argues the 'inconclusive' outcome is the most valuable finding, as it demonstrates how pre-registration prevents selectively interpreting favorable data as proof of genuine trading edge.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in