GPT-4o, Claude 3.5 Sonnet, and Llama 3 Tested for Smart Contract Security Auditing
AI researcher Saranyo Deyasi benchmarked three large language models — OpenAI's GPT-4o, Anthropic's Claude 3.5 Sonnet, and Meta's Llama 3 70B — on their ability to detect security vulnerabilities in a Solidity smart contract. The test focused on identifying reentrancy and integer overflow flaws, evaluating each model on detection accuracy, JSON output compliance, response latency, and hallucination rate. Claude 3.5 Sonnet emerged as the overall winner, correctly identifying both vulnerabilities with zero hallucinations, though it was slightly slower than GPT-4o. GPT-4o delivered the fastest structured output but incorrectly flagged a non-existent issue as a critical vulnerability. Llama 3 70B, run locally via Ollama, offered the lowest latency and no API costs but missed one vulnerability and failed to return clean JSON output.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in