Modal Leads Serverless GPU Cold Start Race at 1.8s, RunPod Wins on Cost
A 2026 benchmark study compared cold start latencies and per-second costs across major serverless GPU platforms including Modal, RunPod, Replicate, Together AI, and Lambda Labs. Modal recorded the fastest median cold start at 1.8 seconds on an A100 40GB GPU, attributed to its filesystem and memory snapshotting technology. RunPod Serverless emerged as the most cost-efficient option at roughly $2.59 per hour, making it well-suited for batch processing and high-throughput workloads. Replicate posted the slowest true cold start at 6.5 seconds and the highest equivalent hourly rate, while Together AI offered instant responses through pooled inference but without scale-to-zero flexibility. The full benchmark data, hardware configurations, and testing scripts have been published openly on GitHub for independent review.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in