How to Size PTU Deployments for GPT-5.6 Luna on Microsoft Foundry
Microsoft Foundry's Provisioned Throughput Units (PTUs) offer dedicated, fixed processing capacity for model deployments, providing tenant-exclusive resources and latency SLAs unlike standard pay-as-you-go plans. PTU sizing requires calculating Normalized TPM by combining effective input and weighted output token volumes, then dividing by a model-specific Input TPM per PTU constant. For GPT-5.6 Luna on a Global Provisioned deployment, that constant is 30,000 Input TPM per PTU, with a minimum allocation of 15 PTUs. A sample workload of 1,000 peak RPM with 1,200-token prompts and 200-token responses requires 80 PTUs without caching, dropping to 60 PTUs when a 50% prompt cache rate is applied. PTU quota is allocated per subscription and region, meaning capacity in one region or deployment type cannot be shared with another.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in