Hetzner Launches Experimental LLM Inference API Built on Its Own Infrastructure

German cloud provider Hetzner has launched an experimental LLM inference service offering an OpenAI-compatible API hosted on its own hardware. The service is currently in an exploratory phase with no billing, no SLA, and no production guarantees, as Hetzner aims to gauge user interest and system scalability. The only available model at launch is Qwen3.6-35B-A3B-FP8, a 35-billion-parameter Mixture-of-Experts model supporting text, images, and a 262K context window. Early tests conducted on July 23, 2026 showed a median time to first token of 153ms and around 224 output tokens per second, though these figures reflect a single-client snapshot rather than any guaranteed performance. Users can access the API by generating a token via Hetzner's Experiments dashboard and pointing any OpenAI-compatible client at Hetzner's base URL.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in