Large Language Models
Jul 24, 2026
Hetzner explores LLM inference with experimental API for testing purposes
Jul 24, 2026
AI Summary
Hetzner is conducting experiments with a new API for large language model (LLM) inference, currently featuring a single model, Qwen/Qwen3.6-35B-A3B-FP8. The initiative aims to assess user interest and system capabilities without any production guarantees or billing.
- Hetzner is testing an LLM inference API that is compatible with OpenAI's standards.
- The current model available for testing is Qwen/Qwen3.6-35B-A3B-FP8, which has 35 billion parameters and accepts both text and images.
- The API is in an experimental phase, lacking billing, service level agreements (SLAs), and production guarantees.
- Users can create an API token through the Experiments dashboard and integrate it with their applications using standard OpenAI client libraries.
- The initiative is aimed at understanding user demand, system scalability, and feature requirements.
- Hetzner's approach may leverage its efficient hardware acquisition and operational capabilities to potentially enter the low-margin inference market.
- The current GPU offerings from Hetzner may not support larger models, raising questions about future hardware investments for scaling the service.
- The experiment is seen as a way to utilize spare GPU capacity and generate revenue, but its long-term viability depends on further developments in hardware and model offerings.
llminferencehetznerai researchtechnology