Hetzner Inference
AssessPlatforms
Hetzner's experimental OpenAI-compatible LLM inference API.
Why it's here
Placed in Assess: 2 article(s) of evidence from 1 source(s), led by product launches, with 2 in the last 30 days. Confidence 31%.
Evidence (2)
- 7Hacker News·8/3/2026open_sourceAirLLM Runs 70B Models on a Single 4GB GPU
AirLLM is an open-source inference approach that claims it can run a 70-billion-parameter language model on a single 4GB GPU by loading model layers on demand. The project drew attention on Hacker News because it lowers hardware requirements for large-model inference, though practical performance and tradeoffs depend on the implementation.
- 5Hacker News·7/24/2026product_launchHetzner tests OpenAI-compatible LLM inference
Hetzner has launched an experimental LLM inference service on its own infrastructure, exposed through an OpenAI-compatible API. The service currently offers only one model, Qwen/Qwen3.6-35B-A3B-FP8, and Hetzner says the goal is to learn demand, scaling behavior, feature priorities, and load characteristics before any production release.