Trendora

Hetzner Inference

Assess

Platforms

Hetzner's experimental OpenAI-compatible LLM inference API.

Why it's here

Placed in Assess: 2 article(s) of evidence from 1 source(s), led by product launches, with 2 in the last 30 days. Confidence 31%.

Evidence (2)

  • 7Hacker News·8/3/2026open_source
    AirLLM Runs 70B Models on a Single 4GB GPU

    AirLLM is an open-source inference approach that claims it can run a 70-billion-parameter language model on a single 4GB GPU by loading model layers on demand. The project drew attention on Hacker News because it lowers hardware requirements for large-model inference, though practical performance and tradeoffs depend on the implementation.

  • 5Hacker News·7/24/2026product_launch
    Hetzner tests OpenAI-compatible LLM inference

    Hetzner has launched an experimental LLM inference service on its own infrastructure, exposed through an OpenAI-compatible API. The service currently offers only one model, Qwen/Qwen3.6-35B-A3B-FP8, and Hetzner says the goal is to learn demand, scaling behavior, feature priorities, and load characteristics before any production release.