AirLLM is an inference-focused package that lets very large open models run with far less GPU memory by loading model layers or experts on demand instead of keeping the whole model in memory.
Sign in to read the manager-framed brief.
Sign inAI-explained · grounded in each repo's README