All relevant open-source models you can run locally with Ollama, LM Studio, or llama.cpp – with memory requirements, context window, license, and German quality. The memory filter shows which models fit your hardware. Whether running locally is worth it at all is covered in the in-depth guide.
The largest openly available model there is: 2.8T parameters, 104B of them active. It needs around 1.5 TB of memory, and there it holds its own against the closed frontier models.
Largest open model
1M context
Native MXFP4
DeepSeek V4-Pro
98
1,600B · 49B active
880 GB
1M
good
MIT
1457
$1.60
DeepSeek's new flagship: 1.6 trillion parameters, three reasoning tiers and a 1M context for large servers.
Frontier class
1M context
MIT licence
GLM-5.2New
97
753B · 40B active
415 GB
1M
good
MIT
—
$0.00
The successor to GLM-5, with five times the context window and an MIT licence, so no usage restrictions at all. Built for long, multi-step tasks.
1M context
MIT licence
Long tasks
Kimi K2.6
96
1,000B · 32B active
560 GB
256K
good
Modified MIT
1460
$0.95
Moonshot's flagship: one of the strongest open models, for large servers only.
Top open weights
Frontier class
MoE-fast
DeepSeek-V3
95
671B · 37B active
400 GB
128K
good
DeepSeek License
1358
$0.26
A top-tier model for dedicated multi-GPU servers.
Frontier class
Server hardware
GLM-5
94
744B · 40B active
410 GB
200K
good
MIT
1458
$0.60
Zhipu's frontier MoE for dedicated servers, freely available under the MIT licence.
Frontier reasoning
MoE-fast
MIT licence
DeepSeek V4-Flash
93
284B · 13B active
155 GB
1M
good
MIT
1436
$0.09
The lean V4 variant: frontier reasoning with a 1M context for workstations.
1M context
MoE-fast
MIT licence
Qwen3 235B-A22B
92
235B · 22B active
140 GB
32K
strong
Apache 2.0
1375
$0.45
A near-frontier MoE model for workstations with plenty of memory.
Near-frontier
MoE-fast
Kimi K2.7 Code
91
1,000B · 32B active
560 GB
256K
good
Modified MIT
—
$0.71
Moonshot's coding-focused variant of Kimi K2.6: for long-horizon software engineering agents.
Coding specialist
Agentic
Very fast (MoE)
Mistral Medium 3.5 128B
88
128B
75 GB
256K
strong
Modified MIT
—
—
Mistral's dense 128B flagship: instruction-following, reasoning and coding in a single model.
Strong all-rounder
256K context
EU model
Nemotron 3 SuperNew
87
120B · 12B active
70 GB
1M
good
NVIDIA Nemotron Open Model
—
—
NVIDIA's open reasoning model, built from Mamba and transformer layers. 12B active parameters keep it fast, and the context window reaches 1M tokens.
1M context
Agentic
Mamba hybrid
Llama 3.3 70B
86
70B
43 GB
128K
good
Llama Community
1318
$0.10
A large model on par with earlier cloud flagships.
Flagship class
Highly capable
gpt-oss 120BNew
86
117B · 5.1B active
80 GB
128K
good
Apache 2.0
1352
$0.04
OpenAI's large open model, sized to fit exactly one 80 GB card. Apache 2.0, with no restrictions on commercial use.
Fits one H100
Tool calling
Native MXFP4
Mistral Small 4 119B
85
119B · 6.5B active
70 GB
256K
strong
Apache 2.0
—
$0.08
An MoE model with 119B total and 6.5B active parameters: a general-purpose model and a reasoning model in one.
Hybrid reasoning
Very fast (MoE)
256K context
Llama 4 Scout
84
109B · 17B active
62 GB
10M
good
Llama 4 Community
1321
$0.10
A massive context window (10M tokens) for entire mountains of documents.
10M context
MoE-fast
Multimodal
Devstral 2 123B
83
123B
74 GB
256K
good
Modified MIT
—
$0.40
Mistral's large coding agent: 123B parameters for demanding software engineering tasks.
Coding specialist
Agentic
256K context
Qwen3.6 27BNew
82
27B
18 GB
256K
strong
Apache 2.0
—
$0.30
A dense 27B model that beats its own 397B predecessor on coding benchmarks. Currently the best pick for a single 24 GB card.
Image & video
262K context
Strong at code
Qwen3-Coder-Next
80
80B · 3B active
52 GB
256K
good
Apache 2.0
—
$0.12
An MoE coding model: 80B-class coding quality at 3B speed.
Coding specialist
Very fast (MoE)
Tool calling
Qwen3 32B
78
32B
21 GB
32K
strong
Apache 2.0
1347
$0.08
The most capable model that still fits on a 24 GB GPU.
Strongest dense model
Reasoning
Gemma 4 31B
77
31B
20 GB
256K
strong
Apache 2.0
1451
—
The Gemma flagship: first-class German, image understanding and a free Apache licence.
Top German
Multimodal
Apache 2.0
QwQ 32BNew
76
32.5B
21 GB
128K
strong
Apache 2.0
1336
—
A reasoning specialist that writes out its intermediate steps. Strong at maths and logic, slower to answer because it thinks first.
Reasoning
Maths & logic
Apache 2.0
Gemma 3 27B
74
27B
18 GB
128K
strong
Gemma Terms
1365
$0.08
One of the best options for German-language tasks.
Top German
Multimodal
Qwen3 30B-A3B
72
30B · 3B active
20 GB
32K
strong
Apache 2.0
1327
$0.12
An MoE model: the quality of a 30B at the speed of a 3B.
Very fast (MoE)
Reasoning
gpt-oss 20BNew
71
21B · 3.6B active
16 GB
128K
good
Apache 2.0
1317
$0.03
OpenAI's open model for consumer hardware. Just 3.6B active parameters keep it fast, and MXFP4 training gets it running from 16 GB.
Tool calling
Reasoning levels
Native MXFP4
Mistral Small 3.1 24B
70
24B
16 GB
128K
strong
Apache 2.0
1303
$0.35
A European model with an excellent speed-to-quality ratio.
EU model
Fast
Tool calling
DiffusionGemma 26B-A4BNew
69
25.2B · 3.8B active
14 GB
256K
strong
Apache 2.0
—
—
Google's first open diffusion LLM: generates text in parallel token blocks instead of one token at a time, over 1,000 tokens/sec at 25B-class quality.
Diffusion LLM
Extremely fast
256K context
Gemma 4 26B-A4B
68
25.2B · 3.8B active
14 GB
256K
strong
Apache 2.0
1438
$0.00
The first MoE model in the Gemma 4 line: 25B-class quality at the speed of a 4B model, with native tool calling.
Very fast (MoE)
Tool calling
256K context
Devstral Small 2 24B
67
24B
16 GB
256K
good
Apache 2.0
—
—
Mistral's coding agent at 24B: specialised in codebase exploration and multi-file edits.
Coding specialist
Agentic
256K context
Qwen3 14B
64
14B
11 GB
32K
strong
Apache 2.0
—
$0.12
The sweet spot: strong, versatile, runs on mid-range GPUs.
Reasoning
Tool calling
Ministral 3 14B
63
14B
11.5 GB
256K
strong
Apache 2.0
—
$0.20
The sweet spot of the Ministral 3 line: strong German, image understanding and 256K context.
Multimodal (image)
256K context
Strong German
Gemma 4 12B
62
12B
9.5 GB
256K
strong
Apache 2.0
—
—
The new Gemma generation: strong German and image understanding, now under Apache 2.0.
Strong German
Multimodal
256K context
Phi-4 14B
60
14B
11 GB
16K
solid
MIT
1256
$0.07
Microsoft's specialist for logic, maths and programming.
Reasoning
Maths & code
Gemma 3 12B
58
12B
9.5 GB
128K
strong
Gemma Terms
1342
$0.05
Outstanding German and image understanding at moderate requirements.
Strong German
Multimodal
GLM-4.1V 9B ThinkingNew
56
10B
8.5 GB
64K
solid
MIT
—
—
A vision-language model with a built-in reasoning mode: matches 72B-class models on image and document understanding.
Vision (image)
Reasoning
MIT licence
Qwen3-VL 8B
54
8B
7.5 GB
256K
good
Apache 2.0
—
$0.12
Sees images and documents: Qwen3 with vision for local OCR and analysis.
Vision (image)
Long context
Tool calling
Ministral 3 8B
53
9B
7.5 GB
256K
strong
Apache 2.0
—
$0.08
A versatile 8B model from Mistral's new Ministral 3 line with image understanding and a huge context window.
Multimodal (image)
256K context
Tool calling
Qwen3 8B
52
8B
7 GB
32K
strong
Apache 2.0
—
$0.12
An excellent all-round model for everyday enterprise use.
Reasoning
Tool calling
Multilingual
EuroLLM 22BNew
48
22.6B
15.5 GB
32K
strong
Apache 2.0
—
—
A larger EuroLLM generation from the UTTER project with 22.6B parameters. Unlike the 9B version and Teuken, it ships with a 32K context window.
35 languages
Made in Europe
32K context
Llama 3.1 8B
46
8B
6.5 GB
128K
solid
Llama Community
—
—
The classic, with a vast ecosystem of specialised variants.
Widely adopted
Many fine-tunes
Gemma 4 E4B
41
8B
4.5 GB
128K
good
Apache 2.0
—
—
A compact multimodal Gemma 4 model: image and audio input at a low memory footprint thanks to Per-Layer Embeddings.
Efficient (PLE)
Image & audio
128K context
Qwen3 4B
38
4B
4 GB
32K
good
Apache 2.0
—
—
A strong entry point that already runs smoothly on small devices.
Great value
Tool calling
Ministral 3 3B
37
3.8B
4.5 GB
256K
good
Apache 2.0
—
$0.10
A compact Mistral model with image understanding and a large context window for low-end hardware.
Multimodal (image)
256K context
Small
Gemma 3 4B
36
4B
4.5 GB
128K
good
Gemma Terms
1303
$0.05
Understands images and offers a large context window.
Multimodal (image)
Long context
EuroLLM 9BNew
34
9.2B
7.5 GB
4K
strong
Apache 2.0
—
—
A European multilingual model from the UTTER project, trained on every official EU language. Held back by the same 4K context as Teuken.
35 languages
Made in Europe
Apache 2.0
Gemma 4 E2B
33
5.1B
3 GB
128K
good
Apache 2.0
—
—
A tiny edge model from the new Gemma 4 generation with image and audio understanding thanks to Per-Layer Embeddings.
Efficient (PLE)
Image & audio
For low-end hardware
Llama 3.2 3B
30
3B
3.5 GB
128K
solid
Llama Community
1167
$0.05
A lightweight model with a large context window.
Long context
Edge-ready
Teuken 7BNew
30
7B
6 GB
4K
strong
Apache 2.0
—
—
Trained by Fraunhofer IAIS in the OpenGPT-X project, centred on German and the other EU languages. A 4K context window and a 2023 knowledge cut-off limit what it can do today.
All 24 EU languages
Made in Germany
Apache 2.0
Qwen3 1.7B
24
1.7B
2.5 GB
32K
good
Apache 2.0
—
—
A compact all-rounder with decent German.
Small
Multilingual
Gemma 3 1B
18
1B
2 GB
32K
good
Gemma Terms
—
—
A tiny model for simple tasks and low-end hardware.
Very small
Blazing fast
No model matches these filters. Raise the memory filter.
LMArena (as of 09/13)
What changed recently
How the numbers come together
Transparency first: these figures are carefully estimated reference values for real-world use, not lab benchmarks. The verification date sits below the table; the icon next to each model links to its official model card.
Memory requirement
The memory value applies to the Q4_K_M quantization (the default in Ollama and LM Studio) including a small context window. Rule of thumb: around 0.55 GB per billion parameters plus some overhead. Models trained quantisation-aware from the start (gpt-oss, Kimi K3) carry the vendor’s own published minimum instead, or the size of the published files where the vendor never published one.
Mixture-of-Experts
For MoE models, the total parameter count determines the memory requirement (all experts sit in RAM), while only the active parameters determine the speed. Hence the speed advantage at the same size.
Strength score
The strength is a rounded reference value (0–100) that increases with model size. It helps when comparing within this list, but it is not an official benchmark score.
Note on the context window: the Qwen3 text models from 4B upwards run natively with 32K tokens and can be extended to up to 128K via YaRN (the 1.7B model stays at 32K). The table shows the native value.
Frequently asked questions
Which local AI model is best for German?
For German-language tasks, the Gemma models are considered particularly strong, followed by Qwen3 and Mistral. The table shows the German quality per model as an assessment.
What does the context window mean?
The context window is the amount of text, in tokens, that a model can process at once. 32K is enough for most tasks; long documents benefit from 128K or more. More context costs additional memory.
Can these models be used commercially?
That depends on the license (see the “License” column). Apache 2.0 and MIT allow free commercial use; licenses such as the Llama Community License or the Gemma Terms come with restrictions you should verify against the original license when in doubt.
What is an MoE model?
In a Mixture-of-Experts model, all parameters sit in memory, but only a fraction of them (the active parameters) compute per token. This makes it as fast as a much smaller model while requiring the memory of the large one.
From model to production system
Run your model locally and Corporate LLM turns it into a production system: RAG, agent system, skills, and connectors. 100% GDPR-compliant.