Last updated: 2026-08-13

AI benchmark: today’s best Nordic-language models

Over 350 AI models are evaluated every morning — the strongest performers for Nordic language quality, speed and value are presented here.

Read weekly reports → Compare models → ChatGPT vs Claude → Gemini vs ChatGPT →

How to read the benchmark

The benchmark is built for practical decisions, not only rankings. Start with your task, then compare language quality, price, speed and stability.

Writing or working in a Nordic language?

Start with the language score and hvilkenAI score. A high total score is useful, but language quality matters most when the answer will be published or shared.

Using AI often or through an API?

Look at price per million tokens and value score. A lower-cost model can be best when you run many tasks or build a product.

Need fast answers?

Look at tokens per second and stability. A fast model feels better in chat, support, automations and tools where waiting time matters.

Best in class

🏆
Highest score
Google: Gemini 3.5 Flash Lite
8.4/10
🇳🇴
Best language quality
Anthropic: Claude Haiku 4.5
9.7/10
Fastest
OpenAI: GPT-5.6 Luna Pro
1477 t/s
💰
Cheapest (score ≥ 3)
Meta: Llama 3.2 1B Instruct
$0.03/1M
📊
Best value
IBM: Granite 4.0 Micro
Value 305.7
🔗
Best orchestrator
Anthropic: Claude Haiku 4.5
Orch 9.7/10

Is premium worth it?

Premium language score
8.3/10
Mid-range language score
9.1/10
Price difference
~6×

For Nordic text and simple tasks, mid-range models often perform very well. Premium is most useful for complex reasoning, long documents and high-precision work.

All results

# Model Tier t/s TTFT language quality Instr Score Orch. Value EU Price/1M
1
Anthropic: Claude Haiku 4.5
anthropic
Stable
Mid-range 148 105 ms 9.7 10.0 8.3 9.7 9.7 🇪🇺 EU $1.00
≈€0.87
2
Mistral: Mistral Small 3
mistralai
Budget 353 50 ms 9.3 10.0 8.4 9.3 161.2 🇪🇺 EU $0.05
≈€0.04
3
Google: Gemini 3.5 Flash Lite
google
Stable
Mid-range 219 64 ms 9.3 10.0 8.4 9.3 31.2 ~EU $0.30
≈€0.26
4
Mistral Large 2407
mistralai
Stable
Mid-range 283 69 ms 9.3 9.3 8.3 8.7 4.6 🇪🇺 EU $2.00
≈€2
5
Meta: Llama 3.2 1B Instruct
meta-llama
Stable
Budget 626 53 ms 9.3 7.3 6.0 6.8 208.3 $0.03
≈€0.03
6
Google: Gemma 3 4B
google
Stable
Budget 324 58 ms 9.0 8.7 7.9 7.8 147.2 ~EU $0.05
≈€0.04
7
Cohere: Command R+ (08-2024)
cohere
Stable
Premium 212 93 ms 8.7 10.0 7.6 8.7 3.7 $2.50
≈€2
8
IBM: Granite 4.0 Micro
ibm-granite
Stable
Budget 232 83 ms 8.3 10.0 6.6 8.3 305.7 $0.02
≈€0.02
9
Perplexity: Sonar Pro
perplexity
Stable
Premium 79 189 ms 8.3 10.0 7.3 8.3 3.0 ~EU $3.00
≈€3
10
Claude Opus 5 (Fast)
anthropic
Stable
Premium 195 99 ms 8.3 9.3 7.2 7.8 0.9 🇪🇺 EU $10.00
≈€9
11
OpenAI: GPT-5.6 Luna Pro
openai
Mid-range 1477 10 ms 8.0 9.3 7.5 7.5 78.8 $0.10
≈€0.09
12
OpenAI: GPT-5.6 Sol Pro
openai
Premium 1003 19 ms 7.7 10.0 7.3 7.7 1.8 $5.00
≈€4

Response time over the last 14 days

How we test

We evaluate over 350 AI models and present the best results every morning. The score combines language understanding, quality, speed, price and stability. The exact weighting is proprietary.

Read more about the methodology →

Frequently asked questions

We evaluate over 350 available models and present the best results from the latest benchmark. The selection updates dynamically as the market changes.

The language score shows how well the model answers in the right language with strong language quality. It is shown as a normalised number.

The instruction score shows how well the model follows the task it receives. We show the result as a simple, normalised score.

The score combines language understanding, quality, speed, price and stability. The exact weighting is proprietary.

Best value points to models that deliver strong results relative to price. The exact calculation is proprietary.