Writing or working in a Nordic language?
Start with the language score and hvilkenAI score. A high total score is useful, but language quality matters most when the answer will be published or shared.
Over 350 AI models are evaluated every morning — the strongest performers for Nordic language quality, speed and value are presented here.
Read weekly reports → Compare models → ChatGPT vs Claude → Gemini vs ChatGPT →
The benchmark is built for practical decisions, not only rankings. Start with your task, then compare language quality, price, speed and stability.
Start with the language score and hvilkenAI score. A high total score is useful, but language quality matters most when the answer will be published or shared.
Look at price per million tokens and value score. A lower-cost model can be best when you run many tasks or build a product.
Look at tokens per second and stability. A fast model feels better in chat, support, automations and tools where waiting time matters.
For Nordic text and simple tasks, mid-range models often perform very well. Premium is most useful for complex reasoning, long documents and high-precision work.
| # | Model | Tier | language quality | Instr | Score | Price/1M |
|---|---|---|---|---|---|---|
| 1 | Anthropic: Claude Haiku 4.5 anthropic | Mid-range | 9.7 | 10.0 | 8.3 | $1.00 ≈€0.87 |
| 2 | Mistral: Mistral Small 3 mistralai | Budget | 9.3 | 10.0 | 8.4 | $0.05 ≈€0.04 |
| 3 | Google: Gemini 3.5 Flash Lite google | Mid-range | 9.3 | 10.0 | 8.4 | $0.30 ≈€0.26 |
| 4 | Mistral Large 2407 mistralai | Mid-range | 9.3 | 9.3 | 8.3 | $2.00 ≈€2 |
| 5 | Meta: Llama 3.2 1B Instruct meta-llama | Budget | 9.3 | 7.3 | 6.0 | $0.03 ≈€0.03 |
| 6 | Google: Gemma 3 4B google | Budget | 9.0 | 8.7 | 7.9 | $0.05 ≈€0.04 |
| 7 | Cohere: Command R+ (08-2024) cohere | Premium | 8.7 | 10.0 | 7.6 | $2.50 ≈€2 |
| 8 | IBM: Granite 4.0 Micro ibm-granite | Budget | 8.3 | 10.0 | 6.6 | $0.02 ≈€0.02 |
| 9 | Perplexity: Sonar Pro perplexity | Premium | 8.3 | 10.0 | 7.3 | $3.00 ≈€3 |
| 10 | Claude Opus 5 (Fast) anthropic | Premium | 8.3 | 9.3 | 7.2 | $10.00 ≈€9 |
| 11 | OpenAI: GPT-5.6 Luna Pro openai | Mid-range | 8.0 | 9.3 | 7.5 | $0.10 ≈€0.09 |
| 12 | OpenAI: GPT-5.6 Sol Pro openai | Premium | 7.7 | 10.0 | 7.3 | $5.00 ≈€4 |
We evaluate over 350 AI models and present the best results every morning. The score combines language understanding, quality, speed, price and stability. The exact weighting is proprietary.
Read more about the methodology →We evaluate over 350 available models and present the best results from the latest benchmark. The selection updates dynamically as the market changes.
The language score shows how well the model answers in the right language with strong language quality. It is shown as a normalised number.
The instruction score shows how well the model follows the task it receives. We show the result as a simple, normalised score.
The score combines language understanding, quality, speed, price and stability. The exact weighting is proprietary.
Best value points to models that deliver strong results relative to price. The exact calculation is proprietary.