Correction to our measurements

14 July 2026

We have found and corrected two errors in how we measure. They affected the results we published up to today, and they affected them in a way that made our rankings less reliable. Here is what was wrong.

Error 1: Correct Scandinavian text was recorded as the wrong language

Answers written in correct Norwegian, Swedish or Danish were in many cases recorded as the model having failed to answer in that language. The error did not hit models evenly: it was worst on short answers, and therefore skewed between models.

The consequence is that our language results measured something other than what we said they measured. Some models were marked down for language that was perfectly fine.

Error 2: The strongest models were not allowed to finish answering

Models that reason through a task before answering had their answers cut short. They were then measured as though they had failed the task — despite never being given the chance to complete it.

This hit several of the strongest models on the market. Worse: some of them dropped out of our selection on the basis of measurements they were never allowed to produce, and so had no way back in.

What we have done

Both errors are fixed.

We store the original answers from every model, which meant we could recompute the entire history from source. Measurements destroyed by the second error have been discarded — we do not let them count against the models, because the error was ours.

The models that were wrongly dropped from the selection have been put back into testing.

We have also introduced a stricter assessment of language quality, which measures how well the language is actually written — not merely whether the answer is in the right language.

Earlier weekly reports have been withdrawn

The weekly reports we published up to today were built on numbers we now know were wrong. They have therefore been withdrawn in full. We are not leaving them up with retroactive corrections, because they cannot be honestly reconstructed.

Why we are publishing this

hvilkenAI is a measurement service. Our value is that the numbers are right. When they are not, the only thing worth anything is that we say so.

We have no affiliate agreements, sponsored placements or commercial ties to the AI providers we test. That remains true, and it is precisely why we have no reason to hide this.

See today’s benchmark →