Skip to content
Back to the pressroom
Research

Klang releases an open dataset for Swedish dialects

The Knäck Klang challenge became Klang Dialects: an open dataset with 1,804 recordings for evaluating Swedish speech recognition.

On 14 September 2026, we released Klang Dialects, an open dataset and benchmark for Swedish speech recognition. The recordings came from Knäck Klang, where participants read three sentences in their own dialect. The harder Klang found it to transcribe their speech, the more points they earned.

Swedish dialects are underrepresented in commonly used speech recognition benchmarks. Klang Dialects lets researchers and developers explore how the technology handles different ways of speaking Swedish. Contributions come from across the country, with the most recordings from southern Sweden.

The sv-clean version contains 1,804 recordings with quality-checked reference transcripts. Our research team manually reviewed around 1,000 recordings to make sure the text matches what was actually said. The dataset is freely available on Hugging Face, along with documentation on how to use it.

Get the dataset on Hugging Face

For press & media

Covering Klang?

Find our logos and brand guidelines here. Get in touch for interviews, images, and questions about Klang.

Get in touchGet to know Klang

The Klang brand

Logos, symbols & guidelines

Explore brand resources