Skip to content
Introducing Klang Pianissimo

Pianissimo.

Swedish speech-to-text at 3,600× realtime.

Klang Pianissimo is an open speech-to-text model for Swedish, fine-tuned from NVIDIA’s Parakeet v3. It achieves a 4.5% word error rate on Common Voice and transcribes an hour of audio in one second.

Try Pianissimo

Try transcribing Selma Lagerlöf, or try the model on your own audio.

Connecting to the model…

Selma Lagerlöf

Silvergruvan

0:0033:36
Or try your own audio

Drop a file onto the player. Up to 90 minutes. Audio is sent to Klang for transcription.

Small model.
Big difference.

Pianissimo achieves word error rates comparable to KB-Whisper large, at 64 times the speed.

Speed and accuracy in one picture

Faster to the right. Fewer errors toward the top.

Scroll sideways to see the full chart.

Mattias Fält’s model comparison. Faster transcription to the right, fewer word errors toward the top. Pianissimo extends the frontier to 2,500× realtime with word error rates comparable to KB-Whisper large.
Read our research

60 hours of audio.
In one minute.

1 second

to transcribe an hour of audio

Audio transcribed in one minute of processing
One tile = one hour of audio

Get Pianissimo.

Pianissimo is an open model, built for sovereignty. Run it on your own infrastructure or use it through our partner Berget AI. You choose where your audio is processed.

Hugging Face

Run on your own infrastructure.

Download the weights, adapt the model and build on it. You choose the hardware and where your audio is processed.

You choose where the model runs.

European AI needs models that others can build on. Open weights let you run Pianissimo locally, choose a European provider and switch as your needs change.

Built something with Pianissimo? Tell us about it at [email protected].

Klang

AI for conversations that matter.

Klang makes your conversations searchable and usable - with transcriptions, summaries, and insights. Built in Europe, with trust at its core.