Skip to main content

IBM Launches High-Speed Granite Speech 5.0 Turbo CTC Models

25 AUGUST 2026·2 MIN READ·1 SOURCE·Official source

IBM has released two new 470-million-parameter English speech recognition models in its Granite Speech family. The models prioritize unprecedented processing speeds, capable of transcribing over 3.5 hours of audio in a single second.

IBM Launches High-Speed Granite Speech 5.0 Turbo CTC Models

Key takeaways · 3

  • 01

    Models process speech at over 12,600 RTFx on NVIDIA H200 GPUs.

  • 02

    The Apache 2.0 version scores a 5.00% aggregate Word Error Rate.

  • 03

    The models feature a compact 470-million-parameter architecture.

Model Capabilities

IBM released two new compact, 470-million-parameter English speech recognition models in the Granite Speech family. [1] These models achieve speeds of over 12,600 RTFx on an NVIDIA H200 GPU. [1] Using batched inference, the systems can transcribe more than 3.5 hours of speech in one second. [1] A streaming speech recognition WebGPU demo is available for Chrome and Edge browsers. [1]

Licensing and Performance

The granite-speech-5.0-470m-turboctc model uses a smaller training dataset and carries an Apache 2.0 license. [1] This model scores a 5.00% aggregate Word Error Rate. [1] The granite-speech-5.0-470m-turboctc-nc version includes additional training data and is released under a CC-BY-NC-SA-4.0 license. [1] The noncommercial model achieves a 4.85% aggregate Word Error Rate, showing an advantage on the SPGI Speech test but a disadvantage on the chunked Earnings22 test. [1]

What it means

The focus on speed combined with low parameter count positions the Granite Speech 5.0 Turbo CTC models as highly efficient transcription tools. By offering both Apache 2.0 and noncommercial variants, IBM provides options tailored to enterprise integration versus experimental research. The models compete effectively in terms of Word Error Rate while prioritizing sheer throughput on hardware like the NVIDIA H200, matching leading open-source alternatives. What the sources don't address: How the models handle complex acoustic environments or languages beyond English.

The release provides extremely fast, low-parameter speech transcription models suitable for high-volume audio processing. The Apache 2.0 version allows broad commercial integration with highly competitive accuracy.

Why it matters
Daily session

Turn this story into practical AI skill after launch.

Get the release link for daily sessions built around your role and industry.

Join the waitlist

How this developed

  1. 25 August 2026

    IBM Launches High-Speed Granite Speech 5.0 Turbo CTC Models

  2. 25 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.