MAI-Transcribe-2 from Microsoft and Grok Voice Transcribe 2.0 from SpaceXAI are the best value picks. Both get about 2 words wrong in every 100, and both cost $0.10 an hour of audio.

If you want to run a model on your own computer for free, pick Qwen3-ASR 1.7B.

Scores and prices were collected on October 1 and 2, 2026.

The short answer

Speech-to-text means software that listens to audio and writes down the words. Here is what to pick for each need:

  • Lowest price with top accuracy: MAI-Transcribe-2. It costs $0.10 an hour until December 31, 2026, and gets 2.0 words wrong in 100.

  • Same price, other company: Grok Voice Transcribe 2.0. It costs $0.10 an hour and gets 2.3 words wrong in 100.

  • Steady price, no end date: Scribe v2 from ElevenLabs. It costs $0.22 an hour and gets 2.2 words wrong in 100.

  • Fastest: Nova-3 from Deepgram. It works through audio about 556 times faster than real time. It costs $0.0043 a minute.

  • Free on your own computer: Qwen3-ASR 1.7B. It is free to download and covers 52 languages.

If you only need clean English recordings turned into text, you can stop here.

Hosted models side by side

A hosted model runs on the company's computers. You send the audio and pay for each hour or minute. The speech, image and video page lists every model LLM Waves tracks.

The error rate below is the share of words a model gets wrong. Lower is better.

A rate of 2.0% means 2 wrong words in every 100.

Model

Company

Words wrong per 100

Speed factor

Price as the company lists it

Cost for 100 hours

MAI-Transcribe-2

Microsoft

2.0

373.9

$0.10 an hour (until 2026-12-31)

$10.00

Scribe v2

ElevenLabs

2.2

87.5

$0.22 an hour

$22.00

Grok Voice Transcribe 2.0

SpaceXAI

2.3

173.4

$0.10 an hour

$10.00

GPT Transcribe

OpenAI

3.3

44.2

$0.0045 a minute

$27.00

Voxtral Mini Transcribe 2

Mistral

3.6

79.9

$0.003 a minute

$18.00

Universal

AssemblyAI

3.8

120.9

$0.15 an hour

$15.00

gpt-4o-transcribe

OpenAI

4.0

36.1

$0.006 a minute

$36.00

Chirp 3

Google

4.3

28.5

$0.016 a minute

$96.00

Nova-3

Deepgram

5.2

555.9

$0.0043 a minute

$25.80

Table note: words wrong and speed factor come from the Artificial Analysis speech to text page, collected 2026-10-01. Speed factor means how many seconds of audio the model writes out in one second, so higher is faster.

Prices come from each company's own pricing page; AssemblyAI lists Universal under the name Universal-2. The cost for 100 hours is worked out in the next section.

MAI-Transcribe-2 gets 2.0 words wrong in 100 and costs $10.00 for 100 hours. Chirp 3 from Google gets 4.3 wrong and costs $96.00 for the same 100 hours. Nova-3 is the fastest by far, at a speed factor of 555.9.

Where these numbers come from

The error rates and speed scores for hosted models come from Artificial Analysis, a company that tests speech models on the same audio. The error rates for free models come from the Open ASR Leaderboard, a public scoreboard run on Hugging Face.

LLM Waves Research did not run these models itself. Every price comes from the company that sells the model, collected on October 1 and 2, 2026. Each source is listed on the methodology page.

The two scoreboards use different audio. A score from one cannot be compared with a score from the other. This comparison ranks on words wrong first, then on price, then on speed.

How the cost for 100 hours is worked out

Each company lists its price differently. Some charge by the hour and some by the minute. To compare them, convert every price to the cost of 100 hours of audio.

  • Price per hour: multiply by 100. MAI-Transcribe-2 costs $0.10 an hour, times 100 hours, is $10.00.

  • Price per minute: multiply by 60 to get one hour, then by 100. Nova-3 costs $0.0043 a minute, times 60, is $0.258 an hour. Times 100 hours, that is $25.80.

Chirp 3 costs $0.016 a minute. Times 60 is $0.96 an hour, and times 100 is $96.00. That makes Chirp 3 9.6 times the cost of MAI-Transcribe-2.

Extra features cost more on top. AssemblyAI and others charge extra for things like telling speakers apart.

The cheapest accurate picks

MAI-Transcribe-2 and Grok Voice Transcribe 2.0 share the lowest paid price in the table.

Microsoft released MAI-Transcribe-2 on September 3, 2026, as a public preview. A public preview means the model is open to everyone but is still being tested.

The Microsoft announcement sets the price at $0.10 an hour through December 31, 2026. It does not say what the price will be after that. The post says 60 languages in one place and 43 in another. The exact count is unclear.

The model also labels who is speaking and gives each word a start and end time.

SpaceXAI lists Grok Voice Transcribe 2.0 at $0.10 an hour for batch work in its model documentation. Batch work means you send a whole recording and wait for the text, instead of getting words live.

The steady price pick

ElevenLabs lists Scribe v2 at $0.22 an hour on every plan. Scribe v2 gets 2.2 words wrong in 100. That is between the two cheapest picks.

The price has no end date on the ElevenLabs page. That matters if you plan a budget past December 31, 2026, when the MAI-Transcribe-2 offer ends.

The fastest pick

Nova-3 from Deepgram has the highest speed factor in the table, at 555.9. That means one second of work writes out about 556 seconds of audio.

The Deepgram pricing page lists Nova-3 at $0.0043 a minute for one language on pay-as-you-go. Pay as you go means you pay only for what you use, with no plan.

For many languages at once, the price is $0.0052 a minute.

Nova-3 gets 5.2 words wrong in 100. That is the highest error rate in the table. Pick it when speed matters more than every word being right.

Free models you run yourself

An open weight model is one you can download and run on your own computer. You pay nothing to the maker, but you need a strong graphics card and someone to keep it running.

Model

Words wrong per 100

Speed score

Languages

Licence

Qwen3-ASR 1.7B

4.95

819.96

52

Apache 2.0

Canary Qwen 2.5B

5.23

867.09

1

CC BY 4.0

Cohere Transcribe

5.40

906.56

14

Apache 2.0

Parakeet TDT 0.6B v2

5.48

6,024.67

1

CC BY 4.0

Table note: all figures come from the Open ASR Leaderboard with paid models hidden, last updated 2026-10-02.

The speed score shows how many times faster than real time the model works. A licence is the set of rules for using the model. CC BY 4.0 means you must credit the maker. Apache 2.0 does not ask for that in the same way.

Qwen3-ASR 1.7B makes the fewest mistakes of the free models and covers the most languages.

Parakeet TDT 0.6B v2 is about 7 times faster than Qwen3-ASR 1.7B, but it only works in 1 language.

The speed gap works out like this. 6,024.67 divided by 819.96 is about 7.3. The error gap is 5.48 minus 4.95, which is 0.53 words in 100.

A model with a lower score that is not ranked

The Artificial Analysis page shows one model with an even lower error rate: Fun-Realtime-ASR-preview from Alibaba Cloud, at 1.7 words wrong in 100.

The page lists its price as $0.00, and the word preview means it is still being tested. It is not clear what it will cost once the test ends, so it is left out of the picks above.

What we did not measure

This article did not run a latency test, a throughput test, or an output quality test against any model discussed. It does not cover live, word by word transcription.

It does not test noisy or accented audio beyond what each scoreboard uses. It also does not check how often a model makes up words during silence.

Language coverage for most hosted models is not part of either scoreboard.

Which model for which reader

  • Developers trying a model for the first time: Universal from AssemblyAI. It costs $0.15 an hour, and the AssemblyAI pricing page offers up to 185 hours of recorded audio free for new accounts.

    It gets 3.8 words wrong in 100, so it is the wrong pick if transcripts go out without a check.

  • Startups watching cost: MAI-Transcribe-2 at $0.10 an hour. It is wrong for you if you need a price fixed past December 31, 2026.

  • Big companies: Scribe v2 at $0.22 an hour, with a price that has no end date and 2.2 words wrong in 100.

  • Large amounts of audio: MAI-Transcribe-2 or Grok Voice Transcribe 2.0. At 10,000 hours, $0.10 an hour comes to $1,000. Chirp 3 at $0.96 an hour comes to $9,600.

  • Speed first: Nova-3, with a speed factor of 555.9.

  • Running it yourself: Qwen3-ASR 1.7B, free under Apache 2.0, 52 languages.

  • English only and running it yourself: Parakeet TDT 0.6B v2, the fastest free model on the board.

  • Languages other than English: not enough data. Neither scoreboard here ranks hosted models by language.

What is the most accurate speech to text model?

Among models with a stated price, MAI-Transcribe-2 is the most accurate, at 2.0 words wrong in 100.

Fun-Realtime-ASR-preview from Alibaba Cloud scores lower, at 1.7, but it is still in testing and its future price is not clear.

What is the cheapest speech to text API?

MAI-Transcribe-2 and Grok Voice Transcribe 2.0 are the cheapest in this comparison, both at $0.10 an hour. An API is the way your app sends audio to the company and gets text back.

The MAI-Transcribe-2 price lasts until December 31, 2026.

Is Whisper still the best free model?

No. The original Whisper large v3 from OpenAI is not in the top 12 free models on the Open ASR Leaderboard, updated October 2, 2026.

A version tuned by TheStageAI ranks 8th, at 5.42 words wrong in 100. Qwen3-ASR 1.7B holds the top spot, at 4.95.

Check two models on your own audio

Pick two models: the most accurate one in your language, and the cheapest one that is fast enough for you. Run 10 minutes of your own audio through both.

Then count mistakes in names, numbers, and dates first, since those hurt most. The scoreboards narrow the list, but your own audio picks the winner.