OpenAI Drops Faster Speech API, Yet Still Loses to Alibaba
The king of generative AI tried refreshing its audio transcription arsenal with slash-and-burn pricing, but benchmark charts show OpenAI is still eating dust behind competitors.
OpenAI rolled out two brand-new speech-to-text models via its API aimed at developers needing fast text conversion. The flagship batch processor, GPT Transcribe, tackles uploaded audio files and corporate archives, while GPT Live Transcribe handles streaming voice interfaces and real-time audio streams with minimal latency.
The offline version crunches recorded audio at 34 times faster than real-time playback while accepting custom vocabulary lists to prevent corporate jargon and proper nouns from turning into gibberish. Meanwhile, the live variant focuses on stripping out delay during phone calls and video conferences without tripping over multi-language inputs.
Benchmark data from Artificial Analysis reveals that GPT Transcribe hits a 3.3% word error rate, marking a slight upgrade over its predecessor. However, that score leaves OpenAI lagging behind rivals, as Alibaba leads the pack at a razor-sharp 1.7% error rate, followed by ElevenLabs at 2.2%, Microsoft at 2.4%, and Gemini tieing Mistral at 2.8%.
To offset the accuracy gap, API pricing was slashed down to $0.0045 per minute for batch transcriptions and $0.017 per minute for live streaming.
Throwing cheap prices at mediocre accuracy seems to be the current strategy when Chinese tech giants and specialized startups are quietly outperforming the market leader in basic listening comprehension.
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.