Apple's new SpeechAnalyzer API has been put to the test against Whisper and its predecessor, with surprising results [1]. The benchmark, which included 5,559 standard test utterances, found that Apple's SpeechAnalyzer is the most accurate on-device speech engine, beating Whisper Small and the legacy SFSpeechRecognizer [2]. The results show that SpeechAnalyzer has a word error rate (WER) of 2.12% on clean speech and 4.56% on noisy speech, outperforming Whisper Small and SFSpeechRecognizer [3]. The new API also runs roughly three times faster than Whisper Small, making it an attractive option for developers and users who require fast and accurate speech-to-text transcription.
The benchmark was conducted using the same production code paths and text normalization as Inscribe, a private on-device AI workspace, to ensure accurate and comparable results [4]. The findings have significant implications for developers and users who rely on speech-to-text transcription, particularly those using Apple devices. With the new API, Apple devices offer amazing speech-to-text transcription capabilities, with speeds that are dramatically faster than rival tools like Whisper [5]. However, it's essential to note that the benchmark was conducted using English read speech, and the results may not be applicable to other languages or types of speech [6].
The full benchmark results, including the raw transcripts and methodology, are publicly available for review and validation [7]. In conclusion, Apple's SpeechAnalyzer API has set a new standard for on-device speech recognition, offering significant improvements in accuracy and speed over its predecessor and rival tools like Whisper. As the technology continues to evolve, it will be exciting to see how developers and users leverage this powerful API to create innovative applications and services.

