Local speech-to-text models
10 current open-weight models from 8 labs. Latest version of each family only.
Models
ModelLabParamsYour deviceScore
Voxtral Realtime ArabicMistral AI4.43B—VibeVoice ASR StreamingMicrosoft2.81B—Granite Speech 5.0 TurboCTCIBM473M—VibeVoice ASR BitNetMicrosoft323M—MOSS Transcribe DiarizeOpenMOSS909M—MOSS Transcribe PreviewOpenMOSS2.42B—Cohere Transcribe ArabicPreviewCohere2.07B—Nemotron 3.5 ASRNVIDIA638M—Higgs Audio v3 STT v2Boson AI8.91B—MiMo V2.5 ASRXiaomi7.62B—Newest first.
Min memory is the smallest listed option. ● stated by the publisher · ◇ estimate from file size, an 8K-token context and 0.5 GB for the runtime. — means no figure is available.