Whistle: Speech to Text in 16.9 MB

69 points | by gmays an hour ago

12 comments

  • kamranjon 2 minutes ago

    Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.

  • INTPenis 21 minutes ago

    I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.

    I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.

    But every single sound he makes with his mouth ends up on the page too.

      ComputerGuru 15 minutes ago

      Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.

      Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...

  • joewhale 18 minutes ago

    I initially read this as whistle to text, which would be way cooler.

  • aidotguru a minute ago

    eager to see if working in android phones

  • andy_ppp 26 minutes ago

    Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!

    I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!

  • mrkn1 8 minutes ago

    love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages.

  • armcat 10 minutes ago

    Those are insane benchmarks at this size. Well done!

  • tecleandor 26 minutes ago

    Spanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent...

      kaoD 24 minutes ago

      Spanish from where? Here (Castilian Spanish) it seemed to work fine.

  • saturn8601 21 minutes ago

    Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.