Google's Sign Language AI Reaches Phones

Google DeepMind's new model lets Deaf users sign to their phone anywhere they'd normally type, starting with ASL on the Pixel 11.

Google's Sign Language AI Reaches Phones

For decades, hearing people have enjoyed talking to their phones. Dictation, voice search, chatty assistants that mostly understand you. Meanwhile the world's 200-plus sign languages, and the roughly 70 million Deaf and hard of hearing people who use them, have been left out of that convenience. Google DeepMind wants to change that, and it has just shipped its first attempt to real devices.

What SL2T actually does

The new model is called SL2T, short for sign-language-to-text. It watches someone sign and produces written text, and it now powers two features on the Pixel 11: sign-to-text dictation in Gboard (Google's keyboard) and Live Transcribe (its captioning app). It launches with American Sign Language to English, with more languages and devices promised later, all at no extra cost.

The idea is simple and useful. Just as a hearing person can dictate a message instead of typing, a Deaf user can now sign to their phone anywhere they would normally type. That means signing a web search, drafting a document, asking Gemini to run a task, or replying in a live conversation. Testers reported that signing in ASL felt faster and more natural than typing in English.

Why this is harder than it sounds

You might assume this is just speech-to-text with a camera. It is not. Two things make sign language much trickier. First, sign languages are not "English on the hands." They are full, independent languages with their own grammar and vocabulary, so the job is genuine translation, not a word-for-sign swap. Second, meaning comes from simultaneous movements of the hands, arms, torso, head, and face. Tracking all of that accurately at high frame rates is a demanding computer vision problem.

This is also why earlier gadgets like sign language gloves fell short. They captured hand shapes but missed the rich, whole-body, spatial nature of the language.

How the model works

SL2T was trained on over 100,000 hours of data spanning more than 50 sign languages, with about a quarter of it in ASL. Training across many languages and skill levels together helped the model learn shared structure and beat single-language versions.

On privacy, there is a neat design choice. Rather than sending raw video to a server, an on-device model tracks points on the signer's body, called pose landmarks, and only those geometric coordinates get sent for translation. The original video can be discarded right away.

SL2T also skips "glosses," the word-by-word sign labels that older systems relied on. Glosses miss things like facial expressions and how signers use space, so translating straight from landmarks removes artificial vocabulary limits and lets quality improve as data grows. On the FLEURS-ASL benchmark, SL2T scored 70 BLEURT (a measure of translation quality), which Google says is well above previously reported results.

The team is candid about limits. Errors still crop up with rare signs, fast fingerspelling ("prey" became "grey" in one example), passive constructions, and tense without context. They also worked on practical issues like latency, avoiding false output when no one is signing, supporting the roughly 10% of signers who are left-handed, and handling one-handed signing for when your other hand is holding the phone.

Built with the community

Google stresses this was built with the Deaf community, not just for it. Deaf collaborators shaped the project from the start, including Deaf Googler Sam Sepah, who helped conceive it. An AI Sign Language Advisory Committee of global Deaf organizations and experts helps guide deployment. The team co-authored a joint impact report laying out what the technology can and cannot do, and plans to keep that practice for future releases.

What's next

ASL input on a phone is the starting point, not the finish line. The team says it is working on more sign languages, sign language generation (turning text back into signing), and stronger underlying AI. The goal is parity with spoken and written languages across digital tools.

For now, the honest framing matters most. This is an early 1.0 with real gaps, but it is the first time a sign language model has left the lab and landed in everyday consumer apps. If the accuracy keeps improving, signing to your phone could become as ordinary as talking to it.

Subscribe to BuzzBelow

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe