Google DeepMind brings sign-to-text AI to Gboard and Pixel 11
A new multilingual model turns sign-language movements into text for messaging, search and live conversations on supported Google devices.

Google DeepMind is moving sign-language translation from research into consumer software with SL2T, a model that converts signing into streaming text. The technology initially supports American Sign Language to English and powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, with more devices and languages planned.
For Deaf and hard-of-hearing users, the feature is designed to make signing a direct alternative to typing. Users can sign messages, web searches, documents and Gemini requests, while Live Transcribe enables signed replies during conversations without requiring text input.
A different AI problem
SL2T is not a speech-to-text system adapted for hands. Sign languages have their own grammars, and meaning can depend on simultaneous hand, body, facial and head movements. DeepMind says the model was trained on more than 100,000 hours spanning over 50 sign languages, with about one-quarter of that data representing ASL.
The system also takes a privacy-conscious approach. An on-device MediaPipe Holistic model extracts body pose landmarks, and the translation service receives geometric coordinates rather than the original camera footage. SL2T translates those landmarks directly into text instead of relying on intermediate “gloss” labels, which the team says can miss important linguistic structure.
Beyond benchmark performance, the development work focused on practical deployment concerns, including low streaming latency, avoiding hallucinated output when nobody is signing, support for left-handed users and one-handed signing while holding a phone. Google reports a 70 BLEURT zero-shot result on the FLEURS-ASL benchmark, which it describes as a substantial improvement over earlier reported scores.
For AI builders, the launch illustrates how multimodal systems can combine on-device perception, privacy-preserving representations and language translation to address accessibility needs that conventional speech interfaces do not cover.
Source: Google DeepMind Blog
Comments
Log in to join the discussion