Putting sign language AI into users’ hands
2026-08-20 · Google DeepMind
Putting Sign Language AI into Users’ Hands
Introduction
Google DeepMind’s Sign Language Team has introduced SL2T, a massively multilingual sign-language-to-text translation model. This marks a significant breakthrough in both quality and generality, bringing sign language AI out of the lab and into consumer products for the first time.
The model currently powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language (ASL) to English. Additional devices and languages will be supported in the future.
Deaf and hard-of-hearing users can now sign to their phones in any situation where they would normally type — searching the web, drafting messages or documents, or asking Gemini to solve queries and perform tasks. In Live Transcribe, users can sign their responses during conversations instead of typing back and forth.
According to testers, signing in ASL is faster, more natural, and more delightful than typing in English.
Why Sign Languages Matter
Sign languages are the primary languages of Deaf communities worldwide and form the cornerstone of Deaf cultural identity. There is significant diversity among deaf individuals in their proficiency across signing, speaking, reading, and writing. Therefore, supporting access across all communication modalities is essential.
Sign language processing offers deaf users benefits comparable to those hearing users gain from spoken language technologies. It also creates new opportunities to bridge the communication gap between Deaf and hearing communities.
Despite this potential for positive social impact, progress in sign language AI has been slow due to both technical complexity and widespread misconceptions about how sign languages function.
Core Challenges in Sign Language AI
Sign language translation presents two fundamental challenges compared to spoken language transcription:
- True Translation Requirement: Sign languages are independent natural languages with their own distinct grammars and lexicons. They cannot be treated as sequential mappings from signs to words in the same language. This requires genuine machine translation rather than simple transcription.
- Complex Visual Understanding: Meaning is conveyed through simultaneous movements of the hands, arms, torso, head, and face. Accurately tracking these fine-grained, whole-body movements at high frame rates is an extremely demanding computer vision task.
Early attempts such as sign language gloves were fundamentally limited because they treated sign languages as “English on the hands,” failing to address the need for sophisticated visual perception and full language translation. SL2T was specifically designed to overcome both challenges.
How SL2T Works
SL2T was developed through a user-centric, culturally informed approach combined with massive data scaling. The model was trained on over 100,000 hours of data spanning more than 50 sign languages, with roughly a quarter of the data in ASL. Joint training across diverse languages, dialects, and proficiency levels enables the model to learn shared underlying structures, outperforming single-language models in experiments.
Privacy-First Design: To protect user privacy, SL2T processes sign language as a sequence of pose landmark locations rather than raw camera footage. An on-device MediaPipe Holistic model extracts these geometric coordinates, which are sent to the server for translation. The original video is discarded immediately.
Direct Translation Architecture: Unlike previous work that relied on intermediate “gloss” annotations, SL2T translates directly from landmark sequences to text. Glosses fail to capture rich, non-linear aspects of sign languages such as non-manual markers and spatial constructions. By removing this artificial bottleneck, translation quality scales directly with data volume.
Performance and Real-World Optimizations
SL2T achieves state-of-the-art results on key benchmarks. On the FLEURS-ASL (sd-test) benchmark for ASL-to-English translation, it reaches a remarkable zero-shot BLEURT score of 70 — significantly higher than any previously reported result.
Beyond academic metrics, the team focused heavily on practical usability:
- Minimizing streaming latency for real-time interaction
- Preventing hallucination when users are not signing
- Ensuring fairness for the approximately 10% of signers who are left-handed
- Improving performance for one-handed signing, which is common when holding a smartphone in the other hand
These optimizations ensure SL2T delivers a reliable and accessible experience in everyday scenarios.
Significance
By delivering SL2T directly to users’ devices, Google DeepMind has taken a major step toward making advanced AI tools equally available to Deaf and hard-of-hearing communities. The technology not only enhances individual productivity and communication but also contributes to greater digital inclusion on a global scale.