Technology designed for accessibility has often been created without enough input from the people who ultimately depend on it. A new development from Google DeepMind is attempting to change that approach by putting Deaf people at the center of a major breakthrough in American Sign Language technology.
The company has introduced SL2T, an artificial intelligence model capable of translating American Sign Language into written text. The technology is being integrated into Gboard and Live Transcribe on Google’s Pixel 11 devices, giving ASL users another way to communicate with their phones.
A New Era for American Sign Language Technology
Translating ASL is significantly more complicated than simply recognizing hand movements.
American Sign Language is a complete language with its own grammar and structure. Meaning can be communicated through handshapes, movement, facial expressions, body position and the use of physical space.
Google’s new system is designed around that reality. Instead of reducing signs to simplified labels before translating them, SL2T uses computer vision to track movements across the signer’s face, body and hands.
The system tracks 130 points, creating a detailed representation of how a person is signing.
Built With Deaf People, Not Just for Them
One of the most important aspects of the project is the involvement of the Deaf community throughout its development.
The concept came from Sam Sepah, a Deaf Google employee who has spent years working on technology for Deaf and disabled communities. Google also established an AI Sign Language Advisory Committee made up of Deaf organizations and subject-matter experts.
Deaf users were involved in testing and evaluation, helping the development team understand how the technology performs in real communication situations.
This user-centered approach represents an important shift in accessibility technology.
How Google’s Sign-to-Text AI Works
SL2T was trained using more than 100,000 hours of sign-language data covering more than 50 sign languages, with a significant portion focused on ASL.
The technology is designed to recognize different signing situations, including one-handed signing and left-handed users. This is particularly useful when someone is holding a phone while communicating.
The system also takes privacy into consideration. Rather than sending raw camera footage to Google’s servers, the device converts the visual information into movement coordinates before the data is used for translation.
That technical approach allows the system to analyze complex movements while limiting the amount of visual information that leaves the device.