NOV 06 to 08 | 2025 SÃO PAULO EXPO

Thursday to Saturday
10 AM to 7 PM

Caixa lotteries

Lip-reading equipment helps people with speech difficulties

Scientists develop glasses that can read your lips without you having to emit a single sound

The art of lip reading has fascinated psychologists, computer scientists, and forensic experts. In most cases, experiments involve someone reading another person's lips or an artificial intelligence reading a human's lips via a phone app.

But a different type of experiment is being developed in Laboratory of Intelligent Computer Interfaces for Future Interactions (SciFi) from Cornell University.

A team of scientists has developed a speech recognition system that can identify up to 31 English words. But EchoSpeech, as the system is called, is not an app—it is a pair of seemingly normal glasses.

Agreed described in a report, the glasses are capable of reading the user's own lips and helping those who cannot speak to perform basic tasks, such as unlocking the phone or asking Siri to turn up the TV volume without having to emit a single sound.

The glasses are equipped with two microphones, two speakers, and a microcontroller so small that it practically blends in. Their operation seems like magic, but in reality, the technology used is sonar.

More than a thousand species use sonar to hunt and survive, like whales, which can send out sound pulses that reflect off objects in the water. The sounds return so the animal can process these echoes and build a mental image of its environment, including the size and distance of surrounding objects.

EchoSpeech works in a similar way, except that the system does not rely on distance. It tracks how sound waves (inaudible to the human ear) travel across the face and how they hit various moving parts of the face.

The process can be summarized into four main steps:

1. The small speakers (located on one side of the glasses) emit sound waves.

2. As the user pronounces various words, the sound waves travel across the face and strike various “articulators,” such as the lips, jaw, and cheeks.

3. The microphones (located on the other side of the glasses) collect these sound waves

4. The microcontroller processes them along with the device the glasses are paired with

But how does the system know how to attribute a given word to a given facial movement? To do this, researchers used a form of artificial intelligence known as a deep learning algorithm, which teaches computers to process data in the same way the human brain does.

“If you train enough, you can look at someone's mouth, without hearing any sound, and infer the content of their speech,” says the lead author of the study, Ruidong Zhang.

The team used a similar approach, except that instead of another human inferring the content of their speech, the team used a pre-trained artificial intelligence model to recognize certain words and match them with a corresponding “echo profile” of a person's face.

BEYOND ENGLISH

For now, EchoSpeech has the vocabulary of a child. It can recognize all 10 numerical digits. It can understand directions such as “up,” “down,” “left,” and “right”—which, according to Zhang, could be used to draw lines in computer-aided design software. And it can activate voice assistants like Alexa, Google, or Siri, or connect to other Bluetooth-enabled devices.

ECHOSPEECH TRACKS HOW SOUND WAVES TRAVEL ACROSS THE FACE AND HOW THEY HIT VARIOUS MOVING PARTS OF THE FACE.

Zhang says that increasing the system's vocabulary to 100 or 200 words should not pose any specific challenge with current AI. But any number above that would require a more advanced AI model, which would piggyback on existing speech recognition research.

This is an important step, considering the team wants to pair the system with a voice synthesizer and help people who cannot speak vocalize sound more naturally and efficiently.

For now, EchoSpeech is a prototype with an intriguing concept and tremendous potential to help people with disabilities, but the team does not expect it to be ready for use in the next five years. And it will only work for the English language.

“The difficulty is that every language has different sounds,” says François Guimbretière, co-author of the study, who is French. Different sounds can mean different facial movements. But it also depends on the types of languages the AI model is trained on.

“There is an effort to apply other languages as well, so that not all technology is geared toward English,” says Guimbretière.

SOURCE:
Fast Company Brasil

Share this Content

Comments

LAST DAYS FOR
THE END OF 1st Batch!

Days
Hours
Minutes
Seconds
Open chat
1
Reatech
Hello 👋
Can we help you?