How AI could help us talk to animals

notes.

How AI could help us talk to animals

Source: How AI could help us talk to animals, Vox, 9:06, uploaded 2024-07-31, category Science, playlist index 845.

In the 1980s, Joyce Poole noticed that an elephant sometimes called one member of her family and received an answer from that individual whilst the other elephants ignored her. A second call could draw a response from another elephant. Poole had a useful observation and no way to test whether the calls were addressed to particular animals. Decades later, she worked with Mickey Pardo, who recorded calls alongside the identity of the caller, the likely receiver, and the surrounding behaviour.

The pair turned each recording into numbers and fed nearly 500 calls into a statistical model. Given the acoustic structure of a new call, the model predicted which elephant would receive it more often than chance would allow. Vox presents this as evidence that African savanna elephants give one another names. The result comes from a modest experiment with a defined question. It does not produce an elephant dictionary. It shows that a pattern which human observers could not detect may sit inside the calls.

The field recording problem

Animal communication research usually joins three kinds of work. Researchers record vocalisations, observe the animals and the context around them, and sometimes play a sound back to measure the response. AI enters each part of this process, although the first difficulty appears in the recordings themselves.

Field audio often contains several animals calling at once. The result resembles the cocktail party problem in human speech recognition, where a listener tries to follow one voice through a room full of conversations. The video uses Deep Karaoke as the human-language example. Researchers trained the model on music whose vocals and instruments had been recorded separately, then on the fully mixed versions, until it could separate new recordings into their component sounds.

Similar methods can isolate an individual call from a recording of macaque monkeys. That gives researchers a cleaner object to study, and it also changes what a playback experiment can do. Generative sound models can learn from many examples of an animal’s recording and produce a new version. A researcher could use such a model to create controlled sounds for a playback rather than choosing from a small archive of calls already captured in the field.

The video calls these supervised learning because people label the training examples. In the elephant study, human observations supplied the information that sat beside the acoustic data. The model found a relation in the combination that direct observation had missed, yet the labels also set the boundary of what the model could find.

Yossi Yovel’s study of Egyptian fruit bats makes the boundary more visible. His statistical model trained on 15,000 vocalisations and identified the caller, the context, the behavioural response, and the animal to whom a call was addressed. The researchers annotated the recordings by hand. Yovel points out that this remains a restriction because human labels carry human assumptions into a bat’s communication. A model can discover patterns beyond our hearing whilst still learning from the categories we already know how to name.

From labels to shapes

This is why some researchers place greater hope in self-supervised models. The video explains them through the example of ChatGPT. A language model can take a large body of unlabeled text, find recurring patterns, and sort words into relationships without a person assigning every sentence to a category first. In this account, every language has a shape that a model can discover.

Aza Raskin, who co-founded the Earth Species Project, describes that shape as a field of relationships among words. Words with related meanings sit near one another, and the directions between them encode other relations. The familiar example is that the relation between man and king resembles the relation between woman and queen. The Earth Species Project visualises this arrangement for the 10,000 most common English words.

The proposed next step comes from work on machine translation. In 2017, researchers found that the shape of one language could be matched to the shape of another. A word such as dog could occupy a roughly corresponding position in both spaces, even when the model had no paired translation for each example. The video treats this as a possible route towards animal communication: a model might find a structure shared by human and nonhuman signals without a human first supplying a phrase-by-phrase equivalent.

Animal communication includes more than sound, which makes that idea harder. Dolphins and elephants recognise themselves in mirrors, a sign of self-awareness in the video’s account. Aza Raskin points to image models such as DALL-E and Midjourney for a possible extension. Text, images, and sound can be represented as related shapes and then aligned, which could let a model search for relations across the different senses an animal uses.

The hope is that the points where those communication spaces overlap might reveal shared experiences. That language belongs to a research programme. The video presents the geometry and the cross-modal models as reasons for hope, not as a working translation system.

Validation and the limits of shared language

Self-supervised learning still needs validation. People refine a model by grading its answers, yet grading becomes difficult when the model describes a form of communication that humans cannot interpret. A fluent output can look meaningful because it fits our expectations. The researchers also warn that people may expect too much overlap between human and nonhuman lives and imagine a conversation before they have established what the signals refer to.

The video shows a cat demonstration in which a human says, “How are you doing, dude?” and an AI turns the sentence into a meow. The resulting exchange sounds like a greeting because the human supplies the interpretation. A meow generated from a human prompt gives evidence about the model’s ability to produce a sound. It gives no evidence that the cat receives the intended question or answers it.

Raskin draws a distinction that keeps the project from becoming a promise of instant translation. Humans communicate through language, a specific behaviour that the video says appears unique to us given current knowledge. Other species communicate through sounds, gestures, movement, chemicals, colour, and other channels. Human language can be one case within a wider field of communication, although the two terms should stay separate while researchers work out what the signals do.

The missing data

The immediate task concerns data collection. Researchers need far more labelled and unlabelled recordings. Raskin describes a database containing close to 10,000 individual calls and calls that figure very small. Around the world, researchers are tagging animals and collecting video, sound, and spatial data for models that need to see behaviour together with the signal.

The database changes the scale of the problem. A model cannot infer a communication system from a few striking recordings, especially when calls vary with the caller, receiver, setting, and response. It also cannot escape the conditions under which researchers gather the data. The animals that are easiest to observe will be overrepresented, and the features people know to record may still shape what later counts as a meaningful pattern.

True interspecies communication remains uncertain. The source’s more immediate claim is narrower: machine learning can expose structure in animal communication that human observation misses. Researchers hope that this work will also change how people value and protect the species around them. The closing reminder is plain. Other animals communicate, care for one another, and appear to hold memories of the past and expectations of the future. Those capacities give them a reason to be here and a right to remain here.

Limits of the source’s evidence

Vox names the elephant, bat, macaque, translation, and sound-separation studies in the video or its description, yet the video does not provide their full methods, sample sizes, controls, or results. The figures and technical explanations above therefore remain claims made by the source unless a linked study is read separately. The model examples show pattern detection and sound generation. They do not establish that an animal understands a human sentence, that a model has found the meaning of a call, or that a shared language exists.

Further reading / references

The video description names these materials:

22 paragraphs1,425 words9,033 characters