Listening for Cetaceans: How AI is Advancing Passive Acoustic Monitoring

Cetacean populations face multiple anthropogenic threats: collisions with ships, noise pollution, bycatch in fisheries, and habitat degradation. Monitoring their presence effectively is crucial. Protecting them requires knowing where they are and when. Visual surveys can not solve this at scale.

On the other hand, passive acoustic monitoring (PAM) provides reliable and non-invasive tracking. Hydrophones can record and track marine mammals continuously for months or years, not just hours. However, this creates a massive data explosion which is very hard to analyse. The challenge is to develop classification tools that save experts a considerable amount of time, in order to build reliable and structured databases, the foundation for an AI capable of analyzing all the recordings.

In fact, the ocean never sounds exactly the same twice: depending on the region, the circadian rhythm, tides and other biotic and abiotic parameters, each recording has its own unique acoustic signature. Cetacean vocalizations represent only a small part of this dense soundscape, where they mingle with maritime traffic, geophony, and all other sounds. An AI model trained on data from a single region quickly loses reliability elsewhere: to work everywhere, it must learn from annotated examples that cover this diversity of sites, seasons, and times of day. However, it is impossible to annotate all the recordings by hand, which is why labeling tools are useful, as they allow experts to focus their efforts on the most informative segments. These annotations must have biological significance: simply detecting a peak in energy is not enough; it only becomes usable data once the species, the type of vocalization, and the context have been identified.

First, we define what we are recording, at several levels: the presence or absence of a signal, species, type of vocalization (click, whistle, song, coda), and sometimes behavioral context. Certain highly stereotyped vocalizations serve as ideal references. In blue whales (Balaenoptera musculus), “A-calls” and “B-calls” are low-frequency tonal signals that have long been described in the literature and are emitted in regular sequences, making them easily identifiable on a spectrogram. Humpback whale (Megaptera novaeangliae) songs exhibit a hierarchical structure: sound units are assembled into phrases, which are repeated to form themes. This allows for annotation at multiple scales, from the individual unit to the entire song. Each site, each season, and each population has its own unique soundscape. This well-documented variability is what gives the datasets their robustness and enables the models to apply to any ecosystem.

In practice, the work begins with viewing the data as a spectrogram, on which the expert highlights the signals in both time and frequency. Annotation can be performed using a labeling tool. At Ezako, we naturally use our Upalgo Sound Labeling tool. Upalgo Labeling speeds up the process through pre-annotation by a model with expert correction, active learning to prioritize the most informative examples, and semi-supervised learning to leverage unlabeled data. The reliability of the labels is based on best practices: a written annotation protocol, a predefined taxonomy, double annotation on a sample, and the measurement of inter-annotator agreement. This method enables the rapid annotation of large datasets and results in a structured and reliable database on which to train AI models for sound detection. These models will then facilitate species identification, with applications in conservation and ecology, scientific research, and fisheries management. In this way, every hour of recorded data becomes actionable biological data that benefits cetaceans.