Speech Data Project

Exploring the Igbo language through speech data.

A proof-of-concept project for organising and exploring recorded Igbo speech, connecting individual word recordings with examples of those words used in context.

Explore the Dataset

The project brings together recorded Igbo words, contextual sentences and supporting metadata in a structured dataset that can be explored through a simple web interface.

The Data

Recorded words and language in context

The dataset contains recordings of individual Igbo words alongside sentence recordings in which those words occur. Metadata identifies the target word, English meaning and category for each sentence.

Isolated word recordings

Individual recordings provide repeated examples of target words. Multiple recordings of the same word allow the data to capture variation between recordings.

Contextual sentences

Sentence recordings place the target words in natural linguistic context and connect each recording to its corresponding metadata.

Structured metadata

Each sentence recording is associated with its target word, meaning, category and audio file, making the recordings easier to organise and analyse.

Speech research

The structured recordings provide a foundation for future investigation into Igbo speech, pronunciation, tonal distinctions and language technology.

Capturing tonal distinctions

Some of the recordings focus on words whose meanings are distinguished through tonal and orthographic differences. The dataset keeps these forms separate so they can be examined individually.

àkwà bed
àkwá egg
ákwà cloth
ákwá cry / weeping
égbé hawk
ègbè gun

Explore the recordings

Browse the available target words, isolated recordings, contextual sentences and associated metadata.

View Dataset