Skip to content

About

Deep learning project for spoken digit classification using CNN and RNN/LSTM models on audio spectrograms. Developed for Aprendizaje Automático II (UNR).

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Spoken Digit Classification with Deep Learning

Aprendizaje Automático II
Tecnicatura Universitaria en Inteligencia Artificial (Universidad Nacional de Rosario)
Máximo Alva, María Sol Aranda
2025

This project explores spoken digit classification using deep learning models trained on audio recordings of digits from 0 to 9.

The goal is to process raw audio signals, transform them into spectrogram representations, and compare different neural network architectures for the classification task.


Overview

The project uses the TensorFlow Spoken Digit dataset, which contains 2,500 audio recordings of spoken digits.

The complete workflow includes:

  • Audio dataset loading and exploration
  • Audio preprocessing
  • Waveform normalization
  • Padding and trimming audio samples to a fixed length
  • Conversion of audio waveforms into spectrograms using Short-Time Fourier Transform (STFT)
  • Training and evaluation of different deep learning architectures
  • Performance visualization using training curves and confusion matrices
  • Evaluation on an additional test set containing new audio recordings

Models

Two different deep learning approaches were implemented and compared.

Convolutional Neural Network (CNN)

The first approach treats spectrograms as image-like representations of audio.

The model includes:

  • Input resizing
  • Data normalization
  • Convolutional layers
  • Max pooling
  • Dropout
  • Fully connected layers

This architecture learns spatial patterns from the time-frequency representation of the audio signals.

Recurrent Neural Network (RNN / LSTM)

The second approach processes the spectrogram representation as a sequence.

The model includes:

  • Input resizing
  • Data normalization
  • Reshaping spectrograms into sequential data
  • Recurrent layers based on LSTM
  • Dense output layers for digit classification

This approach allows the model to learn temporal relationships within the audio representation.


Audio Preprocessing

Each audio sample goes through several preprocessing steps before being used for training:

  1. Extra dimensions are removed from the waveform.
  2. Audio samples are trimmed or padded to a fixed length.
  3. Waveform amplitudes are normalized.
  4. The waveform is converted into a spectrogram using STFT.
  5. Spectrograms are used as input for the neural networks.

This representation makes it possible to analyze how the frequency content of each spoken digit changes over time.


Evaluation

The models are evaluated using several classification and visualization techniques, including:

  • Accuracy
  • Precision
  • Recall
  • F1-score
  • Classification reports
  • Confusion matrices
  • Training and validation loss curves
  • Training and validation accuracy curves

In addition to the original dataset split, both models are evaluated on a separate set of newly recorded audio samples to analyze their ability to generalize to unseen speakers and recordings.

Detailed results, analysis, and conclusions are available in the notebook.


Technologies

  • Python
  • TensorFlow
  • Keras
  • NumPy
  • Matplotlib
  • Seaborn
  • Scikit-learn
  • Jupyter Notebook

Dataset

The project uses the Spoken Digit dataset available through TensorFlow Datasets.

The dataset contains audio recordings of digits from 0 to 9, spoken by multiple speakers.


Project Details

The complete implementation, experiments, visualizations, model architectures, evaluation results, and conclusions are documented in:

📓 audio_mnist.ipynb

About

Deep learning project for spoken digit classification using CNN and RNN/LSTM models on audio spectrograms. Developed for Aprendizaje Automático II (UNR).

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages