Aprendizaje Automático II
Tecnicatura Universitaria en Inteligencia Artificial (Universidad Nacional de Rosario)
Máximo Alva, María Sol Aranda
2025
This project explores spoken digit classification using deep learning models trained on audio recordings of digits from 0 to 9.
The goal is to process raw audio signals, transform them into spectrogram representations, and compare different neural network architectures for the classification task.
The project uses the TensorFlow Spoken Digit dataset, which contains 2,500 audio recordings of spoken digits.
The complete workflow includes:
- Audio dataset loading and exploration
- Audio preprocessing
- Waveform normalization
- Padding and trimming audio samples to a fixed length
- Conversion of audio waveforms into spectrograms using Short-Time Fourier Transform (STFT)
- Training and evaluation of different deep learning architectures
- Performance visualization using training curves and confusion matrices
- Evaluation on an additional test set containing new audio recordings
Two different deep learning approaches were implemented and compared.
The first approach treats spectrograms as image-like representations of audio.
The model includes:
- Input resizing
- Data normalization
- Convolutional layers
- Max pooling
- Dropout
- Fully connected layers
This architecture learns spatial patterns from the time-frequency representation of the audio signals.
The second approach processes the spectrogram representation as a sequence.
The model includes:
- Input resizing
- Data normalization
- Reshaping spectrograms into sequential data
- Recurrent layers based on LSTM
- Dense output layers for digit classification
This approach allows the model to learn temporal relationships within the audio representation.
Each audio sample goes through several preprocessing steps before being used for training:
- Extra dimensions are removed from the waveform.
- Audio samples are trimmed or padded to a fixed length.
- Waveform amplitudes are normalized.
- The waveform is converted into a spectrogram using STFT.
- Spectrograms are used as input for the neural networks.
This representation makes it possible to analyze how the frequency content of each spoken digit changes over time.
The models are evaluated using several classification and visualization techniques, including:
- Accuracy
- Precision
- Recall
- F1-score
- Classification reports
- Confusion matrices
- Training and validation loss curves
- Training and validation accuracy curves
In addition to the original dataset split, both models are evaluated on a separate set of newly recorded audio samples to analyze their ability to generalize to unseen speakers and recordings.
Detailed results, analysis, and conclusions are available in the notebook.
- Python
- TensorFlow
- Keras
- NumPy
- Matplotlib
- Seaborn
- Scikit-learn
- Jupyter Notebook
The project uses the Spoken Digit dataset available through TensorFlow Datasets.
The dataset contains audio recordings of digits from 0 to 9, spoken by multiple speakers.
The complete implementation, experiments, visualizations, model architectures, evaluation results, and conclusions are documented in: