Implementation and study of the core concepts behind Neural Networks and Transformer Decoder models.
- Neural Network
- Forward Propagation
- Backpropagation
- Gradient Descent
- Activation Functions
- Loss Functions
- Token Embedding
- Positional Encoding
- Self-Attention
- Masked / Causal Attention
- Multi-Head Attention
- Transformer Decoder
- Feed-Forward Network
- Layer Normalization
- Residual Connections
- Softmax
- Next Token Prediction
The project demonstrates the flow:
Tokenization → Embedding → Positional Encoding → Masked Self-Attention → Feed-Forward Network → Linear Layer → Softmax → Next Token Prediction
Forward Pass → Loss Calculation → Backpropagation → Gradient Calculation → Gradient Descent → Parameter Update