Skip to content

Repository files navigation

Transformer Decoder Model, Neural Network & Backpropagation

Implementation and study of the core concepts behind Neural Networks and Transformer Decoder models.

Topics Covered

  • Neural Network
  • Forward Propagation
  • Backpropagation
  • Gradient Descent
  • Activation Functions
  • Loss Functions
  • Token Embedding
  • Positional Encoding
  • Self-Attention
  • Masked / Causal Attention
  • Multi-Head Attention
  • Transformer Decoder
  • Feed-Forward Network
  • Layer Normalization
  • Residual Connections
  • Softmax
  • Next Token Prediction

Transformer Decoder

The project demonstrates the flow:

Tokenization → Embedding → Positional Encoding → Masked Self-Attention → Feed-Forward Network → Linear Layer → Softmax → Next Token Prediction

Learning Process

Forward Pass → Loss Calculation → Backpropagation → Gradient Calculation → Gradient Descent → Parameter Update

About

Neural Network, Backpropagation, and Transformer Decoder

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages