A comprehensive, first-principles repository dedicated to mastering Large Language Models (LLMs) and sequence modeling architectures. This workspace chronicles a complete pedagogical journey from raw character-level recurrent networks and LSTM memory systems to custom Byte Pair Encoding (BPE), multi-head causal self-attention, complete GPT-2 decoder blocks, pretraining loops, and instruction fine-tuning.
- π¬ Executive Overview & Philosophy
- ποΈ Repository Architecture & Directory Blueprint
- πΊοΈ Multi-Stage Pedagogical Roadmap
- π Mathematical & Architectural Deep Dive
- π» Hands-On Implementations & Codebase
- π Theoretical Knowledge Base (
LLM Notes/) - π Landmark Research Papers Library (
Theory/) - π Datasets & Benchmark Corpus Catalog (
Datasets/) - π Offline Developer Documentation (
Pages/) - β‘ Installation, Environment Setup & Quickstart
- π‘ Key Takeaways & Architectural Lessons
- π€ Author & Acknowledgments
Modern natural language processing is driven by Large Language Models built upon the Transformer architecture. While high-level frameworks like Hugging Face allow quick prototyping, true mastery requires building every component from scratch using raw tensor operations.
This repository bridges the gap between abstract mathematical theory and concrete code implementation:
- Low-Level Tokenization: Implementing character-level tokenizers, word-level dictionaries, and Byte Pair Encoding (BPE) merge algorithms.
- Vector Space Representations: Deriving continuous token embeddings, geometric cosine relationships, and learned 1D positional vectors.
- Attention Mathematics: Constructing scaled dot-product attention, causal triangular masks, query-key-value transformations, and multi-head parallelization.
- Decoder Architecture: Assembling Pre-Layer Normalization, Gaussian Error Linear Units (GeLU), two-layer Feed-Forward Networks, and residual skip connections into full GPT-2 blocks.
- Autoregressive Pretraining: Aligning input and target sequences for next-token prediction, calculating cross-entropy loss, and optimizing with AdamW.
- Downstream Adaptation: Transitioning pretrained weights to sequence classification tasks and multi-turn instruction fine-tuning.
LLMsPractice/
βββ CodeLLM.ipynb # Root Transformer encoder-decoder time-series & sequence forecasting
βββ StocksLSTM.ipynb # Root deep LSTM stock market forecasting laboratory
βββ README.md # Master documentation & pedagogical guide (~400 lines)
β
βββ Datasets/ # Raw and curated corpora, financial time-series, & LLM evaluation logs
β βββ AAPL.csv, AMZN.csv, GOOG.csv, MSFT.csv # High-resolution stock OHLCV datasets for sequential modeling
β βββ TSLA.csv, Mobiles.csv # Financial time-series and tabular benchmark datasets
β βββ the-verdict.txt, article.txt # Literary and article text corpora for language pretraining
β βββ gpt_sentences.txt, sentence_start.txt # Synthetic prompt generation and completion datasets
β βββ question_set.txt, llm_ans.txt # Instruction-following question sets and target answers
β βββ llama_answers.csv, mistral_answers.csv # Model inference output comparisons and evaluations
β
βββ LLM Codes/ # Source code implementations & runnable notebooks
β βββ Self/ # Independent experimental builds & end-to-end scripts
β β βββ 1. TinyLLM2_CharLevel.ipynb # Character-level autoregressive text generator
β β βββ 2. MemoryLLM3_Stocks.ipynb # LSTM-based sequential indicator network
β β βββ 3. CodeLLM.ipynb, 6. CodeLLM.ipynb # Transformer decoder sequence models
β β βββ 4. TokenLLM.ipynb, 5. TokenLLM2.ipynb# Subword tokenization pipelines & Gutenberg processing
β β βββ Stocks.py # Production-grade stock data preprocessor & LSTM pipeline
β β βββ tokenllm.py # Gutenberg dataset scraping & tokenization engine
β βββ Vizuara/ # 15-part end-to-end GPT implementation series
β β βββ 01. Byte_Pair_Encoding.ipynb # Custom BPE tokenization from scratch
β β βββ 02. Data_Loader_Input_Output_Pairs.ipynb # Sliding window sequence pair generation
β β βββ 03. Token_Embeddings.ipynb # Token embedding lookup tables
β β βββ 04. Vector_Embedding.ipynb # Vector geometry and embedding operations
β β βββ 05. Position_Embeddings.ipynb # Learned 1D positional encodings
β β βββ 06. LLM_Data_Preprocessing.ipynb # Full text-to-tensor preprocessing pipeline
β β βββ 07. Attention_Mechanism.ipynb # Dot-product attention and similarity weighting
β β βββ 08. Trainable_Self_Attention.ipynb # Query, Key, Value parameter matrices
β β βββ 09. Causal_Attention.ipynb # Lower-triangular masked attention
β β βββ 10. Multi-Head Attention.ipynb # Parallel multi-head attention module
β β βββ 11. LLM Architecture.ipynb # Complete GPT-2 transformer block assembly
β β βββ 12. LLM Pretraining.ipynb # Full pretraining loop with loss optimization
β β βββ 13. Model Weights Loaded.ipynb # Loading official OpenAI GPT-2 weights
β β βββ 14. Architecture Classification finetuning.ipynb # Sequence classification head
β β βββ 15. Instruction fine-tuning dataset prep loading.ipynb # Instruction tuning pipeline
β β βββ gpt_download3.py # Direct weights downloader for GPT-2 models
β βββ SQ/ # Specialized implementations
β βββ decoder_transformers_with_pytorch_and_lightning_v2.ipynb
β
βββ LLM Notes/ # 54 comprehensive Markdown architectural write-ups
β βββ GPT/ # 20 notes detailing GPT evolution & design patterns
β βββ Vizuara/ # 34 granular notes covering every mathematical building block
β
βββ Theory/ # Landmark peer-reviewed NLP & LLM research papers
β βββ Attention is All You Need.pdf # Vaswani et al. (2017)
β βββ Improving Language Understanding...pdf # Radford et al. (GPT-1, 2018)
β βββ Language Models are Unsupervised...pdf # Radford et al. (GPT-2, 2019)
β βββ Language Models are Few-Shot Learners.pdf# Brown et al. (GPT-3, 2020)
β βββ LoRA.pdf # Hu et al. (2021) Parameter-Efficient Fine-Tuning
β βββ Efficient Memory Management...pdf # Kwon et al. (2023) PagedAttention
β βββ Gaussian Error Linear Units (GeLU).pdf # Hendrycks & Gimpel (2016)
β βββ Instruction Tuning With Loss Over...pdf # Modern instruction alignment research
β βββ Measuring Massive Multitask...pdf # Hendrycks et al. (2020) MMLU Benchmark
β βββ Neural Machine Translation By...pdf # Bahdanau et al. (2014) Additive Attention
β βββ Visualizing the Loss Landscape...pdf # Li et al. (2018) Loss surfaces & skip connections
β βββ A New Algorithm for Data Compression.pdf # Classic BPE data compression foundations
β
βββ Pages/ # Archived official PyTorch docs & visualizers for offline study
βββ Datasets & DataLoaders, Module, Sequential, Softmax, Dropout, Embedding docs
βββ torch.Tensor.masked_fill_, torch.tril, torch.triu, torch.utils.data docs
βββ register_buffer vs register_parameter documentation
βββ Visualizing A Neural Machine Translation Model (Jay Alammar)
| Stage | Model Paradigm | Implemented Project | Core Mathematical & Architectural Concepts |
|---|---|---|---|
| Stage 1 | Character-Level RNN | TinyLLM |
Character vocabulary, one-hot vectors, hidden state recurrence |
| Stage 2 | Word-Level Embeddings | SmaLLM |
Word tokenization, continuous dense lookup tables, vocabulary frequency thresholding, OOV tokens |
| Stage 3 | Recurrent Memory (LSTM) |
MemoryLLM / StocksLSTM
|
Forget, input, and output gates ( |
| Stage 4 | Attention & Transformer Decoder |
AttentionLLM / CodeLLM
|
Scaled dot-product attention, causal mask ( |
| Stage 5 | Full GPT-2 Pretraining |
Vizuara Series (01β13) |
Pre-LayerNorm, GeLU FFN, weight tying, sliding window data loaders, checkpoint loading |
| Stage 6 | Fine-Tuning & Alignment |
Vizuara Series (14β15) |
Sequence classification heads, instruction tuning, prompt templates, LoRA low-rank adaptation |
ββββββββββββββββββββββββββββββββββββββββββ
β Raw Text Sequence β
βββββββββββββββββββββ¬βββββββββββββββββββββ
βΌ
ββββββββββββββββββββββββββββββββββββββββββ
β Byte Pair Encoding Tokenizer β
βββββββββββββββββββββ¬βββββββββββββββββββββ
βΌ
ββββββββββββββββββββββββββββββββββββββββββ
β Token Embedding + Position Embedding β
βββββββββββββββββββββ¬βββββββββββββββββββββ
βΌ
βββββββββββββββββββββββββ
β Dropout β
βββββββββββββ¬ββββββββββββ
βΌ
ββββββββββββββββββββββββββΊ [Transformer Block] βββββββββββββββββββββββββββ
β β β
β ββββββββββββββ΄ββββββββββββ β
β βΌ β (Residual Skip) β
β LayerNorm 1 β β
β β β β
β βΌ β β
β Masked Multi-Head Attention β β
β β β β
β βΌ βΌ β
β Dropout βββΊ (+) Add Connection β
β β β
β ββββββββββββββ΄ββββββββββββ β
β βΌ β (Residual Skip) β
β LayerNorm 2 β β
β β β β
β βΌ β β
β Feed-Forward (Linear+GeLU) β β
β β β β
β βΌ βΌ β
β Dropout βββΊ (+) Add Connection β
β β β
ββββββββββββββββββββββββ (Repeat N Times) ββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββ
β Final LayerNorm β
βββββββββββββ¬ββββββββββββ
βΌ
βββββββββββββββββββββββββ
β Linear Output Head β
βββββββββββββ¬ββββββββββββ
βΌ
βββββββββββββββββββββββββ
β Logits & Probabilitiesβ
βββββββββββββββββββββββββ
-
Character-Level: Minimal vocabulary (
$V \approx 60\text{--}100$ ), zero out-of-vocabulary (OOV) tokens, but requires extremely long context windows to capture semantic concepts. -
Word-Level: Explicit word semantics, but vocabulary scales unsustainably (
$V > 50{,}000$ ), causing heavy OOV failures on unseen words. -
Byte Pair Encoding (BPE): A subword tokenization algorithm that iteratively merges the most frequent adjacent character or byte pairs in the corpus.
- Generates subwords like
["trans", "former", "##s"]. - Balances vocabulary compactness (
$V = 50{,}257$ in GPT-2) with the ability to represent any arbitrary unicode string without OOV errors.
- Generates subwords like
-
Token Embeddings (
$W_e$ ): Given a sequence of token IDs$\mathbf{x} = [x_1, x_2, \dots, x_T]$ , each index is mapped to a continuous vector: $$\mathbf{e}t = W_e[x_t] \in \mathbb{R}^{d{\text{model}}}$$ -
Positional Embeddings (
$W_p$ ): Because self-attention is permutation-invariant, positional encodings must be injected:$$\mathbf{z}_t = \mathbf{e}_t + \mathbf{p}_t$$ -
Learned Positional Embeddings (GPT-2): Directly optimizes an embedding matrix
$W_p \in \mathbb{R}^{T_{\text{max}} \times d_{\text{model}}}$ . -
Sinusoidal Positional Embeddings (Vaswani et al.):
$$PE_{(pos, 2i)} = \sin\left(\frac{pos}{10000^{2i/d}}\right), \quad PE_{(pos, 2i+1)} = \cos\left(\frac{pos}{10000^{2i/d}}\right)$$
- Bahdanau Additive Attention: Computes alignment score using a feed-forward layer: $e_{ij} = \mathbf{v}a^T \tanh(W_a \mathbf{s}{i-1} + U_a \mathbf{h}_j)$.
-
Scaled Dot-Product Attention: Projects input vectors into Queries (
$Q$ ), Keys ($K$ ), and Values ($V$ ):$$Q = XW_Q, \quad K = XW_K, \quad V = XW_V$$ $$\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}} + M\right)V$$ -
Causal Attention Mask (
$M$ ): An upper triangular matrix preventing tokens from attending to subsequent tokens: $$M_{ij} = \begin{cases} 0 & \text{if } i \ge j \ -\infty & \text{if } i < j \end{cases}$$ -
Multi-Head Attention (MHA): Divides
$d_{\text{model}}$ into$h$ heads of dimension$d_k = d_{\text{model}} / h$ :$$\text{MultiHead}(Q, K, V) = \text{Concat}(\text{head}_1, \dots, \text{head}_h)W^O$$
-
Pre-Layer Normalization: Normalizing inputs before each sub-layer stabilizes forward activations and backpropagated gradients:
$$\text{LN}(\mathbf{x}) = \frac{\mathbf{x} - \mu}{\sqrt{\sigma^2 + \epsilon}} \odot \gamma + \beta$$ -
Feed-Forward Network (FFN): Two linear transformations with a non-linear activation expand the hidden dimensionality by
$4\times$ :$$\text{FFN}(\mathbf{x}) = \text{GeLU}(\mathbf{x}W_1 + \mathbf{b}_1)W_2 + \mathbf{b}_2$$ -
Gaussian Error Linear Unit (GeLU):
$$\text{GeLU}(x) = x \cdot \Phi(x) \approx 0.5x \left(1 + \tanh\left(\sqrt{\frac{2}{\pi}}\left(x + 0.044715x^3\right)\right)\right)$$ -
Residual Skip Connections:
$\mathbf{x}^{(l+1)} = \mathbf{x}^{(l)} + \text{SubLayer}(\text{LN}(\mathbf{x}^{(l)}))$ , mitigating vanishing gradient problems across deep stacks. - GPT-2 Family Architecture Specifications:
| Model Variant | Parameters | Layers ( |
Hidden Dimension ( |
Attention Heads ( |
Head Dimension ( |
|---|---|---|---|---|---|
| GPT-2 Small | 124M | 12 | 768 | 12 | 64 |
| GPT-2 Medium | 355M | 24 | 1024 | 16 | 64 |
| GPT-2 Large | 774M | 36 | 1280 | 20 | 64 |
| GPT-2 XL | 1558M | 48 | 1600 | 25 | 64 |
-
Causal Next-Token Prediction: For an input sequence of length
$T$ , the targets are shifted by 1 position:$$\text{Input: } [x_1, x_2, \dots, x_{T-1}] \longrightarrow \text{Target: } [x_2, x_3, \dots, x_T]$$ - Cross-Entropy Loss: $$\mathcal{L} = -\frac{1}{T-1} \sum_{t=1}^{T-1} \log P(x_{t+1} \mid x_1, \dots, x_t) = -\frac{1}{T-1} \sum_{t=1}^{T-1} \left( \hat{z}{t, x{t+1}} - \log \sum_{v=1}^V \exp(\hat{z}_{t, v}) \right)$$
-
Perplexity Metric:
$\text{PPL} = \exp(\mathcal{L})$ , representing the effective branching factor during next-token selection. - AdamW Optimization: Decoupled weight decay regularization to preserve generalization without corrupting adaptive gradient moment estimates.
-
Greedy Decoding:
$x_{t+1} = \arg\max_i (z_i)$ β deterministic, prone to repetitive loops. -
Temperature Scaling (
$T$ ): Modulates logit sharpness before Softmax:$$P(i) = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}$$ -
Top-$K$ Sampling: Truncates candidate logits to the top
$K$ most probable tokens before renormalizing. -
Top-$P$ (Nucleus) Sampling: Selects the smallest token subset whose cumulative probability exceeds threshold
$P$ .
1. TinyLLM2_CharLevel.ipynb: Character-level neural language model with custom sequence generation.2. MemoryLLM3_Stocks.ipynb: Deep LSTM memory network predicting price trends with technical indicators.3. CodeLLM.ipynb&6. CodeLLM.ipynb: Transformer decoder experiments with sequence-to-sequence targets.4. TokenLLM.ipynb&5. TokenLLM2.ipynb: Automated tokenizer token-ID encoding and decoding flows.Stocks.py: Full standalone stock data preprocessor, feature pipeline (RSI, MACD, Bollinger Bands), and LSTM model.tokenllm.py: Project Gutenberg novel downloader, sentence boundary segmenter, and Hugging Face tokenizer bridge.
-
01. Byte_Pair_Encoding.ipynb: Subword vocabulary extraction and merge rule derivation. -
02. Data_Loader_Input_Output_Pairs.ipynb: PyTorchDatasetandDataLoadersliding context windows. -
03. Token_Embeddings.ipynb: Matrix mapping from token IDs to dense vectors. -
04. Vector_Embedding.ipynb: Vector space properties, cosine similarities, and geometric intuition. -
05. Position_Embeddings.ipynb: Implementation of 1D absolute learned position encodings. -
06. LLM_Data_Preprocessing.ipynb: Complete preprocessing pipeline from raw text to tensor batches. -
07. Attention_Mechanism.ipynb: Raw dot product similarity and attention weight distributions. -
08. Trainable_Self_Attention.ipynb: Query ($W_Q$ ), Key ($W_K$ ), and Value ($W_V$ ) parameter matrices. -
09. Causal_Attention.ipynb: Upper triangular masking usingtorch.trilandmasked_fill_. -
10. Multi-Head Attention.ipynb: Efficient batched multi-head splitting, computation, and recombination. -
11. LLM Architecture.ipynb: Complete GPT-2 block assembly with Pre-LN, GeLU FFN, and residuals. -
12. LLM Pretraining.ipynb: Full pretraining loop with loss tracking, target alignment, and generation check. -
13. Model Weights Loaded.ipynb: Transferring official OpenAI GPT-2 checkpoints into custom PyTorch modules. -
14. Architecture Classification finetuning.ipynb: Replacing the output head for sentiment and text classification. -
15. Instruction fine-tuning dataset prep loading.ipynb: Prompt-response formatting and instruction alignment. -
gpt_download3.py: Direct weights downloader for GPT-2 models (124M,355M,774M,1558M).
decoder_transformers_with_pytorch_and_lightning_v2.ipynb: Modular decoder-only transformer built using PyTorch Lightning with distributed training support.
CodeLLM.ipynb: Comprehensive Transformer modeling notebook for multi-variate sequence forecasting.StocksLSTM.ipynb: In-depth time-series experimentation comparing Recurrent architectures with modern attention.
A structured collection of 20 detailed study guides covering architectural decisions:
LLM01 TinyLLM.md-LLM04 SmaLLM.md: Transitioning from character models to vocabulary-backed word models.LLM05 MemoryLLM.md&LLM10 Stage3 vs Stage4.md: In-depth trade-off analysis comparing LSTMs and Transformers.LLM06 CodeLLM.md-LLM09 Training.md: Pretrained model economics, GPU memory profiles, and training loops.LLM13 Pos Embed.md-LLM16 Handle OOV.md: Input embedding strategies, vector math, and handling unknown tokens.LLM17 Causal Attention Code.md-LLM20 GPT2-Small Block.md: Masked attention mechanics, LayerNorm placement, and the complete 12-layer GPT-2 topology.
34 modular markdown files dissecting every sub-component:
- Concepts: GenAI fundamentals, Pre-training vs Fine-tuning, Encoder-Decoder vs Decoder-only, BERT vs GPT.
- Data & Embeddings: BPE tokenization algorithms, sliding window dataloaders, vector spaces, and positional encodings.
- Attention Evolution: RNN bottlenecks, Bahdanau additive attention, scaled dot-product attention, causal masks, and multi-head parallelization.
- Neural Sub-layers: GELU activation derivations, feed-forward dimension expansion, residual additions, and Pre-LN stability.
- Training & Decoding: Cross-entropy calculation over shifts, perplexity metrics, temperature scaling, top-$K$, and nucleus top-$P$ sampling.
The Theory/ directory houses the core research literature that established modern natural language processing:
| Paper Title | Key Contribution to LLM Architecture |
|---|---|
| Attention Is All You Need (Vaswani et al., 2017) | Introduced the Transformer; replaced recurrence with multi-head self-attention. |
| Improving Language Understanding (GPT-1) (Radford et al., 2018) | Established generative pre-training + discriminative fine-tuning paradigm. |
| Language Models are Unsupervised Multitask Learners (GPT-2) (Radford et al., 2019) | Demonstrated zero-shot task transfer across diverse text distributions. |
| Language Models are Few-Shot Learners (GPT-3) (Brown et al., 2020) | Scaled autoregressive models to 175B parameters; in-context few-shot learning. |
| LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., 2021) | Freezes base weights and injects trainable low-rank rank-decomposition matrices. |
| Efficient Memory Management (PagedAttention) (Kwon et al., 2023) | Non-contiguous KV-cache memory allocation for high-throughput LLM serving. |
| Gaussian Error Linear Units (GELUs) (Hendrycks & Gimpel, 2016) | Probabilistic continuous non-linearity adopted by GPT, BERT, and modern transformers. |
| Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau et al., 2014) | First introduction of soft attention mechanisms in seq2seq translation models. |
| Measuring Massive Multitask Language Understanding (MMLU) (Hendrycks et al., 2020) | Benchmark standard for evaluating world knowledge and multi-task reasoning. |
| Visualizing the Loss Landscape of Neural Nets (Li et al., 2018) | Filter normalization techniques explaining why skip connections smooth loss surfaces. |
| Instruction Tuning With Loss Over Instructions (2023) | Masking instruction prompt tokens during loss computation to improve response alignment. |
| A New Algorithm for Data Compression (1994) | Byte-pair encoding compression algorithms foundational to subword tokenizers. |
- Financial Time-Series (
AAPL.csv,AMZN.csv,GOOG.csv,MSFT.csv,TSLA.csv): Historical price and volume data used to benchmark sequence memory in LSTMs vs causal attention models. - Pretraining Corpora (
the-verdict.txt,article.txt): Clean text corpuses used for vocabulary construction, BPE training, and next-token prediction loops. - Synthetic Prompt Datasets (
gpt_sentences.txt,mistral_phi2_sentences.txt,sentence_start.txt): Evaluation prompts for sampling analysis. - Model Output Evaluations (
llama_answers.csv,mistral_answers.csv,llm_answers.csv): Model response benchmarks for comparative inference studies. - Tabular Datasets (
student_depression_dataset.csv,internet_usage.csv,Mobiles.csv): Multi-modal tabular benchmarks for classification fine-tuning.
Local HTML snapshots preserving critical PyTorch API references and pedagogical guides for offline research:
- PyTorch Core APIs:
nn.Module,nn.Sequential,nn.Embedding,nn.Dropout,nn.Softmax,torch.utils.data.DataLoader. - Tensor Operations:
torch.tril(lower triangular mask),torch.triu, andtorch.Tensor.masked_fill_. - Advanced PyTorch Design:
register_buffervsregister_parameterlifecycle and state dict persistence. - Visual Reference: Jay Alammar's Visualizing A Neural Machine Translation Model (Seq2Seq with Attention).
git clone https://github.com/AvrodeepPal/LLMsPractice.git
cd LLMsPracticepython3 -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activatepip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121 # Or CPU wheel
pip install tensorflow transformers datasets tiktoken pytorch-lightning pandas numpy matplotlib seaborn tqdm scikit-learn- Apple Silicon (M1/M2/M3/M4): Automatically utilizes PyTorch
mpsbackend (torch.device("mps")). - NVIDIA GPUs: Configured for CUDA execution with mixed-precision training (
torch.cuda.amp.autocast).
-
Attention Eliminates Sequential Bottlenecks: Unlike LSTMs whose memory degrades with sequence length, self-attention maintains direct
$O(1)$ computational paths between all tokens. -
Causal Masking is Essential for Generative Decoders: Setting future attention logits to
$-\infty$ ensures the model strictly learns$P(w_t \mid w_1, \dots, w_{t-1})$ . - Pre-LN Stabilizes Deep Training: Placing Layer Normalization before the multi-head attention and feed-forward sub-layers enables training deep networks without gradient explosion.
- Tokenization Directly Dictates Model Capability: Subword tokenization via BPE balances vocabulary size and context window efficiency while avoiding out-of-vocabulary failures.
- Decoupled Weight Decay Matters: Using AdamW prevents regularized weights from distorting adaptive learning rate moments, significantly improving generalization.
Avrodeep Pal
- Passionate about Deep Learning, Natural Language Processing, and Generative AI Architectures.
- Focused on building production-grade machine learning pipelines and demystifying LLMs from first principles.
- GitHub: @AvrodeepPal
- Collaboration: Open issues, discussions, or pull requests are warmly welcomed!
β If you find this repository helpful in your LLM learning journey, feel free to star the repo and share it!