8 complete ML projects covering regression, classification, clustering, deep learning, NLP, and transfer learning.
Each project uses real/standard datasets and is fully runnable.
| # | Project | Model | Dataset | Key Metric |
|---|---|---|---|---|
| 1 | House Price Prediction | Ridge Regression | California Housing | R² ≈ 0.60 |
| 2 | Spam Classifier | Linear SVM | 20 Newsgroups | Accuracy ≈ 98% |
| 3 | Customer Segmentation | K-Means (K=4) | Synthetic Blobs | Silhouette ≈ 0.55 |
| 4 | Image Classification | CNN (PyTorch) | CIFAR-10 | Accuracy ≈ 78% |
| 5 | Sentiment Analysis | Logistic Regression | 20 Newsgroups | F1 ≈ 0.93 |
| 6 | ANN Diabetes Prediction | PyTorch ANN | Pima-like Synthetic | Accuracy ≈ 76% |
| 7 | Transfer Learning | ResNet18 Fine-Tuned | ImageNet → Custom | Feature Extraction |
| 8 | Word Embeddings | Word2Vec + t-SNE | Custom Vocabulary | Cosine Similarity |
ml-projects-collection/
├── 01_house_price_prediction/
│ └── main.py # Linear, Ridge, Lasso comparison
├── 02_spam_classifier/
│ └── main.py # TF-IDF + Naive Bayes + SVM
├── 03_customer_segmentation/
│ └── main.py # K-Means + PCA visualization
├── 04_image_classification/
│ └── main.py # CNN on CIFAR-10 (PyTorch)
├── 05_sentiment_analysis/
│ └── main.py # TF-IDF + Logistic Regression
├── 06_ann_diabetes_prediction/
│ └── main.py # PyTorch ANN with custom DataLoader
├── 07_transfer_learning/
│ └── main.py # ResNet18 fine-tuning vs feature extraction
├── 08_word_embeddings/
│ └── main.py # Word2Vec, cosine similarity, t-SNE
├── requirements.txt
└── README.md
git clone https://github.com/imranalimemon/ml-projects-collection.git
cd ml-projects-collection
pip install -r requirements.txt
# Run any project
python 01_house_price_prediction/main.py
python 02_spam_classifier/main.py
python 03_customer_segmentation/main.py
python 04_image_classification/main.py
python 05_sentiment_analysis/main.py
python 06_ann_diabetes_prediction/main.py
python 07_transfer_learning/main.py
python 08_word_embeddings/main.py- Feature scaling is critical for linear models but not for tree-based models
- TF-IDF + SVM is a surprisingly strong baseline for text classification
- K-Means requires choosing K carefully — use elbow method and silhouette score
- CNNs can learn spatial features automatically from images
- Model comparison is essential — never rely on a single model
- Custom Datasets & DataLoaders in PyTorch give full control over batching and preprocessing
- Transfer learning dramatically reduces training time — freeze layers, retrain classifier
- Word embeddings capture semantic similarity — similar words cluster in vector space
MIT License
Built by Imran Ali — MUET Jamshoro, CS 2026