Skip to content
View diaazg's full-sized avatar

Organizations

@Enhanced-TEVAD

Block or report diaazg

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
diaazg/README.md
Mohamed Diaa Zellagui — AI & Data Science Engineer AI & Data Science Engineer · Computer Vision · NLP · Multimodal Learning · Research that ships as working systems · 3 years shipping mobile apps

I rebuild things before I trust them. Half the projects here started because I wanted to know how something worked, not because anyone asked for them.

Published datasets: 28,600 annotated lines Research papers: 2
Data pipelines: multi-source ETL, end to end Models deployed behind REST APIs Mobile apps: 3 years freelance

🔬 Researcher

Computer Vision NLP Multimodal Learning Vision-Language Models Video Understanding Weak and Self-Supervision Anomaly detection Document AI and HTR Dataset Design and Benchmarking



📄 Papers

RefLAM AraMS-28k — draft

⚙️ Engineer

Model Serving and APIs Distributed Pipelines MLOps and Orchestration LLM Agents and RAG Spatial ML and Causal Inference Data Engineering Interactive Dashboards Mobile Development



📱 3 years freelance


🗣️ Languages

Arabic — Native French - Intermediate English — Advanced

📬 Get in touch

Email LinkedIn WhatsApp
Google Scholar Hugging Face View CV


Open to research and engineering roles in computer vision, NLP, and multimodal ML — onsite or remote.

🎓 AI Research Engineer  ·  📍 Algeria, open to relocation

Pinned Loading

  1. ArchaText/AraMS-28k-Dataset ArchaText/AraMS-28k-Dataset Public

    The largest publicly released line-level dataset of historical Arabic manuscripts — 14 books, 3,043 pages, 28,600 lines, with margin/insertion-anchor annotations for non-linear reading order.

    Python 3

  2. ArchaText/AraMS-Restore ArchaText/AraMS-Restore Public

    Line-level restoration of degraded Arabic manuscripts: a U-Net that makes damaged handwriting readable, scored by whether an OCR model can read it.

    Python 1

  3. Reatail-Loc-London Reatail-Loc-London Public

    A spatial analytics pipeline that scores every Lower Layer Super Output Area (LSOA) in Greater London for retail suitability, using open urban data and machine learning.

    Jupyter Notebook 1

  4. code-comment code-comment Public

    Domain-specific adaptation of CodeT5 for Python docstring generation. Compares pre-trained vs fine-tuned performance using NLP metrics (ROUGE, METEOR, BERTScore), showing significant improvements a…

    Jupyter Notebook 3

  5. image_segmentation image_segmentation Public

    COVID-19 lung segmentation using U-Net and MobileNet-U-Net with MLflow for experiment tracking and model comparison.

    Jupyter Notebook 3

  6. Personnel_chatbot Personnel_chatbot Public

    My LangGraph-built AI assistant that provides conversational access to my professional portfolio. It leverages tool-calling to dynamically fetch and present data from my GitHub, answer questions ab…

    Jupyter Notebook 3