Skip to content
View verydnobl337's full-sized avatar
😁
😁

Block or report verydnobl337

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
verydnobl337/README.md

Hi there 👋

I'm Daniil

Data Engineer focused on building reliable data pipelines, data warehouses, and scalable data processing systems.


Tech Stack

Programming & Data Engineering

Python · SQL · ETL/ELT · Apache Airflow · Apache Spark · PySpark · Apache Kafka

Databases & Storage

PostgreSQL · Vertica · MongoDB · Redis · S3 · HDFS · DWH · Data Lake · Data Vault · OLTP · OLAP

Big Data & Infrastructure

Hadoop · YARN · Parquet · Avro · Docker · Kubernetes · Yandex Cloud

Tools

Git · GitHub


Projects

  • Fintech Data Platform - ETL/ELT pipeline with Airflow, S3, PostgreSQL and Vertica.

  • Cloud Data Platform - Kafka, Redis, PostgreSQL pipeline with Data Vault and incremental data processing.

  • Spark Geo Analytics - Spark and Airflow pipeline for geospatial analytics and recommendation processing.

  • Spark Streaming Pipeline - Streaming data processing project using PySpark, Kafka, and PostgreSQL for personalized promotions.

  • Analytical Data Warehouse - an end-to-end data engineering project that implements a data pipeline from S3 through Apache Airflow to Vertica DWH.

  • Multi-Source DWH Pipeline - multi-source DWH pipeline with Apache Airflow, PostgreSQL, REST APIs, MongoDB, and ETL/ELT processes.


Contact and Links

Pinned Loading

  1. s3-airflow-vertica s3-airflow-vertica Public

    Batch ETL/ELT pipeline using Apache Airflow, S3, Vertica, and SQL with idempotent data loading and analytical data marts.

    Python

  2. cloud-technologies cloud-technologies Public

    Event-driven streaming data platform built with Apache Kafka, PostgreSQL, Redis, and Docker. Implements STG, Data Vault-style DDS, and CDM layers with microservices-based data processing and increm…

    Python

  3. spark-geo-analytics spark-geo-analytics Public

    Apache Spark batch pipeline for user geo analytics, city detection, travel analysis, zone metrics, and friend recommendations, orchestrated with Apache Airflow on YARN/HDFS.

    Python

  4. spark-streaming-project spark-streaming-project Public

    A streaming service built with PySpark Structured Streaming that consumes advertising campaigns from Kafka and matches them with user subscriptions stored in PostgreSQL. The results are written to …

    Python

  5. analytical-dwh analytical-dwh Public

    Analytical Data Warehouse — an end-to-end data engineering project that implements a data pipeline from S3 through Apache Airflow to Vertica DWH. The project includes staging, Data Vault 2.0-style …

    Python

  6. multi-source-dwh multi-source-dwh Public

    Data Engineering project for building a multi-source DWH pipeline with Apache Airflow, PostgreSQL, REST APIs, MongoDB, and ETL/ELT processes. Includes STG, DDS, and CDM layers, data marts, incremen…

    Python