Skip to content
View Muhammad-Ahmad-Waseem's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report Muhammad-Ahmad-Waseem

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Muhammad Ahmad Waseem

ML Engineer Β· Computer Vision, ASR Β· ML Systems & Inference Optimization Β· Robotic Perception

"Building ML systems that ship β€” from training to optimizing low-level CUDA kernels."

LinkedIn Portfolio Email ORCID


About Me

ML Engineer completing M.S. in CS at University at Buffalo (GPA 3.96), with 6+ years of Research Experience, 3+ years shipping production ML systems. My work lives at the intersection of Model Finetuning, GPU kernel optimization, and efficient model deployment.

  • 🎀 CV, ASR & AI β€” UNet, Yolo, DeepLab, Wav2Vec, Conformer, Diffusion Models, LLMs, HuggingFace, Pytorch, TensorFlow, Nemo
  • ⚑ GPU Systems β€” CUDA kernels, TensorRT, QAT/PTQ, AWQ/GPTQ/SmoothQuant, NVCC, PySpark, GCP Dataproc
  • πŸ”¬ Research β€” 7 peer-reviewed publications across IEEE IGARSS, BMVC, SDSC; 20+ citations
  • πŸš€ Open to Applied ML Engineer Β· ML Systems Engineer Β· Research Engineer . Perception roles from July 2026 (STEM OPT)

πŸ”₯ Active Projects

Project Description Status
AutoPhon Phoneme-level speech annotation tool using Wav2Vec 2.0 + Whisper; reduces annotation time 10Γ— 🟒 Active
BiDAAM Post-hoc interpretability framework for diffusion models πŸ“„ Accepted at IEEE SMC 2026
TF-Net Multi-task satellite image segmentation; SOTA 94% F1 on SpaceNet2 + WHU βœ… IEEE IGARSS 2025

πŸ› οΈ Tech Stack

Computer Vision and 3D

UNet Yolo DeepSort DeepLab DSAC HourGlass Colmap

Speech & ASR

Wav2Vec Conformer Whisper CTC Kaldi NeMo

GPU & Systems

CUDA TensorRT Triton nvcc

Quantization & Compression

QAT AWQ GPTQ SmoothQuant LLM.int8

ML Frameworks

PyTorch HuggingFace TensorFlow OpenCV

Languages

Python C++ C CUDA


πŸ“„ Publications

# Paper Venue Year
C.1 BiDAAM: A Bi-directional Attribution Framework for Diffusion Models IEEE SMC 2026
C.2 A Tuning-Fork Network for Improved Building Footprint Extraction IEEE IGARSS 2025
C.3 Evaluating Cooling Efficacy of Urban Green Spaces During Extreme Heat Events IEEE IGARSS 2025
C.4 Unsupervised Landmark Discovery Using Consistency Guided Bottleneck BMVC 2023
C.5 Improved Flood Mapping for Efficient Policy Design IEEE IGARSS 2023
C.6 PD-SEG: Population Disaggregation Using Deep Segmentation Networks IEEE IGARSS 2023
C.7 Estimating Spatio-Temporal Urban Development using AI SDSC 2022

πŸ’Ό Experience

University at Buffalo (SUNY)                    Aug 2024 – Jul 2026
Graduate Research Assistant
Β· ASR    β†’ Reduced PER 80%β†’15% on children's speech via Wav2Vec 2.0 + AutoPhon pipeline
Β· GPU    β†’ CUDA kernels for transformer attention/FFN on A100; profiled vs cuBLAS
Β· Quant  β†’ Outlier-aware QAT for Conformer ASR: 7.17 WER vs 93.7 naive INT8
Β· Bench  β†’ AWQ/GPTQ/SmoothQuant across A100, Jetson, Raspberry Pi

CITY at LUMS                                    Aug 2021 – Jul 2024
ML Research Engineer
Β· CV      β†’ Satellite Image Segmentation, Vehicle Detection and Classification, Route Optimization
Β· Deploy  β†’ TensorRT on NVIDIA AGX Xavier: 5Γ— throughput (3β†’14 FPS); Punjab Safe Cities Authority
Β· Publish β†’ TFNet: A novel architecture for Building Footprint Extraction; 94% F1 on SpaceNet2; IEEE IGARSS 2025
Β· Prod    β†’ City-scale waste routing platform: ~100K liters/month fuel savings (LWMC)

Information Technology University               Sep 2020 – Jul 2021
Graduate Research Assistant
Β· 3D      β†’ DSAC* for camera pose estimation; Colmap for 3D Reconnstruction
Β· CV      β†’ Unsupervised facial landmark extraction; GNN, GCN

Available for full-time roles from July 2026 Β· STEM OPT Β· Buffalo, NY β†’ Open to relocation

Pinned Loading

  1. Building-Detection Building-Detection Public

    This repository is based on my work on Estimating Spatio-temporal urban development using AI at Team Urban Tech LUMS.

    Python 1

  2. Hardware-Accelerator-Design-for-LeNET Hardware-Accelerator-Design-for-LeNET Public

    This repository contains my semester project work on designing hardware accelerator for Handwritten digit recognition model (LeNET).)

    VHDL 1

  3. PyQGIS3-Codes PyQGIS3-Codes Public

    This Repository contains codes used for different operations in QGIS3 using python at Team Urban Tech LUMS.

    Python

  4. T30G_Code T30G_Code Public

    C++

  5. DHA-Dataset DHA-Dataset Public

    The repo contains data for the area of Lahore that was used in the paper TF-Net

    1

  6. rdaam rdaam Public

    Course project for CSE 455 - Intro to Pattern Recognition

    Python