High-performance CUDA inference engine for Qwen models, written from scratch in C++/CUDA
-
Updated
Sep 26, 2026 - Cuda
High-performance CUDA inference engine for Qwen models, written from scratch in C++/CUDA
针对YOLOv11进行fp16和ptq的int8量化,显著提升推理速度(C++) (包含完整模型转换流程和代码)
Ahrireyes 모델 및 파일들 저장소
High-performance LPR system optimized for Indian license plates, achieving 97% character accuracy. Features a hybrid pipeline using YOLOv11 and Mamba-SSM (State Space Models) with built-in regex correction and Beam Search decoding.
Successfully developed a lightweight YOLOv8n structural defect detector trained on 8,162 images across 7 defect classes, converted to ▎ TensorRT FP16 for edge deployment achieving 157.6 FPS — a 40.7% speedup over the FP32 baseline with under 1% accuracy loss.
To associate your repository with the fp16-optimization topic, visit your repo's landing page and select "manage topics."