Dynamic cluster-based data sampling for efficient and long-tail-aware vision-language model pre-training.
-
Updated
Jul 7, 2026 - Python
Dynamic cluster-based data sampling for efficient and long-tail-aware vision-language model pre-training.
`decon`, but with python API binding.
DataComp-style image-text dataset filtering on Backblaze B2: stream WebDataset shards from object storage, score image-text alignment with CLIP (open_clip), write filtered shards + quality metrics back to B2 — no local staging, no database. Full-stack Next.js + FastAPI sample for VLM-pretraining data curation
Add a description, image, and links to the datacomp topic page so that developers can more easily learn about it.
To associate your repository with the datacomp topic, visit your repo's landing page and select "manage topics."