Skip to content
#

pretraining-data

Here are 5 public repositories matching this topic...

Unofficial reproduction of Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training (NeurIPS 2025, arXiv:2504.13161) — search-found mixtures beat uniform baselines +0.014–0.031 STEM at d28; novel finding: selection-mechanism winner's curse; Ascend NPU backend.

  • Updated Sep 22, 2026
  • Python

Do quality filters pull benchmark questions into pretraining corpora? Injects MMLU, GSM8K, GPQA, ARC, HellaSwag, PIQA and TruthfulQA items into a corpus and ranks them with six quality classifiers, including DCLM, FineWeb-Edu and Gaperon.

  • Updated Aug 19, 2026
  • Python

Add this topic to your repo

To associate your repository with the pretraining-data topic, visit your repo's landing page and select "manage topics."

Learn more