Skip to content
r/LocalLLaMA · Communities

My learnings from optimizing training pipeline to go from 36 steps/minute to 47 steps/minute

I got my ML model training pipeline to go from 36 steps/minute to 47 steps/minute by optimizing how they're stored on disk. When training larger ML models on consumer hardware there are many limitations. One is the availability and performance of the storage devices. Everyone would like to have NVMe drives on their sys