How to guide on training Llama-405B using PyTorch distributed APIs
How to guide on training Llama-405B using PyTorch distributed APIs
Been working on a guide for ML folks to upgrade their single GPU training code to multi GPU and multi node. Code diffs and explanations are included. The guide builds up to this final chapter (linked) on how to train a very large model like Llama 3.1 405B on a big cluster with plain pytorch. Everything is just written using the direct pytorch apis (other than the model code which is just using `transformers` models). If there are topics of interest feel free to open an issue in the repo, and contributions are welcome. I'm investigating adding a chapter on tensor parallelism, but it's support in pytorch is still early stages.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
I reproduced Code Llama fill-in-the-middle code completion training
Llama or Alpaca?
Llama 3.2 Interpretability with Sparse Autoencoders
Multi node training of Llama 70B without crying
Terraform llama.cpp on AWS for your own generative APIs
Finetune LLaMA-7B on commodity GPUs using your own text
Llama 2 Uncensored 70B as API
Testing GraphQL APIs Using Karate DSL
Finetune Llama-3.1 2x faster in a Colab
Build a RAG Chatbot using Llama 2 with Replicate and LlamaIndex