Ho

How to guide on training Llama-405B using PyTorch distributed APIs

Hacker News

How to guide on training Llama-405B using PyTorch distributed APIs

Been working on a guide for ML folks to upgrade their single GPU training code to multi GPU and multi node. Code diffs and explanations are included. The guide builds up to this final chapter (linked) on how to train a very large model like Llama 3.1 405B on a big cluster with plain pytorch. Everything is just written using the direct pytorch apis (other than the model code which is just using `transformers` models). If there are topics of interest feel free to open an issue in the repo, and contributions are welcome. I'm investigating adding a chapter on tensor parallelism, but it's support in pytorch is still early stages.

Share card

Actual performance

3points
4comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, single · Missing: mac, agents, macos
88%88% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Hacker NewsStrong engagement from HN community · Strong signals: llama, ide, io · Missing: https docs, excited, just released
57%57% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Indie HackersIH features products with proven revenue · Strong signals: para · Missing: supports, reddit linkedin, podcasting
48%48% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: para · Missing: mobile apps, ios, personal
39%39% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
25%25% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
24%24% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
9%9% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

I
I reproduced Code Llama fill-in-the-middle code completion training72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I reproduced Code Llama fill-in-the-middle code completion training

Hacker News2
Ll
Llama or Alpaca?78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama or Alpaca?

Hacker News6
Ll
Llama 3.2 Interpretability with Sparse Autoencoders74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 3.2 Interpretability with Sparse Autoencoders

Hacker News579
Mu
Multi node training of Llama 70B without crying59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multi node training of Llama 70B without crying

Hacker News1
Te
Terraform llama.cpp on AWS for your own generative APIs71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Terraform llama.cpp on AWS for your own generative APIs

Hacker News2
Fi
Finetune LLaMA-7B on commodity GPUs using your own text63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Finetune LLaMA-7B on commodity GPUs using your own text

Hacker News449
Ll
Llama 2 Uncensored 70B as API79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 2 Uncensored 70B as API

Hacker News18
Te
Testing GraphQL APIs Using Karate DSL66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Testing GraphQL APIs Using Karate DSL

Hacker News3
Fi
Finetune Llama-3.1 2x faster in a Colab74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Finetune Llama-3.1 2x faster in a Colab

Hacker News16
Bu
Build a RAG Chatbot using Llama 2 with Replicate and LlamaIndex52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Build a RAG Chatbot using Llama 2 with Replicate and LlamaIndex

Hacker News9