Deep learning framework from scratch, trains GPT-2 in 3 days
Deep learning framework from scratch, trains GPT-2 in 3 days
I wanted to share Tricycle ( https://github.com/bclarkson-code/Tricycle ), a deep learning framework I built completely from scratch from Autograd to a GPT. I wanted a library that is fast and feature rich enough to train actual models while being simple enough that anyone with a bit of python experience can understand what is going on. The biggest milestone so far is training GPT-2 (124M) on 2.3B tokens in just under 3 days on my GPU (RTX 3090). So far, I've added the following to Tricycle: - An automatic differentiation engine - General matrix operations with einsum - Standard network layers (Dense, ReLU, GeLU etc) - Transformer blocks (MultiHeadSelfAttention and MLP blocks) - Optimisers (SGD, AdamW) - GPT-2 - etc The project is still under active development, I'm in the process of adding mixed precision and multi-gpu support with the goal of scaling up to larger models. To see it in action, the best place to start is train_smol_gpt.py which will train GPT-2 from scratch. Let me know what you think!
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
*From Scratch* Implementations of Deep Learning Concepts (in Numpy)
Deep Learning Framework from scratch (60 step tutorial)
Mariana, the Cutest Deep Learning Framework
Picard – Framework for exploring deep learning architectures
Bezos: build your own (deep) reinforcement learning framework
Bender – a Deep Learning framework for iOS
Simple RPC Framework in Golang from Scratch
GPT-2 from scratch in PyTorch – Video tutorial
Building a GPT from scratch: What I learned and why it mattered
I built a deep learning engine from scratch in Python