Pa

Part 4: Train CIFAR10 to 94% in under 8 seconds on a single A100

Hacker News

Part 4: Train CIFAR10 to 94% in under 8 seconds on a single A100

Hello everyone, It's been two weeks since we moved under the 10 second mark, and as we've made some more progress with (some very hard) work on the issue in the past two weeks, and we've past our internal 8 second benchmark for another release, we're releasing this next update! This update changes the neural network architecture to our own new, custom 8-layer ResNet architecture (dubbed SpeedyResNet) which is extremely simple and fast. We also do some hyperparameter tuning, round the hyperparameters to rounder numbers than they were before, and also change up the learning process a bit by changing how we use our EMA. We also do this by only adding 2 (or 3, depending upon what you're counting) lines of new code! The vast majority of the rest of the work is editing, changing, or simply outright deleting other code. This results in a codebase that is a bit simpler (at least in layout) and faster than before. We also eliminate a hyperparameter that seems to be no longer useful. One downside of these changes is that we do overfit slightly more on longer runs, but that can be mitigated enough with cutout, and we hope to fix this in future releases, as this is not a terrible problem to have when trying to set speed records. We test our code on CIFAR100 without any modifications (other than to the dataloaders to load the correct number of data and the correct number of classes) and show that performance for those two different network sizes is comparable. To do this, we show (at least in rough initial explorations), that both of the small networks matched the performance of SOTA networks in around the same year, and that increasing the base depth of the network by a factor of two improved the performance of both networks by about a year to match respective SOTAs from the same time period. This indicates that this code (hopefully) has some good generalization capabilities beyond just this dataset, though we have not experimented with different image sizes yet (it's rather expensive and the information might get stale very quickly!) There's a lot more in here, but as in previous posts, the mantra of sorts is 'doing the basics and doing them very well'. This goes a whole lot further than it might seem otherwise, having 'the new shiny' when developing neural networks is oftentimes more of a toy and a distraction than sticking with the basics and doing them well. Which is understandably a very difficult thing. That said, as we run out of runway for the 'easier' changes, we will likely need to get more and more creative. But until then, the goal is to stay as simple as possible! If you'd like more info, please do read the release notes, they are very informative although longer. Future releases could have more to do with speed improvements or other things. Additionally, this still should be an excellent researcher's toolbench for prototyping and experimenting with ideas. Many ideas I've been able to implement in 5 minutes or less, most of which are actually running in the code by that point. For some, like sometimes architecture changes, for example, I'm able to get a complete initial go/no go filter answer oftentimes within 1-2 minutes of the idea. I just quickly tweak it, let it run through a few runs, and either have to let it run more runs to see if it is just a noisy answer or if it definitely doesn't work. This is indispensable and part of why I built this tool. It's also partially responsible for the rapid progress in developing this tool -- I'm able to apply the very-rapidly-gained insights from this tool to itself. I'll be hanging around here for questions/comments/etc. I can't answer all of them, but I'll do the best that I can! :D

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: new, single, notes · Missing: mac, agents, macos
85%85% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: para · Missing: supports, reddit linkedin, podcasting
84%84% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, io · Missing: https docs, excited, just released
64%64% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: way, para · Missing: mobile apps, ios, personal
53%53% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Tr
Train against procrastination42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Train against procrastination

Hacker News9
Tr
Train Stable Diffusion Dreambooth on 1080ti60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Train Stable Diffusion Dreambooth on 1080ti

Hacker News2
Fi
Five C++ single-header libraries by Chobolabs54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Five C++ single-header libraries by Chobolabs

Hacker News2
Co
Cookiebro WebExtension can now blacklist single cookies by name46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cookiebro WebExtension can now blacklist single cookies by name

Hacker News4
Di
Disable Ugly Firefox Single Rounded Corner42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Disable Ugly Firefox Single Rounded Corner

Hacker News2
To
Totpgate – Single-packet authorization via TOTP45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Totpgate – Single-packet authorization via TOTP

Hacker News1
Ba
Basic Distributed AI Train Tool42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Basic Distributed AI Train Tool

Hacker News1
ML
MLbot – Train ML models in your cloud with a single command57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

MLbot – Train ML models in your cloud with a single command

Hacker News4
Li
Live Train Webcams35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Live Train Webcams

Hacker News15
Ho
How to use quilt to train an SVM42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How to use quilt to train an SVM

Hacker News1