Th

The problem with the epsilon greedy method

Hacker News

The problem with the epsilon greedy method

I decided I wanted to roll my own AB testing app for Django (https://github.com/crobertsbmw/RobertsAB) when I was finished, I came across this: http://stevehanov.ca/blog/index.php?id=132 Which is a very convincing article on why AB testing sucks and with a few extra lines, you can improve your algorithm to select the best test so you never go back and update your code (yeah right.) I then thought, how many tests does this thing have to run to truly figure out which is best? I made 4 tests with probability of success equalling 1/2, 1/4, 1/5, 1/6 and found that for this algorithm to settle on the best success rate (1/2), it took 91 hits on average with a max of 876 tests. I ran the same test using a standard AB algorithm. Picking whichever test has been tested the least and run that test. It took on average 32 tests to figure out which performed the best with a maximum of 363. On average 3 times better than the greedy epsilon method. I then tried tweaking my success ratios to something a little less dramatic. 1/10, 1/11, 1/12, 1/13. Which just made everything take a LOT longer. The only problem is that in reality you don't know what the best solution is, so you can never know if you have gotten to the "actual" solution. The epsilon greedy method will eventually get there (although you will never know when). And if you are using the standard AB method you will never know if you have arrived at the best option either, especially when we are talking about the difference between 1/20 clicks versus 1/21 clicks. Moral of the story -- AB testing is probably a waste of time. Here is a link to all the tests I ran (python3): https://github.com/crobertsbmw/EpsilonGreedy

Share card

Actual performance

22points
3comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: ios · Missing: supports, reddit linkedin, podcasting
78%78% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
36%36% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: ios · Missing: mobile apps, personal, entrepreneurs
36%36% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: ide, io · Missing: https docs, excited, just released
33%33% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Product HuntUnlikely to reach the leaderboard · Strong signals: using, code · Missing: mac, agents, macos
21%21% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Go
Go-dutchflag – An implementation of the Dutch flag problem in Golang37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Go-dutchflag – An implementation of the Dutch flag problem in Golang

Hacker News3
Th
The AZ Problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The AZ Problem

Hacker News2
So
Someone Else's Problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Someone Else's Problem

Hacker News1
Th
The farmer, wolf, goat and cabbage problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The farmer, wolf, goat and cabbage problem

Hacker News1
Mo
Monty Hall Problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Monty Hall Problem

Hacker News1
Wo
Wolf, Goat and Cabbage Problem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Wolf, Goat and Cabbage Problem

Hacker News1
Gu
Guiderail – How I (partly) solved my procrastination problem46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Guiderail – How I (partly) solved my procrastination problem

Hacker News3
Ar
Are the Riemann Hypothesis and Navier-Stokes the Same Problem?56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Are the Riemann Hypothesis and Navier-Stokes the Same Problem?

Hacker News7
Grubl
Grubl65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The "What's for dinner" problem solved!

Indie Hackerscommitment-full-time
I
I solved the subset sum problem in polynomial time53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I solved the subset sum problem in polynomial time

Hacker News3