An

Analyzing top HN posts with language models

Hacker News

Analyzing top HN posts with language models

Hi HN, I spent a few weeks looking at the top HN posts of all time. This included exploration, clustering, creating visualizations, and zooming in on what (to me personally) seems like some of the best discussions on here. Three things in this post: 1- The interesting groups of HN posts 2- The interactive visualizations that you can explore in your browser 3- The data from this exploration -- this includes CSV of the titles as well as the text embeddings of 3,000 Ask HN articles. Blog post about this whole process here: [1] ============ 1- The interesting groups of HN posts From the exploration, Ask HN proved the most interesting. These are the top four groups of topics I found insightful. Each group contains about 400 posts. - Life experiences and advice threads [2] - Technical and personal development [3] - Software career insights, advice, and discussions [4] - General content recommendations (blogs/podcasts) [5] ============ 2- The interactive visualizations that you can explore in your browser - Top 10,000 Hacker News articles of all time [6] - Top 3,000 posts in Ask HN [7] ============ 3- The data from this exploration CSV file of top 3K Ask HN posts: [8] The sentence embeddings of the titles of those posts: [9] This is a colab notebook containing the code examples (including loading these two data files): [10] ============ If you've ever wanted to get into language models, this is a good place to start. Happy to answer any questions

Share card

Actual performance

117points
43comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Hacker NewsStrong engagement from HN community · Strong signals: hacker news, 000, io · Missing: https docs, excited, just released
72%72% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
71%71% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Indie HackersFits the IH revenue-focused audience · Strong signals: including · Missing: supports, reddit linkedin, podcasting
52%52% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: personal · Missing: mobile apps, ios, entrepreneurs
45%45% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
34%34% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Instella
Instella80%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open 3B language models from AMD

Product Hunt+120Open Source
La
Language Models API via the Cohere Platform63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Language Models API via the Cohere Platform

Hacker News18
Ba
BadSeek – How to backdoor large language models75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BadSeek – How to backdoor large language models

Hacker News461
LLM OneStop
LLM OneStop82%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

All your language models in one place

Product Hunt+6
Im
Implementation of the "Self-Rewarding Language Models" Paper by MetaAI67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Implementation of the "Self-Rewarding Language Models" Paper by MetaAI

Hacker News23
In
Inseq – An Interpretability Toolkit for Generative Language Models68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Inseq – An Interpretability Toolkit for Generative Language Models

Hacker News1
Ex
Explore large language models with 512MB of RAM72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Explore large language models with 512MB of RAM

Hacker News138
TE
TEG, a linguistic game powered by large language models67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TEG, a linguistic game powered by large language models

Hacker News1
Ha
Hackers Guide to Language Models63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hackers Guide to Language Models

Hacker News8
Ph
Phare: A Safety Probe for Large Language Models55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Phare: A Safety Probe for Large Language Models

Hacker News4