Ne

New LLM outperforming GPT-3.5

Hacker News

New LLM outperforming GPT-3.5

Refuel LLM (84.2%) outperforms trained human annotators (80.4%), GPT-3-5-turbo (81.3%), PaLM-2 (82.3%) and Claude (79.3%) across a benchmark of 15 text labeling datasets. It is a Llama-v2-13b base model, trained on over 2500 unique datasets (5.24B tokens) spanning categories such as classification, entity resolution, matching, reading comprehension and information extraction. Here is the interactive demo: https://labs.refuel.ai/playground. Pretty fun to play with!

Share card

Actual performance

6points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: claude, model, new · Missing: mac, agents, macos
88%88% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
TrustMRRFits verified-revenue profile · Missing: mobile apps, ios, personal
57%57% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: llama, io · Missing: https docs, excited, just released
52%52% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Indie HackersIH features products with proven revenue · Missing: supports, reddit linkedin, podcasting
35%35% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
29%29% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
18%18% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
13%13% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

St
Structed LLM outputs via Pydantic with struct-GPT64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Structed LLM outputs via Pydantic with struct-GPT

Hacker News1
LL
LLM OSINT, letting GPT-4 Google about you46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM OSINT, letting GPT-4 Google about you

Hacker News1
SH
SHOW HN:A New 34B Open Source LLM, Astonishing 78 Score in MMLU (GPT-4 MMLU:83)64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SHOW HN:A New 34B Open Source LLM, Astonishing 78 Score in MMLU (GPT-4 MMLU:83)

Hacker News21
Wo
Worldbuilding Experiments with GPT-336%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Worldbuilding Experiments with GPT-3

Hacker News2
Au
Autosummarized HN (With GPT-3)51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Autosummarized HN (With GPT-3)

Hacker News6
I
I composed a sonata with GPT-3 DaVinci-003 and you can too51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I composed a sonata with GPT-3 DaVinci-003 and you can too

Hacker News3
GP
GPT Classifies HN Titles57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPT Classifies HN Titles

Hacker News6
Vi
Visualized GPT57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Visualized GPT

Hacker News3
Ce
Cerebras-GPT-2.7B finetuned on Stanford Alpaca dataset65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cerebras-GPT-2.7B finetuned on Stanford Alpaca dataset

Hacker News4
Ro
Roleplaying GPT40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Roleplaying GPT

Hacker News3