Ru

Run any Llama model finetune and more, instantly

Hacker News

Run any Llama model finetune and more, instantly

Hi there, looking for feedback on my new project "Featherless.AI" The idea is to allow users to run all the models on hugging face instantly. Via the OpenAI API compatible endpoint. Why? Because its a real chore to download models and spin up GPUs, especially if you want to test multiple models. Not to mention GPUs cost multiple dollars an hour to rent. And if we want more people to use open source AI, we got to make it easier for them to try and play with all of them. So what if instead of spinning up dedicated GPUs per model (which is what every provider is doing) We can startup a LLM server, which multiple users can share, and instantly hot swap it in seconds the model its serving??? Make sense right, thats what we do for apps, or shared CPU PHP apps. --- Turns out thats an incredibly complex problem, that even Hugging Face, together AI and various companies had failed to successfully build. Turns out swapping out GB's of data from disk to GPU in subseconds is hard. And involves optimizing every small thing in the chain - had to setup and tune a high speed storage cluster - find a GPU provider with crazy networking speed, and would allow you to scale up/down - custom write a new pipeline for streaming data from high speed storage cluster - custom write code to stream that data, straight into GPU as fast as possible After all that, you now have a GPU server, that can serve one model at a time, but can switch models quickly. So to make this idea economical (since im not billing anyone $8/hour) - you will need to setup a cluster, to handle multiple request concurrently - while writing routing code, to ensure all the request for the same models go to the same server - unless it hits the limit, then you need to start load balancing between servers assigned to that model - and to backoff the load balancing, so that you can free up servers - to be hot swapped for other request Oh also downloading 450+ models, apparently takes up tons of TB's, of expensive high speed clustered storage. But the end result, a highly dynamic scaling (up or down) infrastructure, to the exact cluster workload, for a large collection of HF models (more models then GPUs of course). All so that people can use "all the open source AI models", for a few dollars a month. While not worrying about token pricing. So do give it a try, the free trial account can test any 8B model, and a subscribe account has full OpenAI API access. And give us some feedback, and maybe even a product hunt vote! PS: This is a work in progress, I yet to rewrite the custom pipeline code for non-llama / non-rwkv models, thats why we are starting with only those 2 architectures first. But we do plan to scale to ALL public models on huggingface. Also downloading the other 1000+ llama models does take a long time.

Share card

Actual performance

7points
2comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: compatible · Missing: supports, reddit linkedin, podcasting
95%95% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, apps, user · Missing: mac, agents, macos
94%94% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, llama, ide · Missing: https docs, excited, just released
63%63% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: apps, month, users · Missing: mobile apps, ios, personal
38%38% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: users · Missing: plus, platform, intuitive
27%27% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
25%25% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ru
Run Llama 3.1 8B in the browser78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run Llama 3.1 8B in the browser

Hacker News23
Ll
Llama or Alpaca?78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama or Alpaca?

Hacker News6
Ll
Llama 3.2 Interpretability with Sparse Autoencoders74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 3.2 Interpretability with Sparse Autoencoders

Hacker News579
I
I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp

Hacker News5
Ll
Llama 2 Uncensored 70B as API79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 2 Uncensored 70B as API

Hacker News18
Fi
Finetune Llama-3.1 2x faster in a Colab74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Finetune Llama-3.1 2x faster in a Colab

Hacker News16
Apollo AI
Apollo AI93%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run local models like Llama on iOS

Product Hunt+280iOS
LL
LLaMA Nuts and Bolts, A holistic way of understanding how LLMs run76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLaMA Nuts and Bolts, A holistic way of understanding how LLMs run

Hacker News3
Fi
Finetune Llama 3.2 Vision in a Colab75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Finetune Llama 3.2 Vision in a Colab

Hacker News10
To
TokenHawk, WebGPU Running LLaMA64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TokenHawk, WebGPU Running LLaMA

Hacker News2