Sw

Swiss Army Llama – A Versatile, FastAPI-Based Multitool for Local LLMs

Hacker News

Swiss Army Llama – A Versatile, FastAPI-Based Multitool for Local LLMs

This project originally started out with a focus on easily generating embeddings from Llama2 and other llama_cpp (gguf) models and storing them in a database, all exposed via a convenient REST api (hence why it was originally called the not-very-catchy name of "llama_embeddings_fastapi_service" when I submitted it a few weeks ago). But since then, I've added a lot more functionality: 1) New endpoint for generating text completions (including specifying custom grammars, like JSON). 2) Get all the embeddings for an entire document--can be any kind of document (plaintext, PDFs, doc/.docx, etc.) and it will do OCR on PDFs and images. 3) Submit an audio file (wav/mp3) and it uses whisper to transcribe it into text, then gets the embeddings for the text (after combining the transcription segments into complete sentences). 4) Integrates with vector similarity library (pip install fast_vector_similarity) to provide an "advanced" semantic search endpoint. This uses a 2-step process: first it uses FAISS to quickly narrow down the set of stored embeddings us cosine similarity, then it uses the vector similarity library to compute a bunch of more sophisticated (and computationally intensive) measures for the final ranking. 5) An endpoint to automatically generate a BNF grammar definition from a sample JSON file or string, and also from a Pydantic data model definition. Grammar files can be directly used with llama_cpp for constrained sampling, an incredibly useful thing when making applications. Also includes code for automatically validating grammar files. 6) An endpoint to view the application logs in a nice web view with helpful coloring, with the ability to download the logs or copy to clipboard. 7) And endpoint to add a new model file by supplying the URL to the model (e.g., a Huggingface URL). Previously, you had to manually edit a function in the code to add a new model. As a result of all these additions, I changed the project name to Swiss Army Llama to reflect the new project goal: to be a one-stop-shop for all your local LLM needs, so you can easily integrate this technology in your programming projects. As I think of more useful endpoints to add (I constantly get new feature ideas from my own separate projects-- whenever I want to do something that isn't covered yet, I add a new endpoint or option), I will continue growing the scope of the project. So let me know if there is some functionality that you think would be generally useful, or at least extremely useful for you! A big part of what makes this project useful to me is the FastAPI backbone. Nothing beats a simple REST API with a well-documented Swagger page for ease and familiarity, especially for developers who aren't familiar with LLMs. You can set this up in 1 minute on a fresh box using the docker TLDR commands, come back in 15 minutes, and it's all set up with downloaded models (I include and ready to do inference or get embeddings. It also lets you distribute the various pieces of your application on different machines connected over the internet.

Share card

Actual performance

1points
2comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: started, para, including · Missing: supports, reddit linkedin, podcasting
95%95% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, dock · Missing: agents, macos, agent
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Missing: plus, platform, intuitive
53%53% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: llama, ide, io · Missing: https docs, excited, just released
49%49% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: para · Missing: mobile apps, ios, personal
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ly
Lyp – The Lilypond Swiss Army Knife46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Lyp – The Lilypond Swiss Army Knife

Hacker News2
Sw
Swiss avalanche casualties visualisation46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Swiss avalanche casualties visualisation

Hacker News5
Th
The IT contractor's Swiss army knife52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The IT contractor's Swiss army knife

Hacker News2
rekordcloud
rekordcloud9%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Swiss Army knife for DJs

Indie Hackersemployees-0
Ll
Llama or Alpaca?78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama or Alpaca?

Hacker News6
Ll
Llama 3.2 Interpretability with Sparse Autoencoders74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 3.2 Interpretability with Sparse Autoencoders

Hacker News579
To
Tokenkit – Convert LLMs to new tokenizers (incl byte-level Llama/Gemma)68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tokenkit – Convert LLMs to new tokenizers (incl byte-level Llama/Gemma)

Hacker News1
LL
LLaMA Nuts and Bolts, A holistic way of understanding how LLMs run76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLaMA Nuts and Bolts, A holistic way of understanding how LLMs run

Hacker News3
IC
ICanHazData – Data URI Swiss Army Knife50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ICanHazData – Data URI Swiss Army Knife

Hacker News257
Si
SimpleNet – A modern BBS moderated by local LLMs47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SimpleNet – A modern BBS moderated by local LLMs

Hacker News1