NO

NOS – A fast, and ergonomic PyTorch inference server

Hacker News

NOS – A fast, and ergonomic PyTorch inference server

Hey HN, we built an inference server called NOS that can run a whole host of open-source AI models (LLMs, Stable Diffusion, CLIP, Whisper, Object Detection etc) all under one-roof. You can pull and run NOS in a few lines of code - simply `pip install torch-nos` and run `nos serve up --http`. You should now be able to talk to NOS over HTTP or gRPC (see README for examples). NOS can run locally on your desktop (with a gaming GPU), in any cloud GPU (L4, A100s, etc) and even on CPUs (without any acceleration). We’ll soon support Apple Silicon, so you should be able to run your AI models locally on a Macbook. Why are you building yet another inference server? Most API server implementations today deeply couple the API framework (FastAPI, Flask) with the modeling backend (PyTorch, TF etc.) - in other words, it doesn’t let you separate the concerns for the backend (i.e. scale-out, memory-efficiency, async/batched execution etc.) from the API (auth, observability, telemetry etc.), especially if you’re looking to build a production-ready application. Why use NOS? We’ve tried to make it very easy for developers to add support for new models and take them to production. Here are a few things we think developers care about: - Simple API over gRPC or REST that supports batched requests, and streaming. - Support any OSS model with custom runtimes with pip, conda and cuda dependencies. - Serve multiple custom models simultaneously on a single or multi-GPU instance. - Local execution means that you control your data, and you’re free to build NOS for domains that are more restrictive with data privacy. - Fully containerized means that you can develop, test and deploy NOS locally, on-prem, on any cloud or AI CSP. - Written entirely in Python, Apache-2.0 License. Try it out! Check out one of our demos in the NOS playground ( https://github.com/nos-playground ) (lots of video models) and let us know what you think!

Share card

Actual performance

3points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, apple · Missing: agents, macos, agent
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, para · Missing: reddit linkedin, podcasting, created
89%89% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, io · Missing: https docs, excited, just released
74%74% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: video, para · Missing: mobile apps, ios, personal
38%38% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: host, soon · Missing: plus, platform, intuitive
33%33% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
24%24% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Mi
Mighty Inference Server64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mighty Inference Server

Hacker News6
An
An RDMA/Infiniband Distributed Cache for Fast Inference and Training66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An RDMA/Infiniband Distributed Cache for Fast Inference and Training

Hacker News13
TQNN
TQNN55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fault-tolerant inference for noisy and imperfect data.

Indie Hackers1ai
Trieve Vector Inference
Trieve Vector Inference66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Deploy fast, unmetered embedding inference in your own VPC

Product Hunt+166API
gR
gRPC server for hnswlib – a fast approximate nearest neighbor search61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

gRPC server for hnswlib – a fast approximate nearest neighbor search

Hacker News3
Fa
Fast prefix search server in Golang51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fast prefix search server in Golang

Hacker News3
Inference Engine by GMI Cloud
Inference Engine by GMI Cloud74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fast multimodal-native inference at scale

Product Hunt+180Developer Tools
We
We just launched MegaAI. It's a 4k30fps, 4W, 4TOPS inference powerhouse69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We just launched MegaAI. It's a 4k30fps, 4W, 4TOPS inference powerhouse

Hacker News3
a
a Rust-based multimodal inference server70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

a Rust-based multimodal inference server

Hacker News1
LL
LLM Inference Requirements Profiler59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLM Inference Requirements Profiler

Hacker News4