Saving Money Deploying Open Source AI at Scale with Kubernetes
Saving Money Deploying Open Source AI at Scale with Kubernetes
Hi HN: I wanted to share this piece I wrote on how I saved our small startup 10s of thousands of dollars every month by lifting and shifting or AI data-pipelines from using OpenAI's API to a vLLM deployment ontop of Kubernetes running on a few nodes with T4 GPUs. I haven't seen alot on the "AI-DevOps" or infrastructure side of actually running an at-scale AI service. Many of the AI inference engines that offer an OpenAI compatible API (like vLLM, llama.cpp, etc.) make it very approachable and cost effective. Today, this vLLM AI service handles all of our batching micro-services which scrape for content to generate text on over 40,000+ repos on GitHub. I'm happy to answer any / all questions you might have!
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Open-source AI assistant for Kubernetes troubleshooting
Monetizable open source AI
Extrapolate – Open-Source AI Aging App
I made an open-source AI Headshot Generator
I built an open-source AI system for drones
InfiniteGpu, An open-source AI computational network
The open-source AI alternative to Gong
CopilotTextarea/> = an open-source, AI-infused react <textarea>
The Internet's Open Source AI Paywall
Open-source AI pentester with exploit-verified findings