I

I built a hypervisor and client for inference on consumer compute

Hacker News

I built a hypervisor and client for inference on consumer compute

I'm the founder of Scalattice, this is my second company, third total product. I'm a 2x founder building some challenging software, some easy software, and some curiosity based tools that I've just always wanted to be a part of! So here is Scalattice. Scalattice is an OpenAI-compatible inference API. You keep the OpenAI SDK, swap base_url + API key, and call open models (Qwen3, Llama 3.3 70B, Gemma 3, DeepSeek R1, etc.). Inference is performed by a provider node on our distributed network of consumer hosted inference machines, we run an open source agent ( https://github.com/scalattice/scalattice-agent ) which works in Rust to run the inference and get the provider paid. In short, anyone can become a provider, offer up their machine with our Windows/Linux Rust agent, and earn some extra cash, or start a farm of machines to make big bucks. I decided to build some innovative behavioural traits to the API for higher performance/security: 1. output vetting - Scalattice Cloud vets the response against replica responses requested on the API header N times. 2. security tiers - Using some split inference, the job is completed in chunks by multiple providers and in part by Scalattice Cloud hypervisor for added security/privacy (at an additional cost) 3. regional policy - I know how important data residency is to developer clients (from my experience with Digital ID Infrastructure) and so I built the platform to let developers specify the region for inference. We pay our providers a majority share (Currently 80%) of every token spent by the developer. We currently back the network with a failover of our own company machines which ensures we never drop a request. If you want to give it a try in a couple of minutes, we are currently running a "Top Up $10 for UNLIMITED QWEN3" for an entire month of API calling, (terms and conditions apply). If you want to give it a try in a couple of minutes: 1. Create a key at https://scalattice.cloud/developers 2. Use the copy-paste snippet on the linked post 3. Live rates: https://scalattice.com/pricing Or, if you want to be a provider (We are really keen to onboard people and get them earning): 1. Create a machine at https://scalattice.cloud/providers 2. Download our Windows app, or use our on-page instructions to curl the agent installation script for Linux. 3. Attach a provider token to the app or Linux agent from the Machine created on the website. 4. Control the machine from the website, select which models you want to offer, and the physical hardware components to use for inference and the hours of operation. 5. Earn! Happy to take feedback - especially on DX, pricing clarity, and what would make you try this instead of Together / Fireworks / OpenRouter / Salad etc. This is my first time posting on HN, please let me know if you think the product has potential, there are obviously trade-offs with non-datacentre inference (latency, capability of cards), but I think this could serve a decent amount of developers very well, because not everyone needs H100 cards for their inference. Thanks again.

Share card

Actual performance

3points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: created, ios, compatible · Missing: supports, reddit linkedin, podcasting
96%96% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, agent, model · Missing: agents, macos, cursor
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, llama, ide · Missing: https docs, excited, just released
50%50% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: ios, month, way · Missing: mobile apps, personal, entrepreneurs
43%43% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, host · Missing: plus, intuitive, reviews
30%30% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: paid · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ke
Keen Compute37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Keen Compute

Hacker News1
ZeroGPU
ZeroGPU61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The compute efficient layer for AI inference

Product Hunt+307API
I
I built a lite LPU that can do inference on Karpathy's MicroGPT62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built a lite LPU that can do inference on Karpathy's MicroGPT

Hacker News18
TQNN
TQNN55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fault-tolerant inference for noisy and imperfect data.

Indie Hackers1ai
oL
oLLM – LLM Inference for large-context tasks on consumer GPUs58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

oLLM – LLM Inference for large-context tasks on consumer GPUs

Hacker News3
I
I built a simple HN client in Deno46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built a simple HN client in Deno

Hacker News3
We
We just launched MegaAI. It's a 4k30fps, 4W, 4TOPS inference powerhouse69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We just launched MegaAI. It's a 4k30fps, 4W, 4TOPS inference powerhouse

Hacker News3
I
I built a fancy ChatGPT client39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built a fancy ChatGPT client

Hacker News3
Honestore
Honestore45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The community of consumer activists

Indie Hackers1community
Co
Compute polynomials twice as fast45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Compute polynomials twice as fast

Hacker News139