Bh

Bhumi–OSS Python Library w Rust Underhead for 2.5x Faster LLM Inference

Hacker News

Bhumi–OSS Python Library w Rust Underhead for 2.5x Faster LLM Inference

Read the full blogpost at https://rach.codes/blog/Introducing-Bhumi (click on reader to see the technical breakdown!) AI inference should be fast, but in practice it’s painfully slow. Inference bottlenecks slow down LLM-powered chatbots and AI workflows everywhere. I built Bhumi to fix that. Bhumi is a Python library designed for developers, yet its performance-critical core is implemented in Rust (via PyO3) for near-native speed. This hybrid approach delivers up to 2.5x faster response times across providers like OpenAI, Anthropic, and Gemini—without changing the underlying model. THE PROBLEM: SLOW AI INFERENCE Most LLM clients suffer from three main issues: 1. Batch Processing Overhead – Clients wait for the full response instead of streaming data as it’s ready. 2. Inefficient Buffers – Default buffer sizes aren’t tuned for AI-generated text. 3. Validation Bottlenecks – Tools like Pydantic slow down structured response handling. Bhumi tackles these challenges with a smarter architecture that blends Python’s ease of use with Rust’s raw speed. HOW BHUMI MAKES AI FASTER 1. Rust-Based Streaming: Python’s async is useful, but integrating Rust through PyO3 brings near-native performance. Streaming inference starts instantly, cutting response times by up to 2.5x. 2. Smarter Buffer Management: Quality-Diversity algorithms (like MAP-Elites) dynamically discover optimal buffer sizes, boosting throughput by roughly 40%. 3. Replacing Pydantic with Satya: Pydantic was a performance sink. I built Satya—a Rust-backed validation library—that accelerates structured outputs dramatically. PERFORMANCE BENCHMARKS: • OpenAI: 2.5x faster response times • Anthropic: 1.8x faster • Gemini: 1.6x faster • Minimal extra memory overhead Bhumi is provider-agnostic, allowing you to switch between OpenAI, Anthropic, Groq, and more with a simple config change. USING BHUMI (WITH TOOL USE & STRUCTURED OUTPUTS) Bhumi makes tool integration effortless. For example, here’s how you can register a weather tool in Python: import asyncio from bhumi.base_client import BaseLLMClient, LLMConfig async def get_weather(location: str) -> str: return f”The weather in {location} is 75°F” async def main(): config = LLMConfig(api_key=“sk-…”, model=“openai/gpt-4o-mini”) client = BaseLLMClient(config) client.register_tool(name=“get_weather”, func=get_weather) response = await client.completion([{“role”: “user”, “content”: “What’s the weather in SF?”}]) print(response[“text”]) asyncio.run(main()) WHAT’S NEXT? I’m actively working on: • Supporting More Providers & Models • Adaptive Streaming Optimizations • Advanced Structured Outputs & Tooling Bhumi is a Python-first library powered by a Rust underhead for performance. Check out Bhumi on GitHub at https://github.com/justrach/bhumi or reach out at me@rachit.ai.

Share card

Actual performance

8points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
92%92% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, user, models · Missing: mac, agents, macos
75%75% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Strong signals: efficient · Missing: plus, platform, intuitive
60%60% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, io · Missing: https docs, excited, just released
54%54% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
49%49% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat, smart · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Py
Python JSON library in Rust, faster than ujson63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Python JSON library in Rust, faster than ujson

Hacker News40
(2
(29x faster)Rapidvalidators - Python's validators lib rewritten in Rust49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

(29x faster)Rapidvalidators - Python's validators lib rewritten in Rust

Hacker News2
Fa
Fast spectogram library in python and rust61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fast spectogram library in python and rust

Hacker News4
Te
TensorBase: 5x~10000x Faster Drop-In/Accelerator for ClickHouse in Rust76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TensorBase: 5x~10000x Faster Drop-In/Accelerator for ClickHouse in Rust

Hacker News19
Ru
Rust (PyO3) deserializer callable from Python55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rust (PyO3) deserializer callable from Python

Hacker News1
Fl
FlashText with Rust for Python70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

FlashText with Rust for Python

Hacker News6
Sp
Spice.ai OSS 1.0 – data query and AI-inference engine built in Rust52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Spice.ai OSS 1.0 – data query and AI-inference engine built in Rust

Hacker News26
rt
rtoml – a TOML library for Python written in rust67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

rtoml – a TOML library for Python written in rust

Hacker News5
Ru
Rust Powered Inference, Ingestion and Indexing with EmbedAnything74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rust Powered Inference, Ingestion and Indexing with EmbedAnything

Hacker News1
Ch
Charset Normalizer library – port from Python to Rust53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Charset Normalizer library – port from Python to Rust

Hacker News4