a Rust-based multimodal inference server
a Rust-based multimodal inference server
We built a production-grade multimodal inference server in Rust for serving vision–language models (image + text → streamed text). The goal was to explore what a Rust-native control plane looks like for modern multimodal inference: continuous batching, KV-aware admission control, predictable behavior under load, and proper streaming semantics. The system exposes an OpenAI-compatible API, supports multi-image inputs, and is designed to degrade gracefully under overload rather than OOM or stall. It’s organized as a single monorepo with a gateway, GPU workers, scheduler, and pluggable engine adapters. We’ve also included a benchmark suite focused on real-world scenarios (TTFT, cancellation, overload, fairness) rather than synthetic tokens/sec numbers. Would love feedback from folks building or operating inference infrastructure.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Rust Powered Inference, Ingestion and Indexing with EmbedAnything
Mighty Inference Server
VLM Inference Engine in Rust
NOS – A fast, and ergonomic PyTorch inference server
Orkhon – ML Inference Framework and Server Runtime (Rust and Python)
Rooster – Personal Web Server with Rust
MetricWhiz a Rust based TUI for NLP binary classification analysis
An example GraphQL server written in Rust
Warp, a Rust-based terminal
Godot and Rust based multiplexer (terminal panes and more)