Jinfer – AI inference engine for the JVM. AI in a jar
Jinfer – AI inference engine for the JVM. AI in a jar
Hi HN, Some would call AI on the JVM quixotic. And so, Quixotic AI was born. jinfer is an inference engine for the JVM: chat, vision, audio, embeddings, reranking, and TTS. No Python runtime, no ONNX, no Docker containers, no sidecar process, no IPC. Finally, AI in a jar. The stack underneath is built for the JVM rather than bolted onto it: toknroll: pure-Java tokenizers, zero dependencies gguf / safetensors: read and write llama.cpp and HuggingFace model formats jam: quantized matmul kernels, competitive with llama.cpp on CPU jota: Tensor API targeting Java, C, CUDA, HIP, Metal, OpenCL, and Mojo It ships with integrations for Spring AI and LangChain4j, and is compatible with GraalVM Native Image for low-overhead, self-contained binaries with millisecond startup. This is an early release (CPU only): I'd especially like feedback on the API surface. Site: https://qxotic.ai Repo: https://github.com/qxoticai/qxotic
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
VLM Inference Engine in Rust
Fast multimodal-native inference at scale
Fault-tolerant inference for noisy and imperfect data.
We just launched MegaAI. It's a 4k30fps, 4W, 4TOPS inference powerhouse
LLM Inference Requirements Profiler
Differential Privacy Inference for Julia
Neuropod – Uber ATG's open source deep learning inference engine
Mighty Inference Server
Jlama – A fast Java inference engine for GPT and Llama models
Melange - pegging AI inference to the cost of the most expensive model