LL

LLGTRT: TensorRT-LLM+Rust server w/ OpenAI-compat and Structured Output

Hacker News

LLGTRT: TensorRT-LLM+Rust server w/ OpenAI-compat and Structured Output

I wanted to share a project I've been working on for the past few weeks: llgtrt. It's a Rust implementation of a HTTP REST server for hosting Large Language Models using llguidance library for constrained output with NVIDIA TensorRT-LLM. The server is compatible with the OpenAI REST API and supports structured JSON schema enforcement as well as full context-free grammars (via Guidance). It's similar in spirit to the Python-based TensorRT-LLM OpenAI server example but written entirely in Rust and built with constraints in mind. No Triton Inference Server involved. This also serves as a demo for the llguidance library, which lets you apply sampling constraints via Rust, C, or Python interface, with minimal generation overhead and no startup cost. Any feedback or questions are welcome!

Share card

Actual performance

6points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, compatible · Missing: reddit linkedin, podcasting, created
76%76% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, context · Missing: mac, agents, macos
75%75% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: io · Missing: https docs, excited, just released
68%68% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: host, interface · Missing: plus, platform, intuitive
36%36% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Mo
Motörhead, LLM Memory Server Built in Rust65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Motörhead, LLM Memory Server Built in Rust

Hacker News16
As
Async Rust Library for OpenAI51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Async Rust Library for OpenAI

Hacker News2
LL
LLMdantic: Structured Output Is All You Need38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLMdantic: Structured Output Is All You Need

Hacker News8
St
Structured output from LLMs without reprompting63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Structured output from LLMs without reprompting

Hacker News174
Ty
Tysm, a Rust crate for using ChatGPT with Structured Outputs42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tysm, a Rust crate for using ChatGPT with Structured Outputs

Hacker News7
Ro
Rooster – Personal Web Server with Rust60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rooster – Personal Web Server with Rust

Hacker News10
xm
xmllm – Structured LLM streaming output using lenient XML parsing60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

xmllm – Structured LLM streaming output using lenient XML parsing

Hacker News5
St
Structured, composable logging for Rust62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Structured, composable logging for Rust

Hacker News1
An
An example GraphQL server written in Rust80%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An example GraphQL server written in Rust

Hacker News74
A
A minimal implementation of LLM output watermarking35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A minimal implementation of LLM output watermarking

Hacker News2