Mo

Morph Reflexes – Multi-head classifiers for agent traces

Hacker News

Morph Reflexes – Multi-head classifiers for agent traces

The most common failures for production agents are behavioral: looping, reasoning leakage, user frustration, and more. Using a frontier model like GPT or Sonnet to judge every turn is too expensive and slow to run at scale. To solve this, we built Reflexes: semantic signals from agent traces, served fast and cheap over API. Built on custom kernels and a custom inference engine forked from vLLM. Under the hood, it is a small LLM architected around multi-head inference. Small models need to be trained for specific tasks, but running 50 separate small models on the same input for 50 tasks makes no sense. How it works: We use a modern LLM with hybrid attention and remove the decode step. We built an inference engine that lets prefill compute be 99% reused from reflex to reflex, similar in spirit to older 2019-era BERT/HYDRA and older multiple-head techniques. we built the inference engine to reuse the KV/cache across inputs and compute across all reflexes. One shared backbone reads the trace once, then many heads classify different signals. Our inference engine reuses the same KV/cache and compute across all reflexes, giving us sub-30ms inference with less than 0.1% overhead for each additional reflex. We took the same high-level idea and did the hard work to make it work with a modern architecture and attention. On it, we can run inference in under 30ms and serve the full request in under 90ms. If you run 4 reflexes or 100, the extra overhead is less than 2ms. Why does optimizing this matter? If you’re even a medium-sized startup, you’re dealing with tens of thousands of agent runs and millions of turns. If you want to track things like user frustration rates over time, frontier LLM-as-judge does not scale. I built a similar stack at Tesla. When ML engineers needed to sample data across petabytes for signals like `is_camera_obfuscated=true`, along with 200 other things, you need to 1) spin them up quickly 2) run at scale efficiently What it is not: A dashboard. 99% of dashboards go unused. 100% API first and made for devs who want to use this to trigger their own stuff. vibetrain a custom reflex in our dashboard, and/or then let it self improve in production: https://www.morphllm.com/dashboard/reflex Docs: https://docs.morphllm.com/sdk/components/reflexes/index I’d love feedback from people running agents in prod: what sorts of things do you wish you could track over time across 100% of turns but cant right now? TLDR: semantic signals from agent traces, super fast, cheap via API

Share card

Actual performance

20points
2comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
99%99% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: para, efficiently · Missing: supports, reddit linkedin, podcasting
90%90% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, io · Missing: https docs, excited, just released
71%71% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: para · Missing: mobile apps, ios, personal
53%53% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: efficient · Missing: plus, platform, intuitive
36%36% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
30%30% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Mo
Move a Cube With Your Head or HeadTracking With WebGL62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Move a Cube With Your Head or HeadTracking With WebGL

Hacker News2
head spa
head spa15%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

head spa

Indie Hackers
Ga
Games Head-to-Head50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Games Head-to-Head

Hacker News9
Ex
Experimental multi-head window manager made for programmatic use51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Experimental multi-head window manager made for programmatic use

Hacker News1
Tu
Turnabout - an iOS puzzler I'm still trying to wrap my head around35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Turnabout - an iOS puzzler I'm still trying to wrap my head around

Hacker News1
My
My 3D Printed T-Rex Shower Head45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

My 3D Printed T-Rex Shower Head

Hacker News3
Fintech Battles™
Fintech Battles™45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The world’s top fintechs go head-to-head

Indie Hackers2design
Co
CompEngine: Head-to-Head videos competitions for extreme sports37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CompEngine: Head-to-Head videos competitions for extreme sports

Hacker News2
Tu
Turning MNIST on Its Head – A NN That Counts53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Turning MNIST on Its Head – A NN That Counts

Hacker News4
Cap
Cap28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

caps to head

Indie Hackerscommitment-full-time