Ca

Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash

Hacker News

Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash

Hey HN, Henry from Cactus here. We submitted Needle 2 here a few weeks ago, and the feedback in the discussion thread was incredibly valuable, thanks! Thanks to all that feedback, we’ve been able to move quickly to release Needle 3 and I'd love to hear what you think again. The key features: 1) Automation (tool calls & structured JSON output): Needle still doesn't chat by design, its quite challenging to pack general capacity into such small models, so we focus on tool calls and structured JSON. If no tool you declared fits the request, you get an empty list back (note for when playing with the demo). 2) Intelligence Laddering: Every layer (2 to 20) is a deployable subnetwork, so one set of weights, 25 to 121 million parameters at 2-bit, shipping as 8-29MB binaries. On a Raspberry Pi 5 it decodes at up to 4k tokens/sec and prefills at up to 10k. 3) Monarch Hadamard MLP: replaces the dense FFN with three learnable Walsh-Hadamard-initialized Kronecker (Monarch) factor pairs interleaved with per-channel diagonal scales, fixed permutations, a SiLU nonlinearity, and a rank-8 input-conditioned gate, so each token gets a fully mixed nonlinear transform of its d_model channels at O(d√d) parameters and compute instead of the O(d²) a dense 4x-expansion MLP would cost. 4) Performance: On Mobile Actions (phone commands, scored on the exact call) the 20-layer model gets 86.0 through the shipped 2-bit binary; LFM2.5 1.2B is at 82.4, Qwen3.5 0.8B at 76.0, Apple's on-device model at 57.6, all at f16. More results on the link, we do not win everywhere ofc. 5) Multilingual: Needle 3 now supports English, French, Spanish, German, Dutch, Italian, Polish, with more languages coming. 6) Finetuning: You can achieve DeepSeek v4 Flash grade performance on a narrow task with just 4L, stress on "narrow task", we found that production users often prefer tuning before production. 7) Triggers: Grounding is a common challenge for tool call, at least for Needle 2, so we added support case-insensitive regular expressions matched against each request to gate false negatives. 8) Confidence: Every response also carries a calibrated confidence score, the minimum of a judgement on the finished call and its decode probability. Act above your threshold, show the call and ask below it, or escalate to a bigger model. 9) Supported Platforms: macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32, Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly, and a WASI component. Thanks for reading and as always, thoughts appreciated!

Share card

Actual performance

5points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, para, ios · Missing: reddit linkedin, podcasting, created
91%91% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, macos, model · Missing: agents, agent, cursor
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: ios, users, way · Missing: mobile apps, personal, entrepreneurs
50%50% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, users, calls · Missing: plus, intuitive, reviews
49%49% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, ide, io · Missing: https docs, excited, just released
41%41% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Wa
Warp – Run DeepSeek v4.1 Flash with 5 GB of RAM at 3.77 tok/s63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Warp – Run DeepSeek v4.1 Flash with 5 GB of RAM at 3.77 tok/s

Hacker News9
DeepSeek-V4-Flash-0731
DeepSeek-V4-Flash-073177%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Frontier agent intelligence at Flash prices

Product Hunt+324API
De
DeepSeek Flash inverted the economics of agent products48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DeepSeek Flash inverted the economics of agent products

Hacker News9
g023 Code
g023 Code63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DeepSeek V4 Flash Optimized Harness w/Ollama vision support

Product Hunt+2
Hu
Hue v4.6.153%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hue v4.6.1

Hacker News2
DD
DDNS v447%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DDNS v4

Hacker News2
Be
Benthos v4.0.047%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Benthos v4.0.0

Hacker News3
Ce
Centrifugo v4 Released47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Centrifugo v4 Released

Hacker News6
DS PHP Edit: DeepSeek v4 Powered
DS PHP Edit: DeepSeek v4 Powered39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Self-hosted, DeepSeek V4-API powered, and stubbornly local

Indie Hackers1ai
Op
Optimizing DeepSeek's NSA for TPUs – A Kernel Worklog48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Optimizing DeepSeek's NSA for TPUs – A Kernel Worklog

Hacker News2