Sw

Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

Hacker News

Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with a bit of secret sauce (hint On-Policy Distillation) and achieved great results (-58% length, 1.95x speed up, <1% accuracy loss) so we wanted to open-source it and hear the feedback of the community. This is the link to the model: https://huggingface.co/ukisai/Swift-Qwen3.8-27b / It's at 80k downloads in 3 days with independent evals here: https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai_... We are also providing a Free Research Purpose API (OpenAI compatible), courtesy of Nvidia who were kind enough to provide us with the GPUs. It's limited at 5RPM. https://ukisai.com/api/swift/v1/models We also made a GGUF (Q1-Q8): https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF and there's also a few nice community quants with even lower/higher precision (Bartowski: https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GG... ). The community also created amazing MLX, NVFP4, W4A16 and Uncensored versions you can find on Huggingface. IMPORTANT: Our training approach is not a replacement for the reasoning effort settings, chat templates or token caps but is complementary and targets a completely separate issue (overthinking and "anxiety-like" reasoning loops prior seen in PTQ, but as far as we identified also prominent in BF16 of this size class LLMs as well). Contrary to popular belief, these specific patterns do not contribute to answer quality when properly targeted. (our thesis being: reasoning length IS extremely important and should NOT be shortened by force, but rather optimized). This is also demonstrated bellow in our xhigh vs medium effort benchmark. The goal is to keep xhigh accuracy while reducing only the unnecessary part of thinking. The TLDR of our thought process, research, training and a link to the Meta paper that inspired us is in the first comment. The benchmarks: Swift-27B vs Qwen3.8-27B (BF16, all benchmarks ran x5, thinking effort xhigh) GPQA-Diamond: 88.4% -> 88.3%, 58% fewer median tokens LiveCodeBench v6: 76.8% -> 81.6% (+4.8pp, due to default truncation in LCB it is not performance gain), 46% fewer median thinking tokens Terminal-Bench 2.1: 66.7% -> 65.8%, 39% fewer median tokens MMLU-Pro: 85.5% -> 85.0%, 28% fewer median tokens C-Eval: 90.0% -> 90.6%, 19% fewer median tokens IFBench: 73.5% -> 71.8%, 51% fewer median tokens AIME 2026: 98.7% -> 94.0%, 50% fewer median tokens HMMT (Nov 2025): 99.3% -> 96.0%, 46% fewer median tokens ERQA (vision): 67.5% -> 66.3%, 55% fewer median tokens Token savings hold at every reasoning effort (mean thinking reduction): xhigh 41%, medium 23%, low 26% (albeit with accuracy loses of 1-4% on medium and 1-2% on low which we need further testing for) Swift at xhigh vs the base's own effort settings on GPQA-Diamond (198 questions x 5 seeds): Base xhigh: 88.4%, 6,642 median tokens Swift xhigh: 88.3%, 2,771 median tokens Base medium: 84.1%, 1,753 median tokens So Swift keeps xhigh accuracy at under half the tokens, and beats base-medium by 4pp at roughly 1.6x its tokens. End note: While we are keen on complete open-source, we still need to keep a part of our training and data private, being a new lab. The license is not Apache 2.0, but it only affects companies >$1M. We hope this does not pose a problem for the community and are open to feedback on it. We want to contribute as much as possible to the community and would really appreciate feedback on our work, quantization or Swift model requests. We are working on Swift 3.8 Flash Next and have so far gotten up to -53% thinking token usage. We have strong indicators our methodology is reproducible on other model families as well and are asking the community which ones you want us to optimize next.

Share card

Actual performance

29points
17comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: created, para, compatible · Missing: supports, reddit linkedin, podcasting
96%96% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: llama, ide, io · Missing: https docs, excited, just released
65%65% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: para · Missing: mobile apps, ios, personal
54%54% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: efficient · Missing: plus, platform, intuitive
26%26% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
21%21% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ph
PhoneNumberKit – a Swift take on libphonenumber43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PhoneNumberKit – a Swift take on libphonenumber

Hacker News4
Au
Autocompletion for Swift on Emacs54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Autocompletion for Swift on Emacs

Hacker News79
A
A Lisp Interpreter in Swift65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A Lisp Interpreter in Swift

Hacker News3
At
Attributed strings in Swift 426%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Attributed strings in Swift 4

Hacker News28
Tr
Trailer Swift43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Trailer Swift

Hacker News32
Po
Populating a UITableView in Swift54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Populating a UITableView in Swift

Hacker News1
20
2048 in Swift44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

2048 in Swift

Hacker News92
Cl
Clojure-ish Lisp interpreter implemented in Swift72%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Clojure-ish Lisp interpreter implemented in Swift

Hacker News2
Ge
Get to grips with NSTimer in Swift43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Get to grips with NSTimer in Swift

Hacker News3
Sw
Swift Playgrounds on steroids43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Swift Playgrounds on steroids

Hacker News1