To

Tokenkit – Convert LLMs to new tokenizers (incl byte-level Llama/Gemma)

Hacker News

Tokenkit – Convert LLMs to new tokenizers (incl byte-level Llama/Gemma)

Hi HN! I’ve been working on transferring LLMs across tokenizers using a new method called ALM (details in our paper [1]). It distills a model trained with one tokenizer into a version using another, enabling things like converting subword models into byte-level ones much more effectively than what was possible before. To make this easy to use, I released tokenkit, a library implementing ALM and other tokenizer transfer methods: https://github.com/bminixhofer/tokenkit . As a demo, I used ALM to create two byte-level instruction-tuned models: - https://huggingface.co/benjamin/Gemma2-2B-IT-Byte - https://huggingface.co/benjamin/Llama3-2-3B-IT-Byte Even though the distillation phase is very short (just 1.2B bytes ≈ 330M subword tokens), the models perform competitively (for example 57.0% MMLU of the byte-level Llama vs. 62.4% MMLU of the original Llama3-3B-Instruct). This approach opens up an interesting direction: we can potentially keep subword tokenization for pretraining (to still squeeze as much text into the model in as little time as possible), but then change to a more user-friendly tokenization afterwards. These models aren’t yet optimized for efficiency, but if you would add self-speculative decoding plus a BLT/DTP-style hierarchical architecture and/or linearized attention, they might also be able to replace subword-based models when speed matters. If you're interested in training your own models, the guide on tokenizer transfer via tokenkit [2] should make it easy. The model cards of the transfers above also contain the exact command used to train them. I’ve been training on fairly limited hardware, so effective transfer is possible even in a (near) consumer-grade setup. [1]: https://arxiv.org/abs/2503.20083 [2]: https://github.com/bminixhofer/tokenkit/blob/main/docs/token...

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, user, new · Missing: mac, agents, macos
80%80% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
76%76% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: llama, ide, io · Missing: https docs, excited, just released
66%66% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Strong signals: plus, friendly · Missing: platform, intuitive, reviews
53%53% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
44%44% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Kr
KraspAI Kompass – keep up with new LLMs53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

KraspAI Kompass – keep up with new LLMs

Hacker News1
Ll
Llama or Alpaca?78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama or Alpaca?

Hacker News6
Ll
Llama 3.2 Interpretability with Sparse Autoencoders74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 3.2 Interpretability with Sparse Autoencoders

Hacker News579
Llama 3.3 70B
Llama 3.3 70B68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 405B-level performance, at a fraction of the cost

Product Hunt+347Open Source
LL
LLaMA Nuts and Bolts, A holistic way of understanding how LLMs run76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

LLaMA Nuts and Bolts, A holistic way of understanding how LLMs run

Hacker News3
Sw
Swiss Army Llama – A Versatile, FastAPI-Based Multitool for Local LLMs49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Swiss Army Llama – A Versatile, FastAPI-Based Multitool for Local LLMs

Hacker News1
Ll
Llama 2 Uncensored 70B as API79%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Llama 2 Uncensored 70B as API

Hacker News18
Fi
Finetune Llama-3.1 2x faster in a Colab74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Finetune Llama-3.1 2x faster in a Colab

Hacker News16
On
OnPanda – Steer LLMs and agents at the token level63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OnPanda – Steer LLMs and agents at the token level

Hacker News5
Schedulist 2.0
Schedulist 2.065%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Unleashing a new level of productivity

Product Hunt+95Android