hist: An overengineered solution to `sort|uniq -c` with 25x throughput
hist: An overengineered solution to `sort|uniq -c` with 25x throughput
Was sitting around in meetings yesterday and remembered an old shell script I had to count the number of unique lines in a file. Gave it a shot in rust and with a little bit of (over-engineering)™ I managed to get 25x throughput over the naive approach using coreutils as well as improve over some existing tools. Some notes on the improvements: 1. using csv (serde) for writing leads to some big gains 2. arena allocation of incoming keys + storing references in the hashmap instead of storing owned values heavily reduced the number of allocations and improves cache efficiency (I'm guessing, I did not measure). There are some regex functionalities and some table filtering built in as well. happy hacking
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
TDD Screencast #2: Implementing Sort
Cluster-Sort
QuadSort, Esoteric Fast Sort
Faster than std:sort and pdqsort
YouTube Sort by Likes
vim-sort-folds – Sort vim folds based on their first line
KillerTask: the solution to AsyncTask implementation (in Kotlin)
0patch, vulnerability micropatching solution
Metis - A Divine Solution for Unix Aliases
My solution to scaling P2Pool