Da

Datetime-bench: which datetime formats LLMs get right (and wrong)

Hacker News

Datetime-bench: which datetime formats LLMs get right (and wrong)

tl;dr * If you need an LLM to parse OR emit a timestamp, use: RFC 3339 ( e.g. 2024-03-26 10:30:00-05:00 ) * python date format also works well * Do NOT use unix epoch or javascript date formats. * Smaller models and non-reasoning models still make a LOT of mistakes in time parsing / formatting. --- There are lots of temporal reasoning benchmarks (like TimeBench, TRAM, etc.) but they test whether models understand time concepts. Nothing on which datetime output format models get right most often. So we just built the benchmark ourselves. We tested 22 models across Google, Anthropic, OpenAI, Qwen, and GLM on 235 scenarios and 7 different formats. The two that surprised us the most were JavaScript Date and unix epoch. JavaScript Date is probably the most commonly used format and it's wrong ~1 in 4 times on parsing. Unix epoch drops to 40% on arithmetic tasks. If you need epoch, just have the model output a string and convert it yourself in code.

Share card

Actual performance

2points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, google, models · Missing: mac, agents, macos
90%90% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: mistakes, ios · Missing: supports, reddit linkedin, podcasting
88%88% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: io · Missing: https docs, excited, just released
60%60% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
37%37% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: ios, google · Missing: mobile apps, personal, entrepreneurs
33%33% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
13%13% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

GP
GPTCache – Redis for LLMs69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

GPTCache – Redis for LLMs

Hacker News7
pr
prompttest – pytest for LLMs34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

prompttest – pytest for LLMs

Hacker News2
Qq
Qq: like jq, but can transcode between many formats62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Qq: like jq, but can transcode between many formats

Hacker News115
La
Labelformat now supports all major vision labeling formats32%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Labelformat now supports all major vision labeling formats

Hacker News1
We
WeatherTunes — The right tunes for right now56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WeatherTunes — The right tunes for right now

Hacker News32
Fi
Find the right syntax for your YAML multiline strings35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Find the right syntax for your YAML multiline strings

Hacker News1
Ta
Taskr - "Tasking Done Right"38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Taskr - "Tasking Done Right"

Hacker News8
“Y
“You’re absolutely right” on a tee54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

“You’re absolutely right” on a tee

Hacker News2
Wo
Wordchum – Lurn 2 Right56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Wordchum – Lurn 2 Right

Hacker News2
We
Web-Based Inspectors for Binary Formats54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Web-Based Inspectors for Binary Formats

Hacker News5