AI

AI that can use Gmail, SMS, Slack, Calendar

Hacker News

AI that can use Gmail, SMS, Slack, Calendar

Hello HN, Dawson and Ethan here from Martin ( https://trymartin.com ). We’ve been building an AI personal assistant (the elusive dream of a real life Jarvis) for about a year now, and we recently launched Martin as a web app. Watch our latest demo here: https://youtu.be/ZeafVF8U7Ts . We’re starting with common agentic tasks for consumers/prosumers - Martin can read/draft emails, make your calendar, text and call others for you, and use Slack. Like any personal assistant, it can also set reminders, track your to-dos, and send you daily briefings. The idea is to eventually tackle everything that an on-call virtual assistant does. 4 months ago, we did a launch HN for Martin's voice-first iOS app. A big piece of product feedback we got was "I don't trust AI to take actions like sending texts/emails on my behalf if it's not 100% reliable." We’re happy to report that Martin's failure rate is now a lot lower than before (though we have a lot more work to do for more complex actions). We have tackled some pretty interesting problems since our last launch, so thought we’d share a couple of them here: First, building a testing suite to concretely measure and improve performance for agents is no trivial task. (We're optimistic that someone might build an awesome system for this one day, but we haven't found one so are doing it ourselves.) Specifically, what we’d like to do is run existing test cases on new implementations of our entire LLM processing flow - not just new prompts - and be able to rigorously say whether we’ve improved and/or where we’ve regressed. This means defining tests in such a way that they’re resilient to major overhauls of code structure, as well as building a testing execution context that mimics production behavior (i.e. a test user with calendar events, emails, contact info). On top of that, all test cases need to be manually and painstakingly written, with expected outputs sometimes being many tens of thousands of characters. On the monitoring side, most of our reliability issues are soft errors which are very hard to programmatically catch. When malfunctions happen, most of the time we learn of it through customer feedback and not any conventional third-party monitoring system. The best we can do without manually sifting through tons of data is to implement rudimentary checks based on behavior patterns which we know historically indicate errors (e.g. making many similar API calls in quick succession, implying rapid failure and retry of function calls) Another problem we keep coming back to is the stateless nature of LLM context (information is not stored latently and needs to be reintroduced at every invocation). Because of how much info Martin needs (product information, user memory, tool definitions, previous messages, platform-specific instructions, etc), we need to carefully manage what information we expose to Martin and how we balance broad context with specific information. Vanilla RAG can’t handle the complexity, so we built custom retrieval and context injection systems for each LLM call. We abstract away some information behind function calls and organize certain tools into modules which share context and instructions. This strategy has helped a lot with reliability. Of course, we're still a long way from Jarvis. Whenever one of us struggles with a technical problem, the other will kindly remind him that "Tony Stark built this in a cave, with a box of scraps!" We’re super pumped about where software is headed. It feels like we’re tinkering with ideas that are on the edge of what’s possible. You can try out Martin on desktop and iOS at https://trymartin.com . We have a 7-day free trial, and if you find it useful we charge $35/month afterwards for unlimited usage. Very excited to hear your thoughts! If you have any ideas around reliability for agents or the future of consumer AI interfaces, we’d love to discuss and trade notes.

Share card

Actual performance

12points
12comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: ios · Missing: supports, reddit linkedin, podcasting
96%96% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, agentic · Missing: mac, macos, cursor
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: ios, personal, month · Missing: mobile apps, entrepreneurs, apps
57%57% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: excited, exist, existing · Missing: https docs, just released, lua
54%54% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, interface, calls · Missing: plus, intuitive, reviews
37%37% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
20%20% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: introduce, real life · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Us
Use Slack for Everything53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Use Slack for Everything

Hacker News60
Bo
Boomerang Calendar - a smart calendar assistant in your Gmail48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Boomerang Calendar - a smart calendar assistant in your Gmail

Hacker News86
Fo
Forgetful Fairy – SMS Calendar Reminders33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Forgetful Fairy – SMS Calendar Reminders

Hacker News2
Clerk.Chat
Clerk.Chat62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SMS for Slack - enable 2-way SMS conversations inside Slack

Indie Hackers13$2,600/moapis
Lo
Locent – Mailchimp for SMS52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Locent – Mailchimp for SMS

Hacker News90
Wo
Wolfrasms, a Wolfram/SMS mashup50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Wolfrasms, a Wolfram/SMS mashup

Hacker News5
Ha
Hacking SMS Voting63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hacking SMS Voting

Hacker News18
Ha
HackerNews via SMS62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

HackerNews via SMS

Hacker News3
Wh
Whimsical – Anything done via SMS52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Whimsical – Anything done via SMS

Hacker News4
Wh
Whimsical – Get anything done, via SMS52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Whimsical – Get anything done, via SMS

Hacker News23