GP

GPTinker – AI developer sidekick that can work on existing codebases

Hacker News

GPTinker – AI developer sidekick that can work on existing codebases

GPTinker is an experimental AI junior developer, my weekend project (with code available at [1]) designed to test the limits of end-to-end coding using GPT-4. TL;DR: Here's a 7-minute video showing the app in action: https://youtu.be/XgMKCeiUDQc Although it has already been demonstrated that GPT models can write new code and even modify existing code based on a prompt, the most potential lies in interacting with existing codebases—an area that I believe has not yet been widely explored. The aim of this project is to see how far we can take this idea and gauge its usefulness. At present, GPTinker is a proof of concept that I've put together in just a few days. While it is currently constrained to Typescript repositories (although this could be easily changed), it can already do some pretty neat things: * Modify existing components and write new ones based on current ones, while maintaining the app's style * Refactor components by extracting parts of the component code into separate components and then integrating the new component where the old code used to be * Make project-wide changes, such as performing a search and modifying all component invocations to support new props or changing their class names * Write unit tests for components and business logic, run them, and fix any issues * Debug and fix configuration issues, install missing dependencies, and prompt you for how to proceed when running a command fails * Perform other tasks that I haven't yet thought of testing GPTinker achieves all this by utilising ideas from Toolformer and LangChain. The model is asked to split the given task into a series of discrete actions and then execute them one by one using one of the built-in commands (tools), such as ListFiles, ReadFile, WriteFile, RunCommand, etc. It writes out JSON that gets parsed by the backend, which then executes the command with the given parameters and feeds the output back to the model. This process continues in a loop until the model decides the task is complete and doesn't emit any new commands. I've already noticed some shortcomings and limitations of this method and GPT-4 in general: * The 8k token limit prevents editing larger files and having longer conversations. However, with the 32k model on the horizon and by implementing some optimizations, this limitation can likely be addressed * Writing out larger files can be slow. This will undoubtedly improve over time (hoping for a "turbo" model soon), and if I manage to implement a PatchFile command so that the model emits only the diffs, it will speed things up and cut down on token usage * Ability to reason about your code—while GPT-4 is generally good at understanding what your code does and how to apply changes to it, its performance can sometimes be hit or miss, depending on the tasks you throw at it. Think of it as a junior-level developer with lots of patience * Being either too strict or too lax with given instructions. Often, it will do only the one thing you asked for, disregarding any potential consequences of doing so. Other times, it might not listen to parts of the instructions (e.g., while investigating the PatchFile command, I couldn't get it to reliably output changes in a diff format) * Lack of knowledge beyond 2021: GPTinker is unaware of how to use certain libraries that have undergone significant changes since then or how to utilize new features of frameworks. This limitation could potentially be addressed by increasing the token limit and adding a command that allows the AI to access documentation on the internet. I'm curious to hear what you guys think about this approach and how it could be improved. PRs are welcome! It might be a dead end, but hey, you won't know unless you try, right? [1] https://github.com/maciej-trebacz/gptinker

Share card

Actual performance

3points
2comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, model, new · Missing: agents, macos, agent
98%98% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: para · Missing: supports, reddit linkedin, podcasting
90%90% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, existing, ide · Missing: https docs, excited, just released
61%61% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: soon · Missing: plus, platform, intuitive
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: video, para · Missing: mobile apps, ios, personal
38%38% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

De
Deslop Your TS Codebases44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Deslop Your TS Codebases

Hacker News1
Gr
Greetings, I'm Odin, your AI sidekick - *what troubles you*47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Greetings, I'm Odin, your AI sidekick - *what troubles you*

Hacker News4
9xChat
9xChat38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Your AI sidekick

Indie Hackers1ai
Chapity
Chapity19%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Your AI Sidekick

Indie Hackers1community
Re
Reverse-engineer OpenSpec specifications from existing codebases69%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Reverse-engineer OpenSpec specifications from existing codebases

Hacker News3
stagewise
stagewise85%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The frontend coding agent for existing codebases

Product Hunt+316Developer Tools
Sidekick
Sidekick67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Next Waze - Powered by Cameras

Indie Hackers2$1,000/mob2b
Zysu
Zysu36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Your Social Sidekick

Product Hunt+18
Zw
Zwsgi, zmq for existing wsgi apps.82%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Zwsgi, zmq for existing wsgi apps.

Hacker News1
Re
Registered developer49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Registered developer

Hacker News1