Ge

Gemini 2 with unrestricted code execution

Hacker News

Gemini 2 with unrestricted code execution

I recently experimented using Gemini 2 Flash and Gemini 2 Flash Thinking with unrestricted code execution in a local sandbox and open sourced the results: https://github.com/gradion-ai/freeact . Here's an example that uses Gemini 2 Flash Thinking as agent that acts via code, executed in a sandbox based on IPython and Docker: https://gist.github.com/krasserm/dcdae47f85ee9922e3284953d07... Gemini's code execution environment is restricted to selected Python libraries like NumPy or SymPy and prevents the model from installing new packages, besides other limitations. While this may be useful for agentic applications in restricted environments, it may prevent agents from adapting to new environments, especially agents that write their actions in code (see CodeAct paper https://arxiv.org/abs/2402.01030 ). Has anyone else experimented with Gemini 2 as a CodeAct agent? I'd be particularly interested in hearing about approaches to unrestricted code execution.

Share card

Actual performance

22points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
95%95% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: gemini · Missing: supports, reddit linkedin, podcasting
57%57% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: open source, ide, io · Missing: https docs, excited, just released
45%45% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
33%33% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
29%29% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
4%4% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ca
Camisole, a secure online judge for sandboxed code execution26%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Camisole, a secure online judge for sandboxed code execution

Hacker News10
Gn
Gnutella – Code Execution Visualization41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Gnutella – Code Execution Visualization

Hacker News3
Se
Sequential – Visualizing JavaScript Code Execution44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Sequential – Visualizing JavaScript Code Execution

Hacker News2
a
a decorator function for sequential execution of concurrent code44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

a decorator function for sequential execution of concurrent code

Hacker News1
Agentic Vision in Gemini
Agentic Vision in Gemini96%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Agentic visual reasoning with code execution

Product Hunt+184Artificial Intelligence
Va
VajraClaw – Deterministic <1µs execution guardrail for AI agents14%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

VajraClaw – Deterministic <1µs execution guardrail for AI agents

Hacker News2
iW
iWF – A new “workflow as code” execution engine47%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

iWF – A new “workflow as code” execution engine

Hacker News68
Re
Remote execution in mgmt29%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Remote execution in mgmt

Hacker News1
NuePrism
NuePrism17%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Where engineering execution becomes measurable

Indie Hackers1ai
Invoker
Invoker49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hermetic AI Execution Workflow Orchestrator

Indie Hackerscommitment-full-time