Gemini 2 with unrestricted code execution
Gemini 2 with unrestricted code execution
I recently experimented using Gemini 2 Flash and Gemini 2 Flash Thinking with unrestricted code execution in a local sandbox and open sourced the results: https://github.com/gradion-ai/freeact . Here's an example that uses Gemini 2 Flash Thinking as agent that acts via code, executed in a sandbox based on IPython and Docker: https://gist.github.com/krasserm/dcdae47f85ee9922e3284953d07... Gemini's code execution environment is restricted to selected Python libraries like NumPy or SymPy and prevents the model from installing new packages, besides other limitations. While this may be useful for agentic applications in restricted environments, it may prevent agents from adapting to new environments, especially agents that write their actions in code (see CodeAct paper https://arxiv.org/abs/2402.01030 ). Has anyone else experimented with Gemini 2 as a CodeAct agent? I'd be particularly interested in hearing about approaches to unrestricted code execution.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Camisole, a secure online judge for sandboxed code execution
Gnutella – Code Execution Visualization
Sequential – Visualizing JavaScript Code Execution
a decorator function for sequential execution of concurrent code
Agentic visual reasoning with code execution
VajraClaw – Deterministic <1µs execution guardrail for AI agents
iWF – A new “workflow as code” execution engine
Remote execution in mgmt
Where engineering execution becomes measurable
Hermetic AI Execution Workflow Orchestrator