A lightweight, hackable AI agent that can control your machine by executing Bash scripts
-
Agent2 controls a machine by writing a bash script at the end of its response. The script will be extracted from the response of the LLM, executed and the output/error will be piped back to the LLM. That's it, nothing more!
-
Agent2 tries to be minimalistic and focuses on the essentials. With only ~130 lines of Python code, Agent2 is short, simple, very light, easy to understand, hackable, easy to extend and still quite capable.
-
Agent2 relies on the host's CLI environment. To ensure Agent2 is productive, provide it with relevant tools and update
context.txtso the LLM knows how to utilize these tools. Under the "How to add tools" section, there is an example.
A human can do a lot with a script/terminal, therefore, an agent can do it as well. The more agents/LLMs advance the less framework is required. Because it's so simple, Agent2 is an agent framework for agents.
Agent2 is programmed in Python, open-source on Github and was written by Lukas Pfitscher (feedback is appreciated: agent2pf@gmail.com)"
- Be in the directory where agent2 should be copied, then download with:
curl -L -o agent2.zip https://github.com/lukaspfitscher/Agent2/archive/refs/heads/main.zip- Extract the directory, remove the .zip file and rename it:
unzip agent2.zip; rm agent2.zip; mv Agent2-main agent2- Only
python3and therequestslibrary are required to run Agent2. To install do:
cd agent2; chmod +x install.sh; ./install.shor just paste the following command into your terminal:
apt update; apt install -y python3 python3-pip python3-requests-
Next add your Openrouter API key in the
agent2.pyfile. -
Optionally, also adjust the token limit in the
agent2.pyfile. -
Launch Agent2 with (Be careful! It can control your system!):
python3 agent2.py- Agent2 can run on any Linux system (directly on your local machine, server, VPS or in an environment: docker, podman...).
- The file
context.txtcontains the context of the model, like "You are Agent2, a..." - Everything written to the terminal or in the
conversation.txt. file is exactly as it is seen by the LLM, except the yellow notes in the terminal for better user visualization. - The user can input after an
INPUT:or use theprompt.txtfile for the first prompt (useful for automation). - For user input, the Enter key is a normal new line, submit with Ctrl+D (standard Unix convention for 'end of input').
- The LLM responds with
LLM:. - The communication between LLM and host is kept simple:
The LLM triggers script execution by writing:
agent2_script_start - Everything after the last
agent2_script_startis interpreted as a bash script - Before every execution the script is reset (path doesn't persist, no environment variables persist).
- After that, the script gets executed in a separate shell and therefore doesn't block the agent.
- The script always starts in the
working_dirdirectory. - The script output is written to the
output.txtfile. - Agent2 waits 0.2 seconds for the command to finish and read from
output.txt. - If the command takes longer, Agent2 can put itself to sleep with its own PID.
- The output/error is piped back to the LLM after a
TOOL:message. - Conversations are saved as plain text (no json) in
conversation.txt. - The text in
conversation.txtis exactly as the model sees it. The conversation includes all the start and stop markers. - If the model doesn't request another script, the user is prompted for further instructions.
- The user can stop Agent2 by pressing Ctrl+C.
Here is an example of a minimal SYSTEM-USER-LLM-TOOL conversation:
SYSTEM: You are Agent2, an Agent that can execute bash scripts
by writing agent2_script_start at the end of your response...
USER: list current files!
LLM: agent2_script_start ls -a
TOOL: . .. boot etc lib run...
LLM: Entries in the current directory: boot etc lib run...
Here is the directory structure of Agent2:
agent2/
├─ install.sh # Installation script
├─ agent2.py # Contains config + whole Python code (single file)
├─ readme.md # Readme / documentation (the file you are currently reading)
├─ context.txt # Context of the model ( like role description )
├─ prompt.txt # The initial prompt
├─ conversation.txt # File where the whole conversation is saved
├─ output.txt # Output and error of the script.sh
├─ pid.txt # Process ID, the agent can be paused or killed by other agents
├─ working_dir/ # Working directory of the agent; starting path of the script
Agent2 relies on the host's CLI environment.
To ensure Agent2 is productive, provide it with the relevant tools
and update the agent's context.txt so it knows how to utilize these tools.
Here is an example of how to give Agent2 web search capabilities.
First install the software as usual:
apt install -y ddgr # Get web search results (DuckDuckGo search;`ddgr -x search_keyword`)
apt install -y curl # Get raw website content
apt install -y lynx # Extract useful text (`curl -s https://www.x.com | lynx -stdin -dump`)Add a note to context.txt so the model knows it can use these tools:
You can search the web with ddgr, curl, lynx
- New agents can be made by simply copying Agent2's directory.
- Guidance can be given in the model context.
- This is kept simple: one program, one agent, one conversation. Integrating multi-agent directly in the program makes everything much more complex.
# Copy current agent
cp "path_agent_dir" "path_new_agent_dir"
# Move in to the new agent dir
cd "path_new_agent_dir"
# Clear existing conversation
> conversation.txt
# Clear the agents working directory
rm -rf working_dir/*
# Add a prompt to the model by writing to the prompt file
echo "You are a subagent, make a cleanup of..." > prompt.txt
# Launch Agent2:
python3 agent2.pyAgent2 doesn't integrate a fixed agent structure. Deciding which agent to spawn is up to the agent itself. Agent2 can do this by itself just prompt it right.
You can just paste the whole readme.md into the context.txt file of the agent. So it knows about itself.
Here is the description how Agent2 behaves:
You are Agent2, the best tinkerer, engineer, scientific researcher and coding agent.
- You possess common sense, are logic-driven, and are helpful.
- You remain concise and precise with your answers.
- No feelings. No guessing. Rely on hard scientific truths and facts.
- You are maximally truth-seeking, even if the subject is controversial.
- You think in a clean, structured way. You plan / make todo lists / think step by step for the requested task if needed.
- Agent2 can occasionally get stuck in a repetitive loop.
There is a built-in counter-mechanism to prevent this.
Set
max_tokensinagent2.pyto limit token/spending. - Mid-Response Triggers: Due to model limitations,
the agent may occasionally include the
agent2_script_startstring while "thinking" or explaining a process. This will prematurely trigger command execution. - The agent may sometimes ignore or forget specific instructions explicitly stated in the initial context (a limitation of the underlying LLM's capabilities).
- If the LLM is not explicitly told 'script executed', it will think it didn't work and repeat itself over and over.
-
No built-in tools: With Bash the agent can use all installed CLI tools, if an additional tool is required, it needs to be installed and added to the LLM's context to make the LLM aware of the tool. This keeps Agent2 minimalistic and modular.
-
CLI interface only: No GUI overhead
-
No guardrails / security: This is done by user restriction and environments (docker, podman...). Linux offers lots of tools to restrict a user.
-
No MCP integration: a CLI tool exists for this "
mcp-cli" -
No Multi-modal: a simple script can handle this:
python-llm,aichat,curl -
No memory / No RAG: not an essential feature
Handling all the control sequences gets too complicated. Writing to a file and executing it is much simpler. With this, the agent can already do a lot. It doesn't need a "pseudo-terminal".
Just writing user: or system: won't work.
Every model needs a specific chat format.
The model is trained on these markers.
Without this format the model behaves terribly!
The correctness of stop tokens and role markers is crucial for stable behavior.
If you change a model you also need to change these markers in agent2.py.
You can look them up on the web for each open-source model.
Yes, just change install.sh to your distro’s installer.
Everything else stays the same.
The PID of the Python program running Agent2 (not the separate shell the agent can execute)
is saved in the pid.txt file.
With it Agent2 can put itself to sleep with:
PID="$(< ../pid.txt)" #read pid from file
kill -STOP "$PID" && sleep 10 && kill -CONT "$PID" This is useful for waiting for commands (like waiting for downloads, monitoring), reminders and counters.
The PID is saved in the pid.txt file so other programs/agents have control over the current agent.
This was a deliberate design choice because it is simple and effective. 0.2s handles the fast commands (ls, cat, echo, etc.) — which are most commands. Integrating command execution flags is complicated and often fails to cover every scenario due to numerous edge cases (some commands continue writing and therefore never finish, or a script can contain multiple programs). For situations requiring longer wait times, Agent2 can put itself into a sleep state.
Yes, but its more lightweight and insecure
The script can do this by itself by adding echo "Shell PID: $$".
For most commands this is not needed because they will finish and the shell closes automatically.
We believe this overcomplicates things. An LLM is fundamentally text in, text out — and we should treat it as such. Because we use completion mode, only open-source models are available. Closed-source models don't publish their chat templates, so they aren't compatible with this project. However, you can easily change to normal json chat conversation— just ask an LLM to modify the agent2.py file.
Turns out you can give the LLM long instructions and it will actually perform better. These LLM context texts are carefully tested to work properly.