BlogTechnical

Code-execution agents — the security review checklist for agents that write and run code

Agents that write and run their own code have an unbounded tool surface. The security review checklist for that capability.

Drel12 min read

Every review checklist for tool-calling agents starts from the same premise: the agent's capabilities are a list. Enumerate the tools in the manifest, scope each one to the minimum permission the task requires, verify enforcement sits at the tool layer rather than the model's reasoning, and the review has covered the agent's blast radius. That premise holds for an agent that calls a fixed set of named functions — a search tool, a database write, an email send — because the set of things the agent can do is fully known before the agent ever runs.

It stops holding the moment the agent can write and execute its own code. A coding assistant with a shell or interpreter tool, a data-analysis agent that generates and runs a Python script against a dataset, an agent with a code-interpreter tool bolted onto an otherwise ordinary chat interface — in each of these, the agent does not select from a fixed list of pre-defined actions. It synthesizes a new action, in the form of a script, at the moment it decides to act. The tool manifest lists one entry: “execute code.” What that code does is not fixed at design time and cannot be fully enumerated by a reviewer reading the manifest before deployment.

This piece is about the review checklist specific to that capability. It assumes the reader already has the fixed-tool permission model in hand — the manifest audit, least privilege by task and deployment context, authorization enforced at the tool layer — and asks what changes when one of the tools in that manifest is “write and run arbitrary code.”

A fixed tool manifest is a list of doors, and a review can check the lock on each one. A code-execution capability is not a door. It is a room the agent can build a new door into, on demand, in a shape the reviewer has not seen yet.

Why this is not the tool-manifest review again

It is worth being precise about what carries over and what does not. The principle of least privilege still applies, and a code-execution capability should still be justified against the specific task the deployment performs — a document Q&A agent has no business having code execution at all, and a data-analysis agent that needs it should have it scoped to that task, not left available across every deployment context the same agent design is reused in. That part of the review is unchanged, and the agentic AI security review still treats manifest scoping and privilege-escalation paths through tool chaining as prerequisite ground.

What does not carry over is the assumption that scoping the tool is the same as scoping what the tool can do. Scoping a “send_email” tool means restricting the recipient domain and requiring approval — the tool's behavior is fixed, and the review constrains its parameters. Scoping a “run_code” tool cannot work the same way, because the tool's behavior is not fixed. The parameter is the program. A reviewer cannot list every program the agent might generate and approve each one in advance, and trying to do so is the wrong shape of control for this capability.

Why static permission review breaks down

Static tool-permission review works by asking, for each tool, “what is this action allowed to do, and is that scope justified?” The question presupposes a finite, nameable set of actions. Code execution violates the presupposition directly: the set of programs a model can generate is not finite in any way a review can walk through, and a script the model writes in response to one user request may bear no resemblance to a script it writes in response to the next.

The consequence is that the control has to move. Instead of asking “what can this tool call do,” the review has to ask “what can the environment this code runs in reach.” That is a question about the sandbox, not about the code — network egress, filesystem scope, installed packages, process and resource limits, and what state (if any) survives after the code finishes running. None of those are properties of any single generated script. They are properties of the execution environment, and they are exactly the properties a review can actually enumerate and verify, because unlike the space of possible programs, the space of environment configurations is small and fixed.

1. Generated code that exfiltrates over an unanticipated path

The most direct consequence of an unenumerable tool surface is that generated code can make network calls a reviewer never anticipated, because the code did not exist to anticipate. A data-analysis agent asked to “clean this dataset and summarize the results” has no obvious reason to make an outbound HTTP request — until the generated script imports a library that happens to phone home for telemetry, or the model, following an injected instruction embedded in the dataset itself, writes a script that posts a subset of the data to an external endpoint disguised as a logging or analytics call.

This is distinct from the indirect-injection escalation path covered for fixed-tool agents, where the risk is a hijacked reasoning loop invoking a tool it already has. Here, the risk is that the generated code itself contains the exfiltration logic — there is no fixed “send data” tool to audit, because the agent wrote its own network call as part of a script whose declared purpose was something else entirely.

2. Generated code that reads outside the intended scope

A code-execution agent working on a specific dataset or repository often runs with a working directory that is narrower than the filesystem the underlying process can technically reach. Generated code that reads a relative or absolute path outside the intended working directory — a sibling directory containing another customer's data in a shared filesystem, a credentials file left readable by the execution user, an environment-variable dump that happens to include API keys injected for an unrelated service — is a routine failure mode, not an exotic one, because nothing about a generated script's syntax distinguishes an in-scope read from an out-of-scope one. The interpreter does not know the intended scope; only the sandbox's filesystem permissions do.

This is the same reason a fixed file-read tool is scoped to specific directories rather than accepting arbitrary paths — but for a code-execution capability, that scoping cannot live in a tool's parameter schema, because there is no parameter schema standing between the model and the filesystem once code is actually running. It has to live in the operating-system-level permissions of the account or container the code executes under.

3. Package installation and supply-chain risk

A code-execution agent that can install packages — pip install,npm install, or an equivalent — inherits the full supply-chain risk profile of the package ecosystem it draws from, and it inherits it at runtime, per session, without the vetting a human engineer would normally apply before adding a dependency to a project. A model asked to solve a data-processing task may reach for a plausible-sounding but unvetted or typosquatted package name, or a genuinely useful package that happens to run arbitrary code at install time via its own build scripts.

This threat is easy to underweight because it looks like a developer-tooling concern rather than a security one. It is a security concern specifically because the installation decision is being made by the model, at runtime, without the review gate a human dependency addition would normally pass through — no pull request, no lockfile diff a security team ever sees.

4. Resource exhaustion from generated code

Not every failure mode here is about data leaving the environment. Generated code that contains an infinite loop, an unbounded recursive call, a fork bomb, or a memory allocation that grows without limit degrades or takes down the execution environment itself — and, depending on how isolated that environment is from the rest of the system, can degrade shared infrastructure serving other sessions or other tenants. This is not necessarily adversarial: a model debugging its own script by iterating in a loop that never converges produces the same resource-exhaustion outcome as a deliberately malicious script would.

Because the code that causes this is, again, not enumerable in advance, the only reliable control is a resource ceiling enforced by the sandbox runtime — CPU time, memory, process count, wall-clock timeout — that applies regardless of what the generated code attempts, rather than any attempt to detect resource-exhaustion patterns in the code itself before it runs.

5. Code that manipulates its own execution environment

The last threat is the one most often missed in a first-pass review: code that writes something intended to survive past the current session, or that alters the execution environment's own configuration to gain more access on the next invocation than it had on this one. A script that writes a scheduled task, modifies a shell profile that a future session in the same environment will source, or writes a file to a location a subsequent session happens to reuse (a shared cache directory, a persisted volume mounted across sessions) can turn a single-session compromise into a persistent one.

The review question here is not about any individual script's content — it is about whether the execution environment itself is ephemeral. An environment that is torn down and rebuilt clean for every session closes this threat structurally, regardless of what any given script attempts. An environment reused across sessions, for performance reasons, has to justify that reuse against exactly this risk.

Review checklist

For any agent with code-execution capability — a coding assistant, a data-analysis agent, an agent with a code-interpreter tool — a review is complete on this boundary when it can answer each of these with a configuration artefact, not a policy statement:

Code-execution sandbox — threat, review question, control

ThreatReview questionControl
Network exfiltrationWhat is the sandbox's own egress policy?Default-deny egress; explicit allowlist of destinations, separate from the agent's own network access
Out-of-scope file readsWhat filesystem can the generated code see?Per-session filesystem scope with no path outside the working directory reachable
Supply-chain riskCan generated code install packages, and from where?Package installation disabled by default, or restricted to a vetted internal mirror with version pinning
Resource exhaustionWhat happens when generated code loops, forks, or allocates without bound?CPU, memory, process-count, and wall-clock limits enforced by the sandbox runtime, not by the generated code
Environment persistenceCan code write anything that survives past this session?Ephemeral, per-session execution environment torn down after use; no writable path that outlives the session
  • Is code execution isolated per-session, or is the environment shared across sessions or users?
  • What is the network egress policy for the execution sandbox specifically — is it distinct from, and narrower than, the agent's own network access for its fixed tools?
  • Is package installation allowed at all, and if so, from what source — a public registry unrestricted, or a vetted internal mirror with version pinning?
  • Is there a static or dynamic scan of generated code before execution, and what does that scan actually catch versus what it cannot catch given the space of possible programs?
  • What is logged about generated code — the script itself, not just the fact that execution occurred — for audit purposes after the fact?
  • Does anything written during execution persist past the session boundary, and if so, is that persistence explicitly scoped and justified?

Where a deployment cannot answer one of these with a specific mechanism, the honest finding is that the control has moved to the wrong layer — the reviewer is being asked to trust the code rather than the environment, which is the one thing a code-execution capability makes structurally untrustworthy to review directly.

Review the sandbox, not the script.

Drel's agentic AI review treats code-execution capability as an environment-scoping question — network egress, filesystem reach, package sourcing, and persistence — evidenced separately from the fixed-tool manifest review.

A note on scope: Drel reviews assessed systems against documented architecture, configuration and intent. It does not ingest live telemetry from production environments. Dispositions reflect the assessed system at the time of review and the re-assessment triggers that govern when the disposition must be revisited.