The Agent Is Not a Chatbot
The Tool-Call Loop as an Execution Primitive
The mechanism underneath every example in this book is small enough to state precisely. A coding agent's runtime holds a loop: assemble the current context (the conversation so far, whatever files or search results have already been pulled in), send it to a model, and receive back one of two things — more text meant for the human, or a tool call: a structured request naming a capability and the arguments for it. A shell command to run. A file to read or write. A URL to fetch. The runtime does not treat these two outputs the same way. Text goes to a chat window, where a human reads it before anything happens. A tool call goes to a dispatcher, which invokes the real capability — forks a shell, opens a file descriptor, opens a socket — and appends whatever comes back to the context before the loop repeats. Nothing about the model's output is inherently more trustworthy in the second case than the first. What changed is what the runtime does with it, and how many steps stand between the model deciding on an action and the action actually happening on the host.
That number, for an unattended agent working through a multi-step task, is frequently zero. This is worth being precise about, because ``tool call'' as a term undersells what is actually happening: a probability distribution over tokens has produced a structured instruction that the operating system will act on, with no compiler, no code reviewer, and no shell prompt waiting for a human to confirm in between. The interception point — the one place in this whole pipeline where a decision can still be made before the side effect occurs — is the moment the tool call is dispatched, and it looks nothing like a content filter. It looks like an authorization check on a request.
SigmaShake's own governance daemon treats this moment as a request/response protocol rather than a conversational turn: every host's native tool-call shape -- including Claude Code's own {tool_name, tool_input} format -- is normalized into one canonical {tool, input, session_id, ...} request before the decision is made, because the thing being decided is the same regardless of which product produced the call. This book calls that daemon SSG (SigmaShake Governance) from here on; Claude Code is one of several agent hosts SSG normalizes this way, and the wire shape below is the same regardless of which one produced the call. The request carries a tool name and a JSON payload; nothing about it resembles a paragraph of natural language waiting to be classified for tone. It resembles what it is: a function call with arguments, aimed at the operating system.
That shape is not an abstraction to take on faith; it is worth seeing with real values, not only field names. A trivial, harmless probe of this same interception point — the one this chapter's closing lab has the reader run directly — crosses the boundary looking close to {"tool": "Bash", "input": {"command": "echo hello"}, "session_id": "...", "client": "claude-code"}, and the daemon's response comes back shaped close to {"decision": "allow"} for an unremarkable command, or {"decision": "block", "rule_id": "deny-recursive-force-delete", "reason": "..."} for the cold-open agent's fixture-triggered delete-and-fetch command — only the resolved argument value differs; the wire shape does not. The daemon's response is a single EvalResult struct serialized to JSON, carrying a decision field and, on a rule match, rule_id and reason fields. Neither request nor response carries the fixture file's prose, the phrase ``sandbox database is corrupt,'' or anything a text classifier was built to recognize — by the time a decision exists, the loop has already reduced the entire event to a tool name, a resolved argument string, and a verdict. Figure 1.1 draws the whole loop this wire shape is one moment inside of, text branch and tool-call branch together, so the asymmetry between them is visible in one picture rather than assembled from separate paragraphs.
The command that intercepts this moment, ssg hook eval, is documented in its own Short help text as a check that ``reads a tool call as JSON on stdin, exits 0 or 2'' -- a binary, non-negotiable dispatch decision made before the shell, file write, or network call the tool call describes is ever allowed to run. Exit 0 and the loop continues; exit 2 and it does not. There is no third path where the underlying command runs partway, and no path where the decision happens after the fact. This is the property that makes the interception point valuable and also the property that makes it narrow: whatever decides at this boundary sees exactly one thing — the resolved tool name and its arguments — and nothing about the multi-turn reasoning, the file the agent read three steps ago, or the plan the model believes it is executing.
The decision space at that boundary is not binary in practice, even though the process exit code is: SSG's engine defines six verdicts -- allow, block, log, shadow, ask, and force -- and the internal code identifier for a denial is literally block, distinct from the rule-authoring vocabulary (and this book's own prose) which calls the same outcome DENY. That naming split is noted once, here, because it recurs throughout the source tree this book cites and is worth not tripping over later: when a later chapter quotes DecisionBlock from Go source, it means the same thing as a rule file's DENY keyword. Six verdicts, not two, is the more important fact — chapter 6 builds an entire policy language around the distinction between silently logging a call, silently permitting it while recording the fact, and stopping to ask a human, because collapsing all of that into allow-or-block throws away exactly the nuance an operator needs to run an agent unattended without either blocking their own legitimate work or rubber-stamping everything.
What ``Autonomy'' Actually Delegates
``Autonomous agent'' is marketing language for a specific, enumerable delegation of authority, and the imprecision of the phrase is doing real work against the reader's own risk assessment. An agent is not granted an abstract quality called autonomy. It is granted a small, concrete set of capabilities, and every attack and every defense in this book is organized around that set.
The capability taxonomy this book uses throughout is drawn directly from the governance engine's own constant set: six values -- execute, read, write, search, agent, and network -- against which every tool a host exposes is classified, regardless of what that tool happens to be named in a given product. A ``Bash'' tool, a ``run_command'' tool, and a shell built into an IDE extension are the same thing to a governance layer built this way: an execute capability. This is the vocabulary chapter 6 formalizes into a rule language; here, the point is narrower. Read the six capabilities as a delegation checklist, not an abstraction:
Run an arbitrary command as the operator's own OS user, inheriting every permission, environment variable, and credential that user's shell would have. Open any file the operator's OS account can open — source code, but also .env files, SSH keys, and browser credential stores, because the filesystem does not know the difference between ``project file'' and ``secret'' without help. Create or modify any file the account can write, including configuration files that change how the agent itself, or other software, behaves the next time it runs. Enumerate and read matching content across a filesystem — easy to underrate as passive enumeration, but it is a reconnaissance primitive that feeds the next tool call with whatever it finds. Spawn or delegate to another agent or subagent, extending the same question — what is this instance actually allowed to do — to every agent it creates. Reach any host the machine's network stack can reach, in either direction: pull a resource, or send one out.
Two of those, read plus network in sequence, are sufficient to exfiltrate every credential on a developer's machine. execute alone is sufficient to do essentially anything the operator could do by typing at their own terminal. When a product description says an agent works ``autonomously,'' this is the concrete claim underneath the word: some subset of these six capabilities, granted without a human confirming each individual use.
The same capability resolves to a different concrete operating-system mechanism depending on the platform the agent is running on, and the differences matter for every containment chapter later in this book:
The mechanism differs; the delegation does not. On every platform, granting execute means granting a process that runs as the operator, with the operator's own reach, the moment a model decides to use it.
Notice what this rule does and does not look at. It never inspects the developer's original instruction (``clean up the failing test directory''), the agent's stated reasoning, or the tone of anything. It inspects the fully-resolved command string the execute capability is about to receive, across three operating systems' equivalent destructive invocations. That is the entire model this book argues for: govern the six capabilities and the arguments they are about to receive, not the English that led there.
Why Prompt Filtering Is the Wrong Control Point for an Action
The natural first instinct, coming from a decade of chat-safety tooling, is to point the same kind of classifier that screens chat messages at everything an agent reads and writes. It is worth being specific about why this instinct, while not worthless, cannot be where the last decision is made for an action.
A prompt- or output-level filter classifies text. It answers a question of the shape ``does this string look like it is asking for, or describing, something unsafe.'' That question was built for a specific surface: the human-to-model message, and the model-to-human reply. Two structural facts about agent tool calls put most of the actual danger outside that surface entirely.
The first is where the dangerous text arrives from. In the cold-open scenario, the instruction that mattered never appeared as a prompt at all — it arrived as the content of a tool result, a fixture file the agent read on its own initiative several turns into an unattended loop. A classifier watching the human-model boundary has no occasion to fire on it, because nothing crossed that boundary. Extending the classifier to also scan every tool result does not close the gap so much as relocate it: the classifier is now being asked to predict, from a block of retrieved text, whether some future tool call several steps downstream will be harmful — which is intent classification applied to content that has not yet caused an action, rather than evaluation of an action that is actually about to happen.
The second is what the question can never resolve, however good the classifier gets. ``Does this sound dangerous'' and ``will this specific, fully-resolved command, run against this specific filesystem, right now, do the damaging thing'' are different questions, and the gap between them is exactly where both failure modes live: benign work phrased in alarming words gets blocked, and a precisely damaging command phrased blandly gets through. SHAKEDOWN's own corpus documents this candidly about the very tool this book uses as its running example. The benchmark's own assessment of indirect prompt injection -- an attacker's instruction arriving through a fetched page, a file, or a search result rather than the user's own message -- states plainly: ``near-zero containment ... the harmful follow-on action may be a legitimate-looking tool call that SSG cannot distinguish from benign use,'' because the payload lives entirely inside the model's context and is invisible to an action-layer rule until, and unless, it produces a tool call that a rule happens to cover. That is not a criticism unique to one product. It is the structural shape of the problem: a classifier operating before the action layer is reasoning about probability of intent; a gate operating at the action layer is reasoning about a concrete, resolved side effect. Only the second kind of reasoning can be made deterministic enough to be a last line of defense.
There is also a tractability gap worth naming directly, because it is the reason the rest of this book spends so little time on classifiers and so much on gates. A text classifier has to generalize over an unbounded space: any phrasing, any language, any level of indirection an attacker can compose, evaluated against a moving target of what ``sounds unsafe'' means in context. The action-layer decision at a tool call is comparatively small and closed: six capabilities, a finite set of arguments each one accepts, and a resolved value for each argument at the moment the call is dispatched. Matching a resolved shell command against a pattern is a narrower, more mechanical problem than judging the safety of an open-ended paragraph, and a narrower problem is one a deterministic system can be held accountable for getting right or wrong — which is exactly the property a security control needs and a semantic judgment call does not reliably have.
None of this makes prompt- and output-level defenses worthless — catching an explicit ``ignore your previous instructions'' before it ever shapes a later tool call measurably shrinks the population of attempts that reach the action layer at all, and chapter 9 gives that work its full, fair treatment, including where it earns its place in a defense-in-depth stack. The claim in this chapter is narrower and non-negotiable: whatever a text classifier decides, an action still needs a decision made about the action itself, at the moment it is about to happen, by something that looks at the resolved command and not the paragraph that led to it.
Read the second rule against the cold-open scenario. A prompt filter reading ``reset the sandbox database'' finds nothing to object to — resetting a database sounds like ordinary maintenance. The rule above does not read that sentence at all. It reads the resolved command against the resolved filesystem path, and it does not care what English produced the request.
The Trust Inversion
Classical software security draws its trust boundary in a stable place: input arriving from a network or a user is untrusted and gets validated; code that actually runs has been written, reviewed, and deployed by someone accountable for it, and is therefore trusted by default. The confused-deputy problem — a program with legitimate authority tricked by untrusted input into misusing that authority on the attacker's behalf — has been a named failure mode in this field since the 1980s, but it described a narrow class of programs with a fixed, enumerable set of possible actions: a print spooler that could be tricked into overwriting a file could still only ever print or overwrite a file.
Readers who have hardened a web application against cross-site request forgery have already met a narrow version of this shape: a browser, trusted by a server because it is carrying a valid session cookie, is tricked by a page the user never meant to trust into submitting a request the user never meant to send. The fix in that world is well understood — bind the request to proof of intent (a token, a re-authentication step) the attacker cannot forge — precisely because the deputy there, a browser issuing HTTP requests, has a small and well-understood repertoire. An agent inverts the stable half of that model. The ``code that runs'' is no longer fixed and reviewed — it is synthesized at runtime by a model conditioned on whatever text most recently entered its context, and an unbounded fraction of that text was never authored, reviewed, or even seen by the operator: a package's README, a GitHub issue, a web page fetched mid-task, a fixture file left behind by an earlier process. Meanwhile the execution context that text ends up controlling is fully trusted by every classical measure — it is, in the overwhelmingly common default configuration, running as the developer's own operating-system user, with the developer's own file permissions, credentials, and network reachability. Untrusted text can now produce fully-trusted-context side effects, and unlike the print spooler, the deputy's possible actions are not a short fixed list — they are close to the full authority of the account it runs as: read anything that account can read, run anything it can run, reach anything it can reach.
The practical consequence is that authorization has stopped traveling with the thing security models have always anchored it to. It used to travel with human identity at the moment of the action: a command ran because a specific person typed it and pressed enter. Now the actual author of a given command is the model, and the model's decision was shaped by whichever text most recently entered its context window — which may or may not be the operator's own words. A governance layer that tries to answer ``did a trustworthy party ask for this'' is asking a question the runtime structurally cannot answer, because provenance of intent is not preserved across the boundary between reading a file and deciding to act on it.
SHAKEDOWN's own corpus is explicit about one asymmetry this creates in practice: SSG denies the read capability on .env and credential paths, but ``the same secret read via cat in a Bash command is an execute call, which the read rule doesn't cover'' -- two different capabilities reaching the identical file, governed by two different rules that a policy author has to remember to keep in sync. This is not a flaw specific to one engine's rule set; it is what happens when the same underlying resource is reachable through more than one capability and the governance layer has to decide, per capability, whether to look.
This rule is the correct response to the trust inversion, and it is worth naming exactly why. It does not attempt to determine whether the instruction to read .env came from the developer's own request, from an injected comment in a fixture file, or from the model's own unprompted initiative — a determination the runtime cannot reliably make. It governs the read action against that specific path, full stop, regardless of provenance. Re-anchoring authorization to the action and the resource, instead of the (unrecoverable) intent behind it, is the only version of this defense that survives the fact that a model's ``decision'' to act is not a trustworthy signal about who really wanted the action to happen. Chapter 4 is where this book turns that single idea into a full map of where, across an entire host, that kind of re-anchored decision can still be made.
What This Book Builds
The rest of Part I finishes the groundwork before this book proposes anything: chapter 2 turns the categories only gestured at here into a full taxonomy anchored to real adversary tradecraft, chapter 3 walks a complete kill chain end to end, and chapter 4 names the seven control points a defense can occupy, with an explicit, unhedged position on fail-open versus fail-closed design.
Part II stays at the interception point this chapter introduced and asks what actually stands there in practice: hook mechanisms and their bypasses (chapter 5), a full policy language for the six capabilities (chapter 6), the hard tradeoff between allowlists and denylists (chapter 7), the specific dangers of the Model Context Protocol (chapter 8), and the honest, complete accounting of where prompt-layer defense helps and where it structurally cannot (chapter 9) that this chapter has deliberately left unfinished. Part III moves below the gate, into the operating system itself, for the cases where the gate's decision was wrong, bypassed, or never consulted. Part IV addresses the network perimeter and the blind spots existing security tooling has for a process that looks, to an antivirus engine, exactly like the trusted developer tool it is. Part V measures all of it honestly, including the SHAKEDOWN benchmark this book's own examples are drawn from. Part VI is applied practice, ending with the reader building a working, if deliberately minimal, tool-call gate from nothing.
None of what follows eliminates the risk described in this chapter's opening scenario. Every defense in this book reduces blast radius, catches a specific class of attack, or narrows a specific window of exposure — and every one of them has a documented way to fail, which later chapters name rather than hide. The agent's tool-call loop is where an unattended model's decision becomes an operating-system action; everything else in this book is an argument about what to do about that fact, not a claim that the fact can be made to go away.
Key Takeaways
- A tool call is a dispatch to a real operating-system capability, not a chat message, and the runtime that turns model output into action is this book's actual subject.
- ``Autonomy'' decomposes into six concrete, governable capabilities — execute, read, write, search, agent, network — and every defense in this book earns its place by governing one or more of them, on a real mechanism that differs by operating system.
- A prompt- or output-level classifier evaluates text; it cannot evaluate the side effect of a fully-resolved command against a real filesystem, and it has no visibility into instructions that arrive through tool results rather than the human's own message.
- The trust inversion is structural, not a defect in any one product: authorization now travels with whatever text most recently entered a model's context, not reliably with the human nominally in control.
- The only version of authorization that survives this inversion re-anchors to the action and the resource — what is about to happen, to what — rather than to an unrecoverable judgment about who really asked for it.
- Containment reduces blast radius; it does not eliminate risk, and no chapter in this book, including this one, claims otherwise.