Preface
Why This Book Exists
A coding agent that can read a file, run a shell command, and fetch a URL is not a bigger autocomplete. It is a program that writes and executes other programs, and the instructions that steer it can arrive from anywhere a model reads text: a pull request description, a package's README, a web page fetched mid-task, the tool description of a server it recently connected to. That single fact is the argument of every chapter that follows, and it is worth stating here, once, before the chapters start layering mechanism onto it.
The capability arrived quickly. Models that could once only suggest a line of code can now plan a multi-step task, decide which tool to call next, read the result, and decide again, without a human confirming each step. Vendors shipped this loop as a product feature, and it is a genuinely useful one: an agent that can run its own tests, fetch its own documentation, and fix its own compile errors saves real engineering time. The capability itself is not the problem. The problem is that the containment around it did not arrive on the same schedule.
Containment here means something specific: a control that sits on the path between ``the model decided to do X'' and ``X happened,'' and that can say no. Before agents had tool access, that control point barely needed to exist, because a human was already standing there — every command a human typed was a command a human had, in some sense, already decided to run. An agent collapses that gap. The decision to run a command and the act of running it can now happen in the same process, in milliseconds, driven by text the agent did not write and the human did not read. Closing that gap — with a hook that inspects a call before it executes, a policy that decides what a workspace may touch, a sandbox that bounds what a process can do even if the policy missed it, a network egress rule that limits where data can leave — is what this book is about.
None of that containment is new in the abstract. Namespaces, seccomp filters, App Sandbox, Job Objects, and DNS-layer egress control predate large language models by years; this book did not invent a single operating-system primitive it describes. What is new is the actor on the other side of these controls: not a human operator making a discrete decision, and not a fixed piece of malware with a static signature, but a process that improvises, driven by untrusted text, at a pace no human review step can match. This book takes existing containment mechanisms and asks a narrower question about each one: does it hold against that actor, and if not, where exactly does it give.
Who This Book Is For
This book is written for three overlapping readers. The first is an engineer who runs a coding agent against a machine they actually care about — a laptop with SSH keys and cloud credentials on it, a CI runner with access to a production deploy key, a workstation with a company's source tree checked out. That reader needs the mechanism, not the marketing: what a PreToolUse hook can and cannot see, what a sandbox profile actually restricts, what an egress rule does when the agent tries DNS instead of HTTP.
The second reader is a security engineer who has been asked to sign off on agent adoption inside an organization, often on a short timeline and with limited context on what these tools do under the hood. This book gives that reader a taxonomy of agent failure modes (chapter 2), a reference architecture with named control points (chapter 4), and an honest account of what each layer catches and what it misses, so a sign-off rests on a threat model instead of a vendor's slide deck.
The third reader is a platform or infrastructure team rolling agents out to a fleet: dozens or hundreds of engineers, a shared policy that has to hold across all of them, and a rollout that fails politically the moment it blocks real work. Chapter 23 is written directly for that reader, but the containment chapters that precede it are the material a fleet-wide policy gets built from.
This book is not for a reader who has never operated a Unix shell or a Windows command line and wants to learn those first; it assumes a competent engineer who is new to this problem, not to computing. It is not an offensive-security manual for testing systems that are not the reader's own — the adversarial techniques inside exist to test a reader's own guardrail, and the safety notice later in this preface says so without qualification. And it is not a document a compliance team can attach to an audit as evidence of certification: nothing in it constitutes a certification, an endorsement, or an audit finding by any standards body, because no such body produced one.
What This Book Does Not Cover
Four topics sit close enough to this one that a reader might reasonably expect them here. Each is excluded deliberately, not by oversight.
Model alignment and safety research — reinforcement learning from human feedback, constitutional training methods, interpretability work aimed at a model's internals — is a different discipline with its own literature, conferences, and open problems. This book treats the model as a component that sometimes does what an attacker wants and asks what stops the consequence of that, independent of why the model did it. A perfectly aligned model still needs a sandbox, because the untrusted text steering it did not come from the model's training.
Offensive security is not covered as a discipline in its own right. Chapter 20 and several lab exercises use adversarial testing tools and techniques, but every one of them is scoped to testing a guardrail the reader built or deployed, against a target the reader owns. This is not a penetration-testing curriculum, and none of its techniques are presented as a way to gain access to a system that belongs to someone else.
Machine-learning model theft and extraction — stealing model weights, membership inference, reconstructing a proprietary model's behavior through repeated API queries — protects the model as an asset. This book's threat model is the agent as an actor with tool access on a machine; it has nothing to say about protecting the weights themselves, and a reader looking for that material should look elsewhere.
Compliance certification is out of scope in a specific sense: this book does not map its controls to the control identifiers of any framework — SOC~2, ISO~27001, FedRAMP, or otherwise — and using it does not produce audit evidence. A security engineer can use the reference architecture in chapter 4 as an input to a compliance mapping exercise, but that mapping is not done here, and no statement in this book should be read as one.
How to Read This Book with an Agent
This book assumes some readers will not read it themselves — they will point their own coding agent at it and ask that agent to look something up. That is a deliberate design goal, not an afterthought: a book about agent containment that only a human can consult is missing an obvious use case.
The manuscript builds to a structured JSON representation and a stdio MCP (Model Context Protocol) server that a reader's coding agent can query directly, instead of relying on whatever the agent remembers about this book from its training data.Running bun run mcp starts a stdio MCP server exposing the manuscript to a reader's coding agent. The server exposes five tools: a full-text search scoped by part or chapter, a section-retrieval call that returns a section's text verbatim by its label, a rule-fetch call that returns a companion .rules file ready to install, a lab-fetch call that returns a lab's objective and its expected verdicts, and a query over the technique-by-operating-system coverage matrix that runs through every chapter. This is shipped as part of the project's own stated purpose: to provide, in addition to the manuscript itself, an ``AI-native reader surface (structured JSON + stdio MCP server).''The project's own description names the structured JSON and stdio MCP server as one of its two build outputs, alongside the manuscript. Register the server the same way you would register any MCP server with your host — the project's README documents the exact invocation — and the agent can pull the current text of a rule or a lab's expected verdict instead of paraphrasing it from memory.
The second convention built for an agent reader is the agent-note block that closes out procedural passages in most chapters. These are written differently from the surrounding prose: imperative, numbered, deterministic, with an explicit platform assumption and an explicit stop condition on every step, and with a destructive action replaced by its probe or dry-run form wherever one exists. A human reader benefits from reading these too, for a reason that has nothing to do with automation: an agent note is the most concrete statement in a chapter of what the author expects an agent to actually do, and it is often the fastest way to check whether a guardrail's documented behavior matches what a reader's own hook or policy file does when it runs.
If you are the agent reading this — follow the numbered steps in an agent note literally, verify each observable outcome before moving to the next step, and stop at the first failure rather than continuing past it. Nothing in an agent note should be read as license to run a step the note itself marks as destructive.
Conventions Used in This Book
Every chapter after this one uses a small, fixed set of typographic conventions, each carrying a specific meaning rather than serving as decoration.
A citation attached to a specific claim about SigmaShake Governance (SSG) or SHAKEDOWN behavior means the author opened the referenced file in this project's monorepo and checked the claim against it — the footnote names the file, and where useful, the line. A footnote of this kind is not decorative: a build step checks every one of them against the real repository before this book ships, and a claim whose citation no longer resolves is treated as a defect, not a stale link.
An operating-system coverage table lists one row for Linux, one for macOS, and one for Windows, each naming the actual mechanism and the actual command or API that implements a technique on that platform. Where a technique genuinely has no equivalent on one operating system, the text says so explicitly, under a labeled ``No equivalent'' note, rather than describing a workaround as if it were the same control.
A boxed warning marks a step that is destructive, irreversible, or costly if run against the wrong target — read it before running the command underneath it. A boxed note marks a secondary detail: useful, but not load-bearing to the argument around it. A boxed agent note, described in the previous section, is written for an agent to follow literally rather than for a human to read for understanding. A boxed lab, exactly one per chapter, names a runnable exercise and the companion artifact that implements it.
Every chapter closes with a short list of key takeaways: complete, actionable sentences, not a compressed restatement of the chapter's headings.
This book's index follows the same discipline as everything else in it: an entry marks a page where a term is genuinely explained or substantively extended, not every page where the word happens to appear. A term this book uses constantly — SSG, tool call, allowlist — is not indexed at every occurrence; it is indexed where the text actually teaches you something new about it. Use the index to find where a concept is argued, not to count how often it is mentioned.
A Reading Path, by Reader
The three readers named above do not need to read this book the same way. An engineer who only runs agents against a single machine they control can read Parts I through III straight through and treat Parts IV and V as reference material to return to once a specific control is in place. A security engineer building a sign-off can read chapter 2 and chapter 4 first, then chapter 22 for the comparison matrix, before deciding how much of the containment detail in between they need to read closely versus delegate to the engineers implementing it. A platform team planning a fleet-wide rollout can start at chapter 23 and work backward into the specific containment chapter each blocked control turns out to need.
(a CI runner, a hosted coding-agent sandbox you do not administer) the kernel-level chapters (chapter 10, chapter 11, chapter 12, chapter 13) describe controls you may not be able to configure yourself. Read them anyway for the threat model, then focus your own work on chapter 5, chapter 6, chapter 7, chapter 15, and chapter 21 — the layers a policy author and a platform operator, rather than a kernel administrator, actually control. Chapter 22 is written to be read on its own, with the layer taxonomy it defines and the comparison matrix it builds; the containment chapters before it explain why the taxonomy is shaped the way it is, but are not required reading to use the chapter's buyer's guide. read chapter 18, chapter 19, and chapter 20 together — they are written as one argument about measurement split across three chapters, not three independent topics.
The Companion Repository
The companion repository holds every rule file, script, and lab this book references, and it ships alongside the manuscript rather than as a separate purchase. Three directories matter. companion/rules/ holds real .rules policy files, written in the rule language chapter 6 describes and checked to parse under the shipped policy engine before this book ships. companion/scripts/ holds per-operating-system hardening and audit scripts — separate linux/, macos/, and windows/ subdirectories, POSIX shell for Linux, zsh-compatible shell for macOS, and PowerShell for Windows — each idempotent and each accepting a dry-run flag before it will touch anything. companion/labs/ holds one directory per chapter's exercise: a README stating the objective, setup, and expected verdicts; the runnable scripts; and an expected.json file recording what a correct guardrail should do, so a reader can check a result against a known answer instead of guessing.
https://book.sigmashake.com is the canonical distribution point for this title: the purchase page, a free sample (this preface and chapter 1), and the current companion repository download all live there. The companion repository is versioned separately from the prose so a script fix does not require a new edition of the book; check the distribution page for the current version and for any correction published against this edition before treating a discrepancy between this text and your own installed ssg as a defect in either one.
A lab follows the same shape throughout the book. A setup script builds an isolated scratch environment — its own temporary directory, its own throwaway policy file — so the exercise cannot touch a reader's real project configuration. An attack (or, in early chapters, an ``observe'') script runs the exercise itself and records what happened. A teardown script reverses everything the setup script built.A lab ``reads and writes only inside a temporary directory it creates,'' does not touch the reader's real .sigmashake/rules/ or global configuration, and ``makes no network calls.'' Read a lab's own README before running it — the specifics of what it isolates and what it requires (a licensed ssg installation, a particular operating system, a particular shell) vary chapter to chapter, and the README is the source of truth for that lab, not this preface.
A Safety Notice
This book teaches adversary tradecraft as it applies to agents — prompt injection payloads, shell metacharacter evasion, hostile MCP tool descriptions, denylist bypasses — and it teaches these techniques deliberately, because a guardrail nobody has tried to break is a guardrail nobody has actually tested. Several chapters, most directly chapter 5 and chapter 20, walk the reader through defeating a control the reader themselves built or deployed. That is the point, not a contradiction of the book's purpose: a rule you wrote and never tried to evade is a rule whose false confidence you have not measured.
Run every technique, script, and lab in this book, and in its companion repository, only against systems you own or are explicitly authorized to test. The labs are built to make this easy to honor — they run inside a temporary directory the setup script creates, they do not modify a reader's real project or global SigmaShake configuration, and they make no network calls to any third party. But ``safe by construction'' describes the lab exercises this book ships, not every technique described in prose; a reader who adapts a technique from this book against a target outside their own authorization is doing something this book does not condone and its author does not endorse. The full disclaimer covering this point is on the copyright page, and it applies to everything that follows it.
If a chapter's subject matter makes you uneasy to run even inside a sandbox, that instinct is a reasonable one to listen to. Skip the lab, read the mechanism, and come back to the exercise once you have a scratch environment you are comfortable using.
Acknowledgements
This book exists because engineers running coding agents against their own laptops, CI systems, and shared repositories kept asking a version of the same question: what actually stops this thing if it goes wrong. Their specific incidents, and their specific ``how would you have caught this'' questions, shaped which mechanisms made it into these chapters more than any single design document did. Thanks are due to everyone who read an early rule file, tried to break a draft hook, or filed the report that turned into one of this book's labs. Any error that remains in these pages is the author's alone.
What Guardrails Achieve
One honest sentence belongs here instead of only at the end of a chapter twenty pages in: the controls in this book reduce the blast radius of an agent that goes wrong, and they buy time to detect that it has. That is what they do. They do not make an agent safe, and no chapter in this book claims otherwise.
A hook that blocks a destructive shell command stops that command, in that shell, written that way. It does not stop the same intent expressed through an encoded argument the rule was not written to decode, an MCP tool that performs the equivalent action inside its own server process where the hook cannot see it, or a credential the agent already had a legitimate reason to read and could move through a channel the egress rule was never told to watch. Every chapter in Parts II through IV names the bypass class for the control it describes, not as a caveat tacked onto the end, but because a control's failure mode is as much a part of understanding it as its success case.
The reason this book has twenty-four chapters and not one is that no single layer is sufficient on its own. Chapter 4 names the layers explicitly and argues that they compose: a hook that misses an encoded command may still be caught by the sandbox that bounds what the decoded command could do; a sandbox escape may still be caught by the network rule that blocks the connection the escaped process tries to make; a network rule that gets bypassed may still show up in the audit log a detection rule reads an hour later. None of these is a wall. Stacked, they are closer to a set of margins — each one narrowing how much an agent that goes wrong can do, and how long it takes before someone finds out. Read the rest of this book with that framing, and none of it will overpromise what it delivers.