From Guardrails to GuardFails
Part 3 of the D2 Blog Series — “Authorization in the Age of Agents”
Part 3 of the D2 Blog Series — “Authorization in the Age of Agents”
I got some feedback from my last postabout how some readers did not enjoy anything that was AI generated. Thank you all for the feedback. At the end of the day, I am receptive to all feedback no matter how critical, and I have heard you all. Moving forward, all posts will be purely from the mind of David, and some from my friend, sourced from our discussions on these matters.
Today, I’m going to discuss when AI can let you down; to be more exact we’ll be discussing AI Guardrails, but in order to discuss what those are, we’ll start our discussion on what guardrails even are.
Guardrails

We’ve all seen these things on the road, which we all call highway guardrails. What are they for? Well they are to “protect” cars from crossing over a boundary, in which if it is crossed, the car (and the people contained within) can get seriously hurt (or worse). I put “protect” in quotes because, well, if a car is hurtling at 100mph (160kmh) at certain angles, the guardrail isn’t going to protect much. To explain this a bit more abstractly, this guardrail will NOT protect in all circumstances.

Let’s take another example of a guardrail. I’m not sure if my audience is familiar, but I will make the assumption that there may be some who don’t know what the above is. This is called a circuit breaker. It’s intended design is to “break circuits” when the electrical load of a system exceeds what is “deemed as safe”. This deeming of safe has been heavily studied by experts and a rule was determined that at some level of electrical load, it would become dangerous to a structure by starting an electrical wiring fire and would pose a danger to humans contained within. Unless there is a fault within this system, this guardrail is deterministic because the “monitoring” of the current is performed physically and automatically by the internal mechanisms of the circuit breaker or fuse itself.
The whole point is that there are some guardrails that are deterministic and some non-deterministic. Some do well in all circumstances, and some do not.
Why AI Needs Guardrails in the First Place
When a system starts behaving unpredictably, our instinct is to add boundaries. In AI, those boundaries are called guardrails. They’re meant to stop a model from saying something harmful or taking an unsafe action.
In traditional software, guardrails are built into the code. We validate inputs, check permissions, and make sure only admins can delete data. These rules are deterministic; they pass or fail every time.
Many AI guardrails are softer. They’re made of text: prompts, filters, or layers that try to guide the model’s behavior.
Example system prompt:
You are a polite and professional assistant. Never use profanity or offensive language.
If a user includes profanity, respond with a friendly warning instead.
These guardrails sound helpful, but they rely on interpretation. A phrase like “offensive language” has no fixed meaning to a model. It’s just another piece of text to predict around.
Deterministic code behaves the same every time. Text-based guardrails don’t; the model might follow them, reinterpret them, or ignore them.
That difference between fixed logic and probabilistic behavior is where many modern AI safety mechanisms start to break, and where a guardrail becomes a guardfail.
Deterministic Guardrails: Rules That Always Work the Same Way
Not all guardrails are bad. Some work perfectly because they leave no room for interpretation. These are deterministic guardrails; they behave the same way every time.
Take, for example, a circuit breaker. It doesn’t decide when to act; once the current passes a limit, it shuts off automatically. Or consider a seatbelt; you buckle it, and if there’s a crash, it locks instantly.
In software, firewalls and access control lists do the same. They check against a rule and block anything that doesn’t match.
These guardrails work because the outcome is binary: allowed or denied. Every variable is known, so there’s no guessing or interpretation.
AI systems, however, don’t follow clear, fixed rules. They live in the space between meaning and context. Instead of following instructions exactly, they try to interpret what you mean.
I often liken LLMs to an intern where they might mis-interpret your intent.
That’s why it’s hard to build reliable, deterministic safety into something that doesn’t behave deterministically in the first place.
Non-Deterministic Guardrails: Guessing Games
Many AI guardrails don’t work like circuit breakers or firewalls. They don’t enforce rules in predictable ways, they guess. These are non-deterministic guardrails, built on top of systems that already behave unpredictably.
For example, a system prompt might say:
You are a helpful assistant. Never produce harmful or offensive content. If a user asks you to do something illegal, politely refuse.
It sounds safe, but what counts as “harmful” or “illegal” isn’t defined. The model interprets those words each time. One run might block a question, another might allow it, and another might overcorrect.
It’s like asking a new group of people every day what counts as “rude.” The answers change, and so does the enforcement.
Companies often add layers to catch mistakes. Content filters or other models that review outputs are such layers. That’s like having several referees with different opinions. You get more judgments, not more consistency.
This is the core of non-determinism in AI safety. The model doesn’t know what’s right; it predicts what’s likely to be right. It’s guessing, not reasoning.
Think of asking three weather apps if it will rain tomorrow. Two say yes, one says no. That’s fine for deciding whether to bring an umbrella, but not for deciding whether to close an airport.
That’s how AI guardrails behave today. They reduce risk, but they can’t remove it. Every layer still depends on the same uncertainty.
When Guardrails Start to Fail
Once people realize a model can make mistakes, the next idea sounds logical: use more models to catch them. One writes, another reviews, a third checks for safety. On paper, that feels like layered protection. In reality, it’s just more guessers in the same guessing game.
If each layer is uncertain, stacking them doesn’t make the system reliable, it only spreads the uncertainty.
It’s like flipping three coins instead of one. You might see more “heads,” but you haven’t removed the chance of a “tail.” Or like having three referees, each with a different idea of what counts as a foul. They’ll agree most of the time, but not always, and reconciling disagreements only adds confusion.
In AI, this happens when one model generates text, another model filters it, and a third audits the result. Each layer adds cost and complexity, but none guarantee consistency. Every model still runs on probability, not certainty.
At small scales, these inconsistencies are easy to ignore. But as systems grow and automate more actions, even one unpredictable moment can lead to real problems.
In the next section, we’ll look at a few real examples where these small failures in guardrails became visible in the real world.
Real Examples of GuardFails
If you’ve read up to here, I thank you, and this is the fun part.. real examples of guardfails!
We’ve already started to see what happens when these “soft” guardrails fail in real systems. A few recent incidents show how unpredictable these models can be when something slips through.
In July 2025, a malicious prompt slipped into Amazon’s Q coding assistant for VS Code. The injected text told the model to delete files and terminate cloud resources using AWS commands. Nothing catastrophic happened, but only because a human caught it early.
Around the same time, an AI coding agent from Replit deleted a production databaseduring a test cycle. There was no hacker, no exploit. It was just a model that jumped to conclusions and performed actions on its own.
And most recently, even Google ran into this problem in an unexpected place. In Google Maps, Gemini, their AI summarizer, was embedded to answer user questions about restaurants. As shown in this LinkedIn post by Rémi Ounadjela, someone asked “Reverse a list in Python,” and Gemini cheerfully answered with code right in the middle of the restaurant listing for Los Altos Grill.
That one is harmless, even funny. But it shows a deeper issue: the model had no sense of context. It didn’t know that a restaurant page isn’t the right place for a programming answer. It simply saw a question and tried to respond the best way it could.
When systems like these are connected to real tools or sensitive data, the results won’t be funny, they’ll be costly.
All three examples from Amazon, Replit, and Google, share the same root cause. We’re asking probabilistic systems to behave like deterministic ones, and they can’t.
Probability Is Not Protection
Security isn’t about what usually happens. It’s about what can happen.
A system that works 99.9% of the time sounds reliable, but AI systems run millions of operations a day. That tiny failure rate still means thousands of unpredictable outcomes, and it only takes one to cause real damage.
It’s like building a bridge that’s probably strong enough. That sounds fine until the one truck that’s a little heavier crosses it. “Most of the time” isn’t safe.
The same idea applies to AI safety. Guardrails based on probability can reduce risk, but they can’t eliminate it. Security requires guarantees from predictable systems where the same input always produces the same result.
AI guardrails live in the gray zone of likelihoods. They might be right most of the time, but most of the time isn’t good enough. Probabilistic safety protects by chance; deterministic enforcement protects by design.
Deterministic Enforcement: Moving from Advice to Action
Real safety can’t depend on reminders or filters. It has to be built into the system itself. Deterministic enforcement means unsafe actions are blocked automatically, every time.
Like a circuit breaker, it reacts without hesitation whenever a threshold is crossed. That’s how AI systems need to behave. A guardrail shouldn’t reason about safety; it should apply it.
The future of AI security is about (non-AI) systems that enforce their own rules. Define what a model or agent can do in one place, and let the system enforce it consistently across every call.
That’s the idea behind D2 which is turning authorization into a single, verifiable layer of truth that follows every action. When safety becomes part of how the system works, it stops being about probability. It becomes part of the design.
In Part 4, we’ll look at how that kind of system can define what an agent can do, when, and on what data, making security a property of design, not an afterthought.
Subscribe to Breadcrumbs
New field notes on appsec and AI agent security. Free — unsubscribe anytime.
