
Share:








Share:




Share:




You have heard about AI workstreams by now, the work you hand to AI to run for you. The sharper version is the AI agent: AI that does not just answer but goes and does the task, deciding what to do and taking steps across your systems on its own. Everyone is building them. The demos are genuinely impressive, and in the back of your mind is a question you have not said out loud. What is this thing, and do I need a whole department before I can say yes to it.
You do not need an IT department for that. You need one thing almost every plan leaves out. Companies bring us these plans every few weeks to review before they commit, and the plan always describes everything the agent will do and almost nothing about what stops it when it does something it was never meant to. The person approving it is rarely an engineer, and they can feel that gap without being able to name it.
That missing piece goes by a bureaucratic name, AI agent governance, and a plainer one we prefer: guardrails. It is the answer to the problem in our companion piece, the AI already running inside most companies that nobody approved. This piece is about what putting guardrails on it actually takes.
The pattern is consistent. A capable model gets wired into your systems, performs beautifully in the demo, then meets a situation nobody scoped. That is when the missing guardrail stops being theoretical.
It reached the news in April 2026 at PocketOS, which runs the booking systems car rental firms use to manage operations. An AI coding agent on a routine task in a testing environment hit a mismatched credential and decided to fix it. It found a key with enough access sitting in an unrelated file and used it to delete the storage holding the entire production database. Every backup went too, because they sat in the same place.
| 9 sec the time the agent took to delete the production database and every backup. |
Asked afterward what it had done, the agent wrote a confession listing the safety rules it had broken. It had been told in writing not to run destructive commands. It ran them anyway, because nothing outside its own reasoning was built to stop it.
The model was working normally. Nothing stood between its reasoning and the systems it could reach. Even Meta, with more security engineering than almost anyone, had an internal agent expose data it was never meant to touch this year for the same reason.
| A capable AI agent with no guardrails is a brilliant stranger with the keys to every room and no one watching which doors it opens. |
The deletion is dramatic and rare. The expensive version is quiet, daily, and shows up in the two places a leader answers for: money and quality. Ungoverned AI costs you four ways, none of which make the news:
The last one is the real cost, the one you cannot budget for. A confident wrong answer that reaches a client outweighs every license fee on the invoice. Guardrails move that discovery earlier, from the customer’s email to the moment before the work left the building.
The System That Puts Guardrails on Every AI Workstream Is an Agent Management System (AMS)
Asked how to keep agents under control, most of the market points at dashboards that watch an agent once it is live and report, after the fact, what it did. That has its place. For an agent that empties a database in nine seconds, a report on the tenth second is a record of the funeral.
Real control is set before the agent runs and enforced on every run. The system that does this has a name. An Agent Management System, or AMS, sits above your agents and manages every one across its life: it sets what each agent may do, checks its work before that work leaves the building, keeps the record of what it did and why, and retires it when it is no longer fit to run.
| The short version: an AMS designs, approves, governs, and monitors every agent from before it runs to after it retires. Guardrails are what it puts around each one. A raw agent is an employee with no manager. An AMS is the manager. |
An agent dropped into unclear processes hesitates the way a new hire would, and one with broad access and no rules moves fast in the wrong direction. An agent is only as safe as the system you run it inside.
That system is the controls layer of our AI operating system, Mustang. We wrote before about one half of what Mustang does, giving your AI your company’s memory so it stops answering like an outsider. The AMS is the other half, so the agent stops acting like a stranger with the keys.
The difference between a raw agent and one running under an AMS shows up in the mechanics of a single run:
| In the run | A raw AI workstream | The same workstream under an AMS |
|---|---|---|
| It produces an answer | Hands it straight to you or your systems. | The answer runs a fixed chain of checks that fails closed, so a result that misses a check is stopped before it reaches anyone. |
| It cannot resolve something | Guesses, and sounds confident doing it. | Stops, marks the gap, and hands back a ranked question for a person to answer. |
| Knowing the job is done | Stops when the prompt runs out. | Runs against a fixed list of everything the job requires, and closes only when every item is produced. |
| What you get back | Text, with no way to see how it got there. | A record tying every output to its evidence and to the person who owns it. |
| Reusing what works | Rebuilt from scratch by whoever builds it next. | Proven steps are registered once and versioned, so any agent calls the same trusted piece. |
The two things that cost you most are handled at the source. Quality stops being luck because the chain fails closed, and spend stops climbing because a proven step gets built once instead of rebuilt by every team.
You do not trust a capable employee by watching every keystroke. You trust them because you can see what they did and why. An AMS gives an agent that accountability, which is what turns a stalled pilot into something that reaches production.
| 50%+ the minimum drop in manpower a job requires across our Mustang engagements, effort per workstream falling 40 to 65 percent by complexity. |
You do not have to build this yourself. Building agents safe to run in production, and the AMS that governs them, is a specific discipline, and the one we practice every day. We build the agents your operation needs, on the models you already use, ChatGPT, Claude, Gemini, and run them under the AMS inside Mustang, on your existing systems, inside your walls, your data never leaving.
What to Check Before You Approve the Next Agent
When the next plan reaches your desk, three things have to be true. You can check them yourself, without being technical:
Get those three right and the agent is something you can run. Miss one and you have a demo going into production on trust. An AMS makes all three true for every agent at once, instead of one plan at a time.
Pick one process and we will run it under Mustang against your real systems for two weeks, and show you the guardrails working before you commit a dollar.
Sources
The Register, April 27, 2026: Cursor-Opus agent snuffs out startup’s production database.
TechCrunch, March 2026: Meta is having trouble with rogue AI agents.
Share:









We’ve helped teams ship smarter in AI, DevOps, product, and more. Let’s talk.
Actionable insights across AI, DevOps, Product, Security & more