October 5, 2026

AI System Security Belongs Where Agents Act, Not in the Model

No items found.

Key Takeaways

  • For agents, AI system security depends less on how well a model resists manipulation and more on what the agent can do inside your apps.
  • Model guardrails are a necessary layer, but prompt injection can talk agents past them. Deterministic controls on actions set the boundary that holds.
  • Each agent needs its own short-lived identity scoped to the person it acts for, rather than a borrowed human session or a standing API key.
  • Start securing agentic AI by mapping which agents already act in your environment and what they reach, since shadow AI is unlikely to disappear completely.

Most AI system security programs protect the model while agents act somewhere else

Most organizations have a well-worn AI review by now, covering model selection, red-team results, data handling, and a privacy sign-off. Then an agent starts reconciling invoices in the finance app through an employee's logged-in session, and the review says little about its actions there.

Most AI reviews end at the model card. The agent's day starts after that.

AI system security covers three things: the models, the data they touch, and, increasingly, the actions they take. The first two have mature playbooks, while the third is where agents live and where most programs are thinnest.

In our view, many widely used AI security approaches are layered around the model, its safety system, and the application a team builds. That made sense when enterprise AI meant a chatbot the team deployed and controlled.

For agents, the UK's National Cyber Security Centre (NCSC) addressed this directly in its August 2026 agentic AI guidance. It advises organizations to implement additional safeguards rather than relying solely on model-level or harness-level protections. A harness is the software that wraps a model and turns its output into actions. Gartner's September 2026 guidance for CISOs says understanding the interaction between models and harnesses is more important than the safety of the model itself.

Those approaches were built for AI that answered questions, and agents now click buttons.

Agentic capability also arrives through software enterprises already run and through tools employees adopt on their own. Copilots appear inside SaaS suites, AI browsers navigate on a user's behalf, extensions read pages, and desktop agents connect to Model Context Protocol (MCP) servers. Gartner predicts 40% of enterprise applications will be integrated with task-specific AI agents by the end of 2026, up from less than 5% in 2025.

As we see it, some of these agents run on models the enterprise didn't select or test. A review centered on the model sees little of what they do.

Readiness is moving more slowly than adoption, according to the Q2 2026 survey of 297 cybersecurity leaders behind Gartner's September 2026 guidance. Of those surveyed, 54% had no defined approach to limit AI agent access or relied on predefined human access. The second half of that finding deserves attention: agents inheriting a person's access is a common default, and that default was designed for people.

Prompt injection makes model guardrails one layer, not the boundary

Teams putting agents to work summarizing supplier pages or triaging shared inboxes are handing them content written by strangers. Somewhere in that content sits an instruction the agent wasn't meant to follow. The stack doesn't flag it, because to the model the instruction looks like more data.

Large language models don't separate instructions from data, and the NCSC has been candid about what follows. In December 2025, it wrote it's "very possible" prompt injection "may never be totally mitigated in the way that SQL injection attacks can be." Asking the model to police its own instructions is a bit like asking the phishing email to flag itself.

Hijacking also works under test conditions, as NIST's Center for AI Standards and Innovation (CAISI) showed in a January 2025 evaluation. Its new red-team attacks raised the success rate against one agent from 11% to 81% on simulated tasks. That's a single model in a controlled environment rather than a real-world rate, though it shows how much room attackers found once they looked.

The consequence scales with access, because a successful injection inherits whatever the agent's tools and sessions can do. That's why the OWASP Top 10 for Agentic Applications placed agent goal hijack and tool misuse first and second on its December 2025 list.

Some security teams have reached for the brake, and Gartner advised organizations to block AI browsers for now, as The Register reported in December 2025. It's an understandable pause, but it isn't a strategy for a workforce that will use agents anyway.

The same NCSC post points toward a durable answer: design protections around "deterministic (non-LLM) safeguards that constrain the actions of the system." Model guardrails stay in place as a useful layer of defense. The boundary that holds is enforced on what the agent actually does.

The control point that holds is the session where the agent does the work

When the control point moves to the session, your team stops debating model trustworthiness and starts deciding what each agent may do. Security teams already know how to answer that kind of question. Gartner's September 2026 guidance frames it the same way, telling CISOs to govern autonomous multiagent systems "based on action privileges rather than model intelligence."

Where enforcement sits determines how much of the agent's work it can see, and three architectural approaches are common today:

  1. Native enforcement in an enterprise browser or workspace sees the page, the data, and the agent's actions as they happen.
  2. Extension-based controls add a valuable layer on consumer and AI browsers, carrying policy to places a managed browser doesn't reach.
  3. Network and proxy controls, built for an era when risk traveled mostly as traffic, see destinations but less of what happens inside the session.

Each approach has a role, and many enterprises run more than one. The closer enforcement sits to the agent's actions, the more precisely policy can decide:

  • Which apps and tenants an agent can reach, including corporate versus personal accounts
  • What data it can read, copy, paste, or upload, with redaction before anything reaches a model
  • Which actions, like send, submit, purchase, or delete, require step-up or human approval
  • Whether incoming content carries prompt injection or jailbreak attempts, detected inline as it loads
  • How each prompt, tool call, and action lands in one audit trail shared by people and agents

An agent asked to update a supplier record can open the vendor portal and change the mailing address under its own scoped identity. It can't paste a customer list into a personal chatbot tenant. If a page it reads nudges it toward a payment, the action waits for the employee who launched it.

Island Enterprise AI is one example of this architecture. It extends the Island environment from people to agents, inventorying agents, MCP servers, and extensions, and issuing just-in-time credentials through the Island MCP Gateway.

AI Protect classifiers run inline to catch injection, jailbreak, and data leakage attempts. People and agents share one policy engine and audit trail, in the Island Enterprise Browser or through the Island Extension on other browsers.

Session-level control is what lets security say yes to agents. When your team can see and constrain an agent's actions in the session, approving it becomes a policy decision instead of a leap of faith. The business gets the agents it asked for, and security keeps a complete record of what they did.

Agents need their own identity, scoped to the person they act for

Teams that search the identity provider for agents usually find a handful of service principals and API keys. Much of what's acting doesn't show up there, because the identity provider only knows about agents someone thought to register. An agent riding a browser session or a developer's token looks, to the identity provider, like the human who signed in.

The NCSC's August 2026 agentic AI guidance is direct on this point. It says agents should be assigned "their own unique identity in a class which differentiates them from human or individual systems." It also calls for credentials "with the shortest possible lifetime" and for monitoring agent activity as a form of user activity.

In practice, that guidance points to three properties for an agent identity:

  • Distinct from the human, so actions are attributable to the agent
  • Bound to the person and task it's acting for, so delegation is explicit
  • Short-lived and scoped to the minimum access the task requires

Standards are still forming, and MCP's authorization specification builds on OAuth 2.1 for HTTP servers. Authorization is optional, though, and local stdio servers pull credentials from the environment instead.

Meanwhile, NIST's National Cybersecurity Center of Excellence (NCCoE) released a draft concept paper in February 2026. It proposes a project showing how identity standards can apply to software agents.

The riskiest agent credential is often an employee's already-authenticated browser session, or an API key sitting in a developer's config file. Identity programs that count non-human identities only in the identity provider miss both. So ask "whose session is this agent using?" before asking "how many agents do we have?"

Audit then has to answer three questions for each action: which agent acted, on whose behalf, and under which policy. If a log can't answer all three, incident reviews turn into archaeology.

Human approval scales only when policy decides which actions earn it

The approval prompt appears for the 40th time before lunch, and the person on the other end clicks it without reading. Most teams have watched this happen with alert queues, and agents are reproducing it at machine speed.

A 2026 position paper from Hugging Face and Data & Society researchers argues users who repeatedly approve agent prompts drift into "approval fatigue." OWASP's agentic list names a sharper risk, human-agent trust exploitation, where confident and polished explanations lead operators to approve harmful actions.

The NCSC's August 2026 guidance recommends human oversight "alongside technically enforced controls" for higher-risk scenarios. Read together, these sources point toward oversight proportionate to risk. One way to structure it is a three-tier model (our synthesis, not a published framework):

  1. Allow and log low-risk, reversible actions like reading public pages or drafting text.
  2. Constrain sensitive actions by redacting, masking, or narrowing data, and by blocking personal tenants.
  3. Require a human for irreversible or external actions like sending, paying, deleting, or sharing outside the organization.

The tiers belong in central policy rather than in each agent's configuration. That way, a new agent inherits them on its first day instead of waiting for someone to write its rules.

Fewer approval prompts make oversight stronger, because when only the third tier reaches a person, each request carries weight.

Route that request to the person the agent acts for, at the moment of action, with the data in view. A security operations queue reviewing the same request hours later is reading history.

Regulation leans the same direction: for high-risk AI systems, Article 14 of the EU AI Act expects human oversight commensurate with risk and autonomy. It also asks overseers to stay aware of "automation bias," the regulatory name for rubber-stamping.

Your first agent inventory should map access paths, not count agents

When the board asks how many agents the company runs, the honest answer today is usually "we're finding out." You can start there without embarrassment, because the count matters less than what each agent can reach.

Shadow AI is the realistic starting point. In Gartner's 2025 survey of 302 cybersecurity leaders, 69% said their organizations suspect or have evidence of employees using prohibited public GenAI.

In a September 2026 post on shadow AI, the NCSC adds the practice is "unlikely to disappear completely." That argues for governing shadow AI instead of chasing it to zero, starting with a practical sequence:

  1. Discover agents, AI browsers, AI extensions, and MCP servers on endpoints and in browsers.
  2. Map each one to the sessions, apps, and data it can reach.
  3. Rank by exposure, weighting the combination of sensitive data, untrusted content, and the ability to act externally.
  4. Govern one high-value workflow end to end before writing a blanket policy.

Gartner's April 2026 guidance backs the third step, treating use cases combining sensitive data, untrusted content, and external communication as a "no-go zone."

Rank by reach rather than by count: a single coding agent holding a production API key outranks 200 chat extensions that only read public pages.

For the first governed workflow, pick something routine, like an agent filling vendor forms or summarizing customer tickets. Everyday work shows where policy creates friction, which a dramatic edge case rarely reveals.

The inventory works because it's taken where agents act, which is where AI system security has to live. A list of approved models tells you what was purchased, while a map of sessions, apps, and data tells you what can happen next.

Governing agents where they act is a shift you can test in one workflow

If you want to pressure-test this model against one of your own agent workflows, we're happy to walk through what we've built. Schedule a walkthrough.

FAQs

What is AI system security?

AI system security protects the models, the data they use, and, for agents, the actions an AI system takes. That last part is where agents change the picture most.

How do you secure an AI agent?

Give it its own short-lived identity scoped to the person it acts for, and enforce policy on what it can do in the session. Model guardrails remain a supporting layer.

Are AI agents a security risk?

They can be, mainly because a hijacked agent inherits whatever access its tools and sessions have. Limiting that access shrinks the risk more reliably than prompt filtering alone.

What are the key security controls for AI agents?

The core set is agent inventory, distinct agent identity, action-level app and data policy, risk-tiered human approval, and one audit trail for people and agents.

Is it safe to let employees use AI agents at work?

Yes, when agents operate inside governed sessions with clear action policies. The NCSC expects shadow AI to persist, so governing agent use tends to hold up better than banning it (see the inventory section above).

Island Team

Island is defining the future of work for people and AI agents. Its enterprise agentic control plane helps organizations enable, govern, and audit agentic workforces alongside people. Island boosts productivity across devices, browsers, applications, networks, and data while protecting sensitive information, simplifying access, and helping enterprises scale AI safely.