← writing

Seven Questions Boards Must Ask Before AI Agents Go Live

2026.09.30

A recent Deloitte survey of over 3,000 business and IT leaders across 24 countries found that only 21% of organisations have a mature governance model in place for agentic AI. A separate global survey of CIOs found that 81% say they have lost oversight of their own AI agents, with the majority citing agents built outside approved systems entirely. Boards are approving budgets for AI agents faster than anyone is building the questions to ask about them.

That gap matters more for agents than it did for earlier waves of software, because an agent does not wait for a person to click a button. It reads, decides and acts, often across systems that hold real customer data and real money. The seven questions below are not a compliance checklist. They are the questions that would have caught the AI security incidents already making headlines this year, before they happened rather than after.

1. What can this agent access, and does it need all of it?

Start with permissions, not capability. An agent built to summarise sales leads does not need standing access to account financials, payment systems or email. In September 2026, researchers at Zenity Labs disclosed a Salesforce Agentforce vulnerability, later named SalesBleed, in which a single subagent's combined access to both lead data and account data was exactly what allowed a hidden instruction in a public web form to result in exfiltrated CRM records. No rule was broken. The access simply should not have sat in one place.

The board-level question is not "is this agent secure?" It is "if this agent's judgement failed completely tomorrow, what could it reach?"

2. What can it do without a human approving first?

Meta's AI security team has proposed a useful working rule for this, referred to as the Agents Rule of Two: an agent session should carry no more than two of three properties at once, namely processing untrusted input, accessing sensitive data, and taking an action that changes state or leaves the organisation. Where a task genuinely needs all three, a human should be in the loop.

For a board, the practical version of this question is a short list: which actions can this agent take entirely on its own, and which of those, if wrong, would be expensive, embarrassing or irreversible?

3. How was it tested, and was that testing adversarial and repeated?

A model that refuses a harmful request once, when asked directly, has told you very little. Research published jointly by OpenAI, Anthropic and Google DeepMind in 2025 tested twelve recently published AI defences using adaptive, multi-turn attacks and defeated most of them with success rates above 90%, despite the majority of those same defences originally reporting near-zero success rates under standard, single-turn testing. Human red-teamers in the same study, working adaptively rather than from a fixed prompt list, defeated every defence tested.

The question to ask is not whether the agent was tested, but whether it was tested by someone actively trying to break it over multiple turns, using the kind of gradual escalation and persuasion techniques documented in recent AI security research, rather than a single static prompt list.

4. Who re-tests it after every change, not just before launch?

Prompt changes, model upgrades, new tool integrations and new data sources all change an agent's behaviour, sometimes in ways nobody predicted. A defence validated in January can fail in July, not because it broke, but because everything around it changed. The relevant governance question is whether there is a named owner responsible for re-testing after every material change, not only before the original go-live date.

5. How would we know if it has already been manipulated?

Unlike a traditional breach, a manipulated agent does not necessarily throw an error. It keeps working, and keeps looking normal, while quietly doing something it should not. This is the essence of memory and context poisoning, a risk category the OWASP Top 10 for Agentic AI names explicitly: a single successful manipulation written into an agent's memory or shared context can go on to influence every subsequent decision that agent makes, and every other agent that later reads the same memory, without needing to be attacked again.

Boards should ask what monitoring exists to catch this kind of drift, not only what defences exist to prevent the initial manipulation.

6. If it does something harmful, how do we contain it, and how fast?

Every organisation with an incident response plan for a data breach should ask whether that plan actually covers an autonomous agent. Can the agent's access be revoked immediately, system-wide, without waiting for a change ticket? Is there a kill switch that someone is actually authorised and available to use? The Microsoft 365 Copilot vulnerability known as EchoLeak and the SalesBleed vulnerability in Salesforce Agentforce were both patched quickly once identified. The organisations affected were not the ones who wrote the patches. Containment speed on the deploying organisation's side is a separate capability from the vendor's ability to fix the underlying flaw, and it deserves its own answer.

7. Which regulatory regime applies to this system, and are we positioned for it?

This is the least stable of the seven questions, and boards should treat it as a standing item rather than a box to tick once. Three frameworks are converging on agentic AI from different directions.

NIST AI Risk Management Framework (United States, voluntary)

Its Govern function requires organisations to assign named accountability for AI risk, rather than leaving it to whichever team happens to build the agent. It does not currently distinguish systems by degree of autonomy, which is why several extension proposals for agentic-specific oversight boundaries are circulating.

ISO/IEC 42001 (international, voluntary, certifiable)

The first international standard against which an organisation can be independently certified for how it manages AI risk. It requires a documented AI impact assessment, named leadership accountability and continual audit, structured similarly to ISO 27001 for anyone already familiar with that standard.

EU AI Act (binding within the EU, and extraterritorial for many deployers)

High-risk system obligations under Annex III were legislated to apply from August 2026. A Digital Omnibus proposal has sought to delay this deadline to December 2027, and as of this article's publication that proposal had not completed the full legislative process. Organisations operating in or serving the EU market should treat this as an actively moving deadline and confirm current status directly, rather than relying on any single date.

None of these three frameworks alone fully covers agentic risk yet. Between them, they give a board enough structure to ask whether accountability, documentation and audit exist at all, which is a lower bar than compliance but a meaningfully higher bar than most organisations currently clear.

The pattern underneath all seven questions

Every incident referenced in this article, EchoLeak, SalesBleed, the adaptive jailbreak research, the memory poisoning research, shares a single underlying shape. The agent did not misunderstand its instructions. It followed them, exactly as designed, from a source nobody had authorised. Access control, human approval on irreversible actions, adversarial testing repeated over time, and a genuine ability to contain a compromised agent quickly are not exotic security measures. They are the ordinary controls any organisation would apply to a new employee with broad access and no track record. Agents deserve the same discipline, applied consistently rather than once at launch.

References

Anil, C. et al. (2024)

'Many-shot jailbreaking', Advances in Neural Information Processing Systems, 37.

Dataiku and The Harris Poll (2026)

Global AI Confessions Report: CIO Edition, 2026. Survey of 685 global CIOs, July 2026.

Deloitte (2026)

State of AI in the Enterprise, and associated board governance surveys. Survey of 3,235 business and IT leaders across 24 countries.

ISO (2023)

ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system. Geneva: International Organization for Standardization.

Meta AI (2025)

Agents Rule of Two: A Practical Approach to AI Agent Security. 31 October.

Nasr, M. et al. (2025)

The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections. arXiv:2510.09023.

National Institute of Standards and Technology (2023)

AI Risk Management Framework (AI RMF 1.0). Gaithersburg, MD: NIST.

OWASP Foundation (2025)

OWASP Top 10 for Agentic AI Security Risks.

Zenity Labs (2026)

SalesBleed: Indirect Prompt Injection and 0-Click Data Exfiltration on Agentforce. 24 September.