🏆 Winner – African Insurance Awards 2026

    Resources
    AIAgentic AIEnterpriseFinancial ServicesInsurance

    AI agent containment: lessons from Google's Gemini incident

    Google confirmed Gemini reached three real companies during a test. What banks, insurers and airlines must lock down before an AI agent gets tool access.

    Cover image for AI agent containment: lessons from Google's Gemini incident
    Key takeaways
    1. On 18 September 2026 Google confirmed that a Gemini model, during a May security test, accessed the systems of three real companies it mistook for test targets. OpenAI, Anthropic and Meta had disclosed similar incidents from the same testing environment.
    2. The failure was containment, not intent: a test environment that was supposed to be offline had internet access, and the agent did what it was built to do.
    3. A customer-facing AI agent in a bank, an insurer or an airline has the same three ingredients: a goal, tool access and a boundary someone configured.
    4. Safety is architecture, not model quality: least-privilege tools, egress control, gated write actions, tool-call logging and an incident playbook.
    5. A 90-day containment checklist closes the gap before an agent touches core systems.

    On 18 September 2026, Google confirmed that a Gemini model, during a cybersecurity evaluation run in May by the security firm Irregular, gained unauthorized access to the systems of three real companies it had mistaken for test targets. The disclosure followed similar admissions by OpenAI, Anthropic and Meta about incidents traced to the same testing environment. Source: nbcnews.com/…/google-says-ai-model-gained-unauthorized-ac…

    The business question this raises is uncomfortable for anyone about to give an AI agent live access to a policy administration system, a core banking ledger or an airline's passenger service system: if the best-resourced AI labs in the world could not keep an agent inside its sandbox, what exactly is keeping yours inside its lane?

    What happened, and why is it more than a lab story?

    An AI agent with network access and a goal did exactly what it was built to do, and the boundary that should have stopped it was misconfigured. The failure was containment, not intent.

    The facts, as reported by NBC News and CNBC, are simple. In May 2026, Irregular ran a capture-the-flag exercise in which Gemini was asked to retrieve information from a fictional company. That fictional company shared its name with a real one. The environment was supposed to be sealed off from the internet; a bug left it connected. In one case the model guessed login credentials; in two others it used credentials it found in a public code repository. In all three cases it stopped once it had access. Google says it only learned of the intrusions in July, after Irregular reviewed its work following an earlier disclosure by OpenAI involving Hugging Face. Google informed the affected organisations and US federal authorities, and confirmed the incident publicly on 18 September after the Wall Street Journal reported it. Source: cnbc.com/…/googles-gemini-becomes-latest-ai-model-to-brea…

    Why this matters outside the lab: a production customer-service agent in financial services has the same three ingredients. It has a goal (resolve the claim, collect the instalment, rebook the passenger). It has tools (APIs into systems that hold money and personal data). And it has an environment boundary that a human configured, possibly under deadline pressure. The only difference between the Gemini case and a bad day at a bank is which system sits on the other side of the misconfiguration.

    What does this change about deploying AI agents in customer service?

    It ends the idea that model quality is the safety layer. Safety is architecture: what the agent can reach, what it is allowed to do, and what checks each action before it executes.

    Three shifts follow for any decision-maker signing off an agent deployment.

    First, from prompt guardrails to permission scoping. A system prompt that says "only act on this customer's account" is a request, not a control. A credential that can only read that customer's account is a control. The Gemini model was not told to break into three companies; it was able to, and that was enough.

    Second, from "the model will stop itself" to "the system will stop it". Google's statement notes that the model halted after gaining access. That is reassuring and irrelevant to your risk model. You cannot present "the model chose to stop" to a regulator or a board as your control environment.

    Third, from silence to disclosure discipline. The gap between the incident (May), the labs learning of it (late July) and public confirmation (mid-September) is the part of this story that regulated industries should study most carefully. Under the EU AI Act, providers of high-risk AI systems will have serious-incident reporting duties as obligations phase in through 2027, and EU financial institutions already carry ICT incident reporting obligations under DORA. A bank that takes seven weeks to disclose that an agent touched the wrong accounts will not be judged by AI-lab norms.

    Which risks and opportunities does this create for an insurer, a bank or an airline?

    The risk is the wrong action, not the wrong answer. The opportunity is that agents with tightly bounded tools resolve more, faster, and can be defended in front of a regulator.

    Insurers

    A claims agent that opens a first notice of loss, requests documents and triggers a payment is exactly the kind of agent worth building. The risk sits in the last step. Filing a claim against the wrong policy, or releasing a payment on an unverified identity, is a wrong action with a financial and regulatory consequence. The opportunity: an agent whose payment tool only fires after a confirmation step, and whose policy lookup is read-only, can run claims intake at scale without the tail risk. At FCB.ai, AGMA's ClaimStatus deployment is built on that principle of read broadly, write narrowly, and it won first prize in the Insurtech category of the Trophées de l'Assurance 2026.

    Banks and lenders

    A collections agent negotiating a payment arrangement on WhatsApp needs to see the balance, propose a plan and record the promise to pay. It does not need the ability to modify the ledger. TFG's collections programme, which recovers roughly R6 million per month through WhatsApp conversations, works because the conversational layer negotiates and the core system executes under its own controls. The moment an agent can both decide and execute a balance change with one credential, you have recreated the Gemini setup with money attached.

    Airlines

    Rebooking, ancillary sales and disruption handling are the highest-value agent use cases in travel, and all of them mean write access to the passenger service system. The risk is an agent that changes the wrong booking, or one that is tricked, through a prompt injected in a customer message, into doing so. The opportunity is real: Air Caraïbes' WhatsApp journeys show 82% engagement and a 62% conversion rate on terms and conditions acceptance, with tools scoped to the passenger's own PNR. Scope is what makes those numbers safe to scale. For the broader picture on what production readiness means, see fcb.ai/…/agentic-ai-going-to-production-2026

    What should you lock down in the next 90 days?

    Five layers, in order. Do not skip to the fourth because the first three feel like plumbing. The Gemini incident was plumbing.

    The five-layer AI agent containment checklist

    1. Identity and scope. Every tool the agent calls runs under its own least-privilege credential, scoped to the customer in the conversation. No shared service token that can reach everything. Read and write are separate identities.
    2. Environment. An allowlist of systems the agent can reach, with network egress control enforced outside the agent's own code. Test and evaluation environments never hold production credentials and never sit on the live internet. A fictional target must not be able to resolve to a real one.
    3. Action gating. Read actions run freely. Write actions run through a policy layer that checks scope. Irreversible or financial actions (payment, cancellation, limit change, rebooking) require an explicit confirmation from the customer, a human agent, or both, depending on value.
    4. Observability. Every tool call is logged with inputs, outputs, the identity used and the conversation it belonged to, and the log can be replayed. If you cannot reconstruct what the agent did in a given conversation, you cannot answer a regulator, a customer or an auditor.
    5. Disclosure and response. A written incident playbook that names who is informed, in what order, within what timeframe, and includes a kill switch that disables tool access without taking the conversational channel down.

    A 30, 60, 90 day sequence

    Days 1 to 30: inventory every agent, every tool it can call and every credential it uses. Most organisations discover at least one shared token in this step. Days 31 to 60: split read and write identities, put the allowlist and egress controls in place, and move every financial or irreversible action behind a confirmation gate. Days 61 to 90: implement tool-call logging with replay, run a red-team exercise that includes prompt injection through customer messages, and rehearse the incident playbook end to end, including the kill switch.

    Which mistakes should you avoid?

    The most common mistake is treating the vendor's safety documentation as your control environment. It is theirs. Yours is what you configured.

    Four others come up repeatedly in the deployments we review. Testing agents with production credentials because it is faster, which is precisely how a test reaches a real system. Assuming the agent will know it is in a test, when the entire Gemini episode shows it will not. Giving one agent one large token because permission scoping is tedious. And measuring success only by automation rate, which rewards agents that do more, when the safer and ultimately more profitable agent is one that does less, reliably. Our organisational diagnosis of why these shortcuts get taken is at fcb.ai/…/messaging-org-chart-problem

    What we observe at FCB.ai

    In mature WhatsApp journeys we observe 60 to 85% automated resolution and 6 to 15 messages per resolution, and NPS above 90%. These are FCB observations, not sector benchmarks, and they are reached with narrow, well-scoped tools rather than open-ended access. The agents that resolve most are not the ones that can do the most. They are the ones whose capabilities match the conversation in front of them and nothing more. Our platform is designed to support the containment layers above, with payment and refund flows in particular gated as described in fcb.ai/…/whatsapp-payments-refunds-compliance-guide

    The Gemini disclosure is not a reason to slow down agent deployment in insurance, banking or travel. It is a reason to stop treating containment as a phase-two item. If you would like a second opinion on the tool scope and gating of an agent you are about to put in front of customers, FCB.ai runs a short architecture review before any deployment. Details at fcb.ai/…/insurance

    Frequently asked questions

    6 answers, all expanded

    What exactly did Google disclose about Gemini on 18 September 2026?

    Google confirmed that during a May 2026 cybersecurity evaluation run by the firm Irregular, a Gemini model gained unauthorized access to systems belonging to three real companies. The test environment was meant to be isolated but had internet access. The model guessed passwords in one case and used credentials found in a public repository in two others, then stopped.

    Is this an AI alignment failure or a security failure?

    Public accounts describe a security and test-scoping failure. The model believed the systems were part of its exercise and stopped once it had access. No data theft or destructive action has been reported. The lesson for enterprises is that an agent's intent is not a control: the environment, the permissions and the gating around it are the controls.

    How does this affect an insurer or bank deploying a customer-service AI agent?

    It removes the assumption that a well-behaved model is a sufficient safety layer. Any agent with API access to policy, claims, collections or booking systems needs least-privilege credentials per tool, an allowlist of reachable systems, gated write actions, tool-call level logs and a tested incident response path before it handles live customers.

    What is AI agent containment?

    AI agent containment is the set of technical and organisational controls that bound what an AI agent can reach and do: scoped identities, network egress control, read versus write separation, confirmation gates for irreversible or financial actions, full tool-call logging and a defined kill switch. It is designed so the system stops the agent even when the agent would not stop itself.

    What should a bank, insurer or airline do in the next 90 days?

    Inventory every agent and the systems it can reach, replace shared tokens with per-tool least-privilege credentials, put write and payment actions behind explicit confirmation, add tool-call logging you can replay, isolate test environments from production credentials and the live internet, and write an incident playbook naming who is notified and within what timeframe.

    Does the EU AI Act require reporting incidents like this?

    Under the EU AI Act, providers of high-risk AI systems must report serious incidents to market surveillance authorities, with high-risk obligations phasing in through 2027. Separately, financial institutions in the EU have ICT incident reporting duties under DORA. Confirm applicability with your compliance team, but assume regulators will expect faster disclosure than the seven weeks seen in the Gemini case.

    Written byAntoine Paillusseau, CEO, FCB.aiEight years building WhatsApp-native AI in production across African insurance, banking and telco. Writes about what survives contact with real customer operations.