It Wasn't Me. It Was the AI
That might work as an excuse. It will not survive an incident review.
An AI assistant suggests insecure code. An agent sends data to the wrong system. A model gives a confident but completely wrong answer. Who is responsible?
AI is a tool. The people who choose it, configure it and approve its work remain accountable for the result.
Responsible AI can sound like a new discipline requiring an entirely new rulebook. Most of it does not. The tools have changed, but the basic rules for data, access, testing and accountability still work.
There's a new agent in town. The old rules still apply.
Production customer data and personally identifiable information should not be copied to developer laptops, local test environments or unapproved services. That rule predates generative AI. A prompt is simply another data transfer, whether somebody pastes in a support ticket, attaches a log or lets an assistant index a repository.
Access works the same way. A developer assistant does not need production credentials because it can write useful code. Give it only the files, tools and permissions needed for the task. Keep access scoped, short-lived where possible and easy to revoke.
Who gave it the keys?
It makes little sense to approve or reject a tool merely because somebody has labelled it "AI". A code-authoring agent confined to development repositories and sandboxed tools presents a different risk from an operational agent with customer data and permission to update production. Govern what it can access and do, not what it is called.
Models can be confidently wrong, manipulated by untrusted text and able to combine harmless-looking permissions into a harmful action. Any agent that can read sensitive data, call privileged APIs or change live systems is a production workload. It should not run on a developer's laptop. It belongs in a controlled environment with managed identity, scoped credentials, restricted network access, central logs and enforceable tool policies.
The controls are familiar: validate input, test output, restrict permissions, record important actions and provide a reliable way to stop the agent. Human approval only helps when the reviewer has enough context and time to challenge the action. Clicking "approve" on hundreds of opaque requests is theatre, not control.
Put down the AI hammer
Not every problem needs AI. If a task has known inputs, clear rules and an expected output, use ordinary code or workflow automation. It will be cheaper, faster, easier to test and more predictable than asking a model to work it out each time.
Models earn their place when the work contains real ambiguity, such as summarising unstructured incident evidence, classifying messy text or drafting options for a person to review. Even then, let the model handle the uncertain part while deterministic code validates data, enforces policy and performs the final action.
Adding an LLM to every process creates cost and uncertainty without necessarily adding value. Good architecture uses the smallest amount of AI needed to solve the problem.
Start useful, then expand
For operational agents, read-only investigation is a good place to begin. An agent can gather alerts, logs, metrics and recent changes, then prepare an evidence-backed summary or draft ticket. That saves time without giving it permission to alter production.
Write access can come later, one action at a time, after the read-only use case has proved useful and safe. Rollouts should be staged, with a small initial group, clear measures and a way back. Feature flags and rollback plans are no less useful because a model is involved.
Buying licences or launching a pilot is not transformation. Each use case needs a business problem, an owner and a measurable result. Shorter lead time, fewer defects or faster incident resolution tell us something useful. Login counts and token usage do not. Adoption sticks when teams understand why their work is changing, where the boundaries sit and when they must step in.
Run agents like production systems
An agent does not get a reliability exemption. Validate data at every tool boundary. Bound retries, timeouts and concurrency. Respect Shopify and other external API limits. Make actions idempotent where possible, and design for dependencies to fail without taking down the customer journey.
This matters most during high-traffic events. An investigation agent must not consume API capacity needed by customers or amplify a small failure through a storm of retries.
Observability should combine standard service measures, such as latency, errors, traffic and cost, with an audit trail of model requests, tool calls, policy decisions and approvals. That trail must not copy sensitive data into logs. When an agent fails, the incident should result in a lasting change to a test, evaluation, permission, policy or architecture. This is how we already improve production systems.
Keep the doors locked. The agent is coming.
Some agents need to run for longer, use several tools or work with sensitive internal systems. Amazon Bedrock AgentCore provides useful building blocks for those cases.
A policy can be more precise than simply allowing or denying an API. An agent might be allowed to call a refund tool, while a deterministic policy blocks refunds above $100 and sends them to a person for approval. The model cannot talk its way around that limit because the gateway enforces it outside the agent.
Those services do not remove our responsibility. IAM permissions still need to be narrow, inputs still need validation and callers must not be able to bypass the gateway. CloudTrail and CloudWatch can record who invoked an agent and what it did, but the logs need their own data controls.
For workloads that need an AWS boundary, models can be accessed through Amazon Bedrock rather than a provider's public API. AWS states that model providers cannot access Bedrock service logs, customer prompts or completions. An agent can still send information to an external tool if we give it permission, so "the data stays in AWS" must be an enforced architecture, not a slogan.
AI can help teams move faster through maintenance, testing, security work and operational investigation. It can also make mistakes faster. We do not need to choose between banning it and trusting it blindly. Give it useful work, the minimum access required and the same engineering discipline expected everywhere else.
If you're giving AI agents real work and want help doing it without handing them the keys, speak to us at [email protected].