
“The failures I worry about most do not require an agent to behave maliciously. An agent can misunderstand an instruction, act on bad context, or follow manipulated input while still operating within the permissions we gave it.”
Naga Chand Putta, Senior Software Engineering Leader in a highly regulated industry.
Engineering teams that put LLM agents into internal workflows tend to review prompts and model outputs closely while giving far less attention to the identity each agent runs under. Naga Chand Putta, a senior software engineer working in a regulated industry, builds LLM agents alongside the services those agents connect to and has published research on secure cloud architecture for financial systems. From his engineering purview, he argues that the permissions behind an agent deserve at least as much scrutiny as the prompt in front of it. “The failures I worry about most do not require an agent to behave maliciously,” he says. “An agent can misunderstand an instruction, act on bad context, or follow manipulated input while still operating within the permissions we gave it.”
That concern shapes how Putta and his colleagues handle every agent before they deploy it to run live workflows in production. Each review starts with the systems an agent’s identity can reach, and the team strips out any access the workflow cannot justify before the agent moves forward. That step matters more in a regulated industry, where a small permission gap or configuration error can expose sensitive data or move money. And because every permission that survives the review has a clear purpose, engineers can defend the setup in front of risk and compliance boards.
Authentication shouldn’t hand out broad authority
Much of the security conversation around AI agents centers on credentials, and short-lived [ephemeral in nature] tokens do shrink the damage a leaked key can cause. But Putta treats the token as the first half of the problem. “That approach reduces the risk of a leaked credential, but the token itself is only part of the control,” he points out. “We also define what the agent is allowed to do after it authenticates,” he continues. Say, for example, if a workflow only requires the agent to read information, “we do not give,” he clarifies, “the same identity permission to modify records, approve changes, or trigger downstream actions.”
Engineers working alongside him review those permissions as a separate step from authentication, and he underscores why that separation matters. “We review those permissions separately because successful authentication should not automatically translate into broad authority,” he shares. And this is because a valid token proves which agent is calling, but it says nothing about what that agent should be allowed to change.
For teams putting their first agents into production, the practical move is to write down the exact verbs each workflow needs, such as read, create, update, or approve. Then grant only those verbs to the agent’s identity and leave everything else out. Least privilege [granting an identity only the access its job requires] is an old principle, and OWASP’s guidance on excessive agency applies it directly to LLM agents. And agents make it urgent, because an agent acts on whatever instructions reach it, and its permissions set the outer limit on what those instructions can do.
Code review agents don’t need merge rights
During our email interview via Qwoted, he points to a workflow that looked low risk on paper and turned out to carry far more authority than anyone intended. “We ran into that issue when we deployed an agent to assist with pull request reviews,” he recalls. “The service account initially had read access to the codebase and also had permission to merge approved changes. Once we tested the workflow more closely, it became clear that the merge permission gave the agent more authority than the review task required.”
He explains that the agent could inspect code well enough, but it had no reliable way to tell if a technically valid change matched the business intent behind the application. So the team made two changes. They removed the merge permission from the agent’s account and kept final approval with a human reviewer. They also required the agent to cross-reference business documentation from domain owners before it produced a review.
“Those two changes solved different problems,” he highlights. “The business documentation gave the agent more context for its analysis, while removing merge access made sure it could not turn that analysis into a production decision on its own.”
The same fix applies well beyond code review. Any agent that produces a recommendation should hand that recommendation to a person or a separate service for execution, so a wrong call stays a draft instead of becoming a deployed change. And better context should improve the quality of the agent’s work without ever expanding what the agent can change.
Let reversible actions run and gate the risky ones
Now, depending on the risk behind each task, the pull request lesson should also help teams decide how much independence an agent should get. “We do not treat a task as low risk simply,” he cautions, “because it is technical. Because a technical action can still cause serious damage if it changes production data, deploys code, modifies infrastructure, or calls a sensitive API.”
Before an agent executes anything on its own, the team looks at the privileges the action requires and the damage a wrong decision could cause. Engineers also check how easily they can reverse the action and how much context the agent has to validate its own call. Many engineers describe that combination as blast radius [the scope of damage one bad action can cause], and Putta turns it into a practical gate as it seems to be the most reasonable and logical way to separate safe actions from risky ones. “For actions with limited impact and a clear rollback path, we can allow more automation,” he reasons. Say, if an action affects production systems, sensitive data, money movement, or another high-impact workflow, “we narrow the permissions and add an approval step before execution.”
Some decisions need a person for a different reason entirely. “And when the decision depends on business or domain context that the agent cannot verify on its own, we keep a human in the approval path rather than expecting the model to fill in that missing judgment,” he adds.
Teams can put this into practice by sorting every agent action into tiers before launch. Reversible actions with a small footprint, such as drafting a summary or labeling a ticket, can run automatically as long as the system logs every one of them. Anything that writes to production or moves money gets narrower permissions and a named human approver. And if an action depends on business knowledge the model has no way to verify, a person should make the final call, even when the task looks routine.
Map agent access before every release
Tiering tells a team which actions need tighter control, but it only works if the team knows every action the agent can actually take. Putta and his colleagues close that gap with a release check. “One practical check has become especially useful for us before release,” he notes. “We map every system the agent can reach and every action its identity can perform, then compare that list with the actions the workflow actually needs. Anything that does not have a clear purpose gets removed.”
For sensitive actions, the team pushes the check further and deliberately feeds the model bad input. “For sensitive actions, we also test what happens when the model receives malformed, misleading, or hostile input so we know the permission boundaries still hold when the model behaves unexpectedly,” he explains, providing a repeatable way to test those boundaries before launch.
This check works best as a standing gate in the release process instead of a one-time audit. Pull the agent’s live IAM policies and API scopes, list them next to the actions the workflow requires, and treat every unexplained entry as a defect that blocks release. Then repeat the comparison after every change to the agent’s tools or integrations, since new connections often come with broader default access than anyone requested.
Make engineers justify every extra permission
Putta’s approach starts from a realistic expectation that agents will sometimes get things wrong even when nobody attacks them, and it puts the effort where that expectation leads. “That is why we spend so much time reviewing those permissions before production,” he underscores. In his view, the safest design is to give the agent only the access needed for the task and require the engineering team to justify every additional capability.
“That last requirement moves the burden to the right place,” Putta posits. Engineers make the case for each permission before launch, so the agent enters production with the smallest footprint its workflow allows and security teams have less excess access to hunt down later. In regulated industries, that written record of justified access also gives risk and compliance reviewers something concrete to evaluate when the next AI feature comes up for approval.Â
Putta further unpacks what regulators already ask for in his useful piece on AI agent access rules under NYDFS, DORA and the EU AI Act.


