What Risk Leaders Should Be Asking When AI Is Building Software

Agentic coding is the use of AI to develop software, and businesses are rapidly adopting it. In 2025, Stack Overflow's survey put professional developers actively using agents at about 32 percent, with 14 percent using them daily. JetBrains' 2026 Developer Ecosystem Survey, fielded between May and July 2026 with more than 15,000 professional developers, found 90 percent using AI coding agents at work at least weekly and 68 percent using them daily. While each survey covered somewhat different populations, the trend is clear: agentic coding has fast become the norm for enterprise software development.

As risk leaders, how should we think about the impact of agentic coding? Start with what Anthropic says about its own internal use. Per the August risk report, Claude now authors a large majority of the code merged into its production codebases. Anthropic also runs a Claude model over proposed code changes, looking for common errors, obvious security vulnerabilities, and mismatches between what a change claims to do and what it actually does. It maintains automated tests that verify security invariants, meaning properties that have to hold no matter what else changes, such as two systems that are never supposed to talk to each other. It sandboxes agent work.

That says a lot about the trust Anthropic places in its own coding agents. The company that builds the model does not take their code at face value, and the checks it describes cover quality and security, not only alignment. If Anthropic won't merge (accept) its agents' output unchecked, that's a floor for the rest of us rather than a ceiling.

Agentic coding does not automatically translate to new scenarios

For the purpose of risk analysis, coding agents differ from the agentic threat actor I described in last week's post. The concern isn't a "privileged insider" acting against us operationally. It's the agent's effect on the controls embedded in our software. A coding agent potentially changes the strength of controls that live in code: authentication, authorization, input validation, rate limiting, logging, encryption, and even business logic.

Every loss event frequency (LEF) estimate you've produced rests on the controls working about as well as they did when you last assessed them. Every risk assessment also has a time horizon, typically a year, and control effectiveness has to hold across it. When your company adopts agentic coding, the control effectiveness you've assessed for your scenarios may decrease as a result of factors I'll discuss below.

This means that agentic coding rarely necessitates the identification of new risk scenarios. Instead, it requires you to reconsider the software-based controls that are relied upon in the scenarios you already have. That's an unusual risk implication for the adoption of a new and transformative technology (which often bring about new scenarios), and it's why its risk impact may be overlooked.

A public case illustrates the point, though it predates modern agentic coding. Australia's communications regulator alleges that a coding error introduced in 2018 disabled an access control on an Optus interface. Optus found and fixed the same error on its main website in 2021, but not on a subdomain. An attacker walked through the gap in September 2022 and took personal data belonging to roughly 9.5 million people. The regulator says Optus had three chances to catch the error and missed all three.

An ordinary change weakened a control, nobody caught and corrected it, and an attacker took advantage of the gap. No new scenario is inferred from this incident. The risk grew because a software defect changed susceptibility.

Coding agents are fallible at greater scale

How does agentic coding affect the controls embedded in our software? Arguably, AI could improve the quality and security features of our code and result in lower susceptibility. Recent testing suggests otherwise, and both technology and risk leaders should be cautious.

Consider the assessment by DryRun Security. They built features with Claude, Codex, and Gemini and found vulnerabilities in 26 of 30 pull requests, which are the proposed code changes developers submit for review before merging. The recurring gaps were insecure JWT defaults and state management failures, missing brute-force protection and rate limiting, and refresh tokens that couldn't be revoked. Broken access control was the most universal failure, appearing with all three agents in both applications. Read those numbers with one detail in mind: the agents built the two applications from feature specifications that deliberately included no security requirements, which is closer to how most prompts are written than to how a regulated codebase is maintained.

These gaps don't always come from coding mistakes. The researchers found that many of the problems originated in design decisions the agents then implemented, which puts them upstream of what pattern-matching scanners look for. But the lesson is that agents may not catch vulnerabilities during coding.

The broader data points the same way. In 2025, Veracode tested more than 100 models on security-sensitive coding tasks and found roughly 45 percent of generated samples introduced one of the ten most common classes of web application vulnerability. Its 2026 update found that rate essentially unchanged, while the same models kept improving on coding benchmarks. In other words, models got better at writing working code without getting better at writing safe code. CodeRabbit reviewed 470 open-source pull requests and found the AI-authored ones carried about 1.7 times the issues of the human-authored ones.

Take all of these with a grain of salt: they are vendor studies, the agent-specific samples are small, and none of them establish anything about your company's codebase. What they establish is that coding agents are fallible in ways that matter to controls. Whether they are more fallible than human developers per change is still unsettled.

Volume, of course, multiplies any risk, and agentic coding drives up the volume of changes. Researchers at Stanford and Carnegie Mellon produced a longitudinal study of one enterprise's AI coding mandate tracked 802 developers and 196,212 pull requests from January 2024 through April 2026. Total pull request volume grew 3.1 times while the number of people acting as reviewers grew 1.5 times, and by the end of the period roughly 90 percent of pull requests were AI-labeled. Faros AI, reading two years of telemetry from 22,000 developers across more than 4,000 teams, found pull requests 51 percent larger, files touched per developer per month up about 150 percent, and 31 percent more changes merging with no review at all, human or automated. Faros sells engineering intelligence software, so read its numbers with that in mind.

If you hold the defect rate steady and raise the volume of releases, you get more defects. That presents a different kind of challenge for our control assessments.

Agentic coding presents a variance management problem

FAIR-CAM separates the controls that act at the moment of a loss event from the controls that keep those first controls working. The second group are Variance Management Controls (VMCs): (1) preventing conditions that degrade a control, (2) identifying them when they occur, and (3) correcting them. (Remember from FAIR-CAM that 2 and 3 are connected by an "AND" condition; correction only happens if detection happens first.)

Agentic coding can weaken all three of these "controls over our controls":

  • Prevention weakens when change frequency rises or when the probability that any given change introduces a variant condition rises.
  • Identification weakens when the capacity to catch those conditions doesn't rise in proportion.
  • Correction weakens, or simply slows, when the queue of defects grows due to a larger stream of new work.

To get a better idea of the impact, consider Cortex's 2026 benchmark, which combines development metrics from multiple organizations with a survey of more than 50 engineering leaders, measured year over year from Q3 2024 to Q4 2025. It found pull requests per author up 20 percent, incidents per pull request up 23.5 percent, and change failure rate up roughly 30 percent. Change failure rate is the share of deployments that end in a rollback. The same report found fewer than half of the organizations surveyed had a formal policy governing AI use (a Decision Support Control), thus compounding the risk.

Clearly, you must reassess the VMCs you've included in any FAIR-CAM analysis where agentic coding is now developing or maintaining software-based loss event controls.

Modern development practices automate VMCs

Software organizations have spent the past decade automating the path a code change takes to reach production. DevOps is the practice of development and operations working as one process rather than handing work across a wall. Continuous integration means every change is merged and tested automatically, many times a day. Continuous delivery means a change that passes those tests is packaged and released the same way every time. The automated sequence that carries all of this is called a pipeline.

For a risk assessor, that machinery is variance management rendered in software. Automated tests and security scanning identify variant conditions as they are introduced, often within minutes of the change that caused them. Required checks, or gates, block a change that fails from going further, which is prevention. Version control, infrastructure as code, which means servers and permissions defined in files that are reviewed like any other code, and automated rollback make changes visible and reversible, which is what correction depends on. None of this is new, and agentic coding doesn't change what the pipeline is for. It changes the volume moving through it.

Two properties of that automation matter for what follows. The first is that it scales with compute rather than headcount. Static and dynamic application security testing (SAST and DAST), dependency and secret scanning, and policy checks on infrastructure code can absorb a doubling of change volume in a way that human review cannot, which is why automation is the only honest answer to agentic velocity. The second is that the coverage is uneven. These tools are strongest on injection and configuration flaws and weakest on business logic, where the flaw is a violation of what the application is supposed to do rather than a pattern a scanner can match. OWASP's verification standard reflects that, requiring manual testing of business logic, authentication, session management, and access control at its higher assurance levels. Authorization falls in between, since dynamic testing against a running application catches some access-control failures that static analysis of the code will not.

For shops that have adopted these practices, a risk assessment must closely evaluate the automated VMCs in place, then identify which controls from your scenarios can't be tested that way. As stated above, many require manual testing.

Indeed, risk may increase the most in organizations that adopt agentic coding without the proper automation and processes to ensure software quality at increased scale (velocity). Those that have already employed the modern software development methods described above are likely better positioned to absorb the impact of agentic coding on software quality. For those that have not, risk leaders should take a closer look at how these risks are being addressed.

Questions to assess variance management

The following eight questions will help you assess the effectiveness of the VMCs provided by your agentic coding pipelines.

  1. How many changes are merging now, and how does that compare with a year ago? Change frequency is a variance prevention factor in FAIR-CAM, because every change is an opportunity for a control to drift. The rate of merged changes, in total and per developer, sets the baseline exposure regardless of who or what wrote them. Ask for enough history to see the trend. A codebase absorbing twice the rate of change it did a year ago needs more identification capacity than it had, and that gap is the finding.
  2. What share of merged changes is agent-generated? Ask this one second, and don't read it as a quality judgment. The public comparisons between agent-written and human-written code disagree with each other, so assuming agents are worse per change isn't supported. What the share gives you is a sampling frame: it tells you where to look when you measure your own defect rate, and it explains a rising change rate. A share climbing quarter over quarter, with no corresponding change in the gates, is the pattern to flag.
  3. Which checks are required before a change can merge, and what share of merged changes actually passed each one? A gate isn't the test itself. It's the rule that stops a change from advancing when the test fails, so a scan whose failures are purely advisory (i.e., they don't sway a decision) is a test without a gate behind it. Look for gates that exist in policy but get bypassed in practice, and for pass rates that slipped as volume rose. Request and review evidence of the decisions being made.
  4. Do the tests cover the specific controls your analyses depend on? Don't accept a generic statement that the team runs static and dynamic scanning. Ask to see the test cases for the authorization rules, rate limits, and logging that your scenarios assume are working.
  5. For code in production, how long does a flagged defect wait before someone looks at it? That interval is your variance identification time, and it only means something next to how fast a loss would accumulate. If nobody can produce the number, treat the interval as unbounded rather than short.
  6. What happened to last quarter's findings? Evidence of correction looks like a defect backlog with ages and severities attached, the share fixed rather than waived, expiry dates on the exceptions granted, and proof that each fix was re-tested rather than closed on someone's word. You cannot count on a control that is supposedly repaired but not verified (tested).
  7. Who designs and runs the tests, and who can change pipeline configuration or branch protection? If the same agent writes the feature and its tests, the check and the thing being checked share an author. It's common to see the same tests applied during coding and again during testing, but independent testing is a must. Look for test suites owned by someone other than the team shipping the change, and ask whether any path to production bypasses the pipeline entirely.
  8. When was the strength of each code-resident control last validated by a test rather than carried forward from a prior assessment? This closes the loop back to your analyses. If the answer predates the organization's adoption of coding agents, every estimate resting on that control is carrying an assumption nobody has checked since. Environment segregation belongs in the same conversation. Keeping development, test, and production separate doesn't make a bad change less likely, but it bounds what a change can impact, which limits loss magnitude rather than loss event frequency.

Get help and get started

If some of this sounds like unfamiliar territory, you're not alone. I've not met many risk professionals who have software development backgrounds. Those who do may not be up on the latest software development methodologies, tools, and practices. Therefore, talk to your CTO or equivalent to get help in understanding how your organization develops, tests, and releases software that provides the controls you've incorporated into your risk assessments.

Then, take two or three risk scenarios that depend on loss event controls embedded in homegrown software. Identify the development teams and pipelines responsible for the software and evaluate the testing and correction of software defects, especially for the controls included in your assessments. This will tell you a lot about what has or hasn't changed and give you a better idea of how to assess the VMCs of other scenarios.

If your organization has begun to adopt agentic coding, it's time to act. The delta in terms of volume and velocity of changes, skillsets and practices, tools, and other aspects of your software pipelines is simply too big to ignore. Assuming the variance rates of your controls are and will remain the same as before is a risky move.

image 37