Overcoming the
limitations of coding
agents in complex
software systems
Strategies for engineering leaders in financial services, silicon design, networking, and other mission-critical industries
THE MODERN ERA OF SOFTWARE ENGINEERING
The volume of code created is growing at an unprecedented rate. But we all know there’s a catch. The dramatic reduction of the cost of creating code has simply moved the bottlenecks downstream, where engineers struggle with code comprehension, debugging, and incident investigations.
If the promised productivity gains are to be realized, AI must be applied to the software delivery cycle, not just code generation. But AI has so far proven itself unequal to effective comprehension and debugging of complex, large-scale codebases.
WHERE TIME IS
SPENT NOW
Writing code has become cheaper, but in mission-critical codebases, that code must still be understood. The more demanding task is understanding what that code does, how it affects existing codebases, and debugging it when an application doesn’t behave the way it’s expected to.
0%
of engineering teams’ time is spent on debugging.
Based on the assumption of an
average 40-hour working week.
How engineering teams’ time is divided between tasks

0
hrs/week
Producing code

0
hrs/week
Debugging issues during development
0
hrs/week
Debugging issues found by customers in production
0
hrs/week
All other tasks
Nine in ten (91%) engineering leaders say code comprehension and debugging take longer in complex C/C++ codebases. As AI-generated code increases, the challenge will grow.
The impact of
coding agents on
release cycles
Coding agents promise faster software delivery, but unless AI is applied across the entire software development lifecycle then the bottleneck simply shifts downstream.
Now that AI is generating most of the code being produced, engineers no longer have the inherent understanding they used to. That makes it easier for defects to escape, and when something inevitably goes wrong, nobody has the knowledge to trace the failure back to its root cause.
0%
of AI-generated code has not been fully comprehended before engineering teams push it to production.
Challenges engineering teams face in accelerating software delivery with AI agents
Agents introduce incorrect code too frequently, creating rework that delays delivery cycles
0%
Engineers are required to spend excessive time reviewing agent outputs compared to how long it takes to generate the code
0%
Agents cannot reliably infer what code will do at runtime, and often hallucinate when debugging more complex codebases
0%
Engineers find it difficult to review the large amounts of code generated by agents, and merge too many AI-generated errors into the codebase
0%
It takes too many prompts and re-runs to identify the root cause of test failures with AI agents
0%
Agents lack the ability to diagnose issues that are hard to reproduce such as intermittent failures
0%
Engineering leaders are in an impossible bind: they’re expected to deliver significant productivity gains from AI without compromising quality or accountability.
But as AI-generated code proliferates, review fatigue grows. The risk is that this will lead to more code pushed to production before engineers have comprehended it, more escaped defects, more unresolved tickets, and further gaps in accountability.
Issues organizations have encountered as a result of their use of AI coding tools
Loss of productivity due to
engineers needing to ‘unpick’ AI-generated code
0%
0%
Budget overruns due to agents
needing to re-run tasks multiple times to reach the correct output
0%
0%
Incorrect diagnosis of the
root cause of an issue due to
a hallucination
0%
0%
Test escapes, serious defects,
or poorly optimized code entering production
0%
0%
Production incident / service
outage affecting internal users or customers
0%
0%
per month
past six months
0%
of engineering leaders say despite being able to produce code much faster, release cycles are no faster than before.
The problems created by coding agents are most acute in the complex environments that power mission-critical industries. Trading platforms, network operating systems, and applications of a similar nature run on legacy codebases with poorly documented interdependencies, while critical institutional knowledge lives only in engineers’ heads.
AI inevitably makes subtle mistakes. But so much code is being generated that it is difficult for engineers to review and understand it properly.
It’s no wonder that so few enterprises have seen a meaningful improvement in release cycles.
Ultimately, the productivity gains promised by AI will remain limited unless intelligent engineers can also apply it effectively to comprehension and debugging.
If they fail to do so, the time engineers save in code generation is simply absorbed by slower investigation and debugging, or lost later through production incidents and security failures.
AI model families engineering teams use to understand and debug complex codebases
Different models have their individual strengths and weaknesses, and there are trade-offs between capability and economy for any given task. Engineering teams use a whole host of models for debugging and comprehension, but GPT and Gemini ride high on their list.
But few engineering teams are using the most capable models for code comprehension and debugging, even in more complex environments.
Nearly two-thirds (63%) of engineering leaders say the cost is prohibitive to authorize the use of higher tier AI models for this work.
Rising costs are a sizable barrier to increasing the use of more powerful AI
0%
Of leaders say the cost is prohibitive to authorize the use of higher tier AI models for debugging complex codebases.
0%
Only one in five would use Mythos for these tasks if it was available to them.
0%
More than half of engineering leaders opt for lower cost AI models for code comprehension and debugging.
Problem-solving
in complex codebases
More than a third (37%) of teams use AI agents for comprehension and debugging only for straightforward codebases.
When they do use AI for code comprehension and debugging in more complex codebases, engineering teams rely on a range of workarounds to give a greater degree of confidence.
0%
of engineering leaders say coding agents struggle to solve difficult problems in large-scale, complex codebases.
Techniques organizations use to mitigate the challenges and risks of using AI agents for code comprehension and debugging
0%
Using a human in the loop to approve any AI agent action
0%
Enriching source code analysis and logs with specific context about what happened at runtime
0%
Developing hypotheses and then testing them on agents rather than asking the agent to identify the cause / solution
0%
Using agents to check the work of other agents in autonomous workflows
Undo enables coding agents to solve the most complex problems on the most complex codebases.
Fully automated root-cause analysis
Undo creates deterministic, self-contained, portable recordings of complete program executions, giving agents the runtime context to reason about dynamic behavior like static code.
With Undo, developers have been able to solve problems up to 100x faster in some of the world’s most demanding software environments including at AMD, AWS, Cisco, and Palo Alto Networks.

Find out how Undo could help your engineering team unlock the potential of AI agents.
EXPLORE THE FULL RESEARCH
Download the full report to explore the findings in greater detail.