How AMD reduces schedule risk by cutting debug time from days to hours

How AMD reduces schedule risk by cutting debug time from days to hours

Image link

For AMD‘s globally distributed Graphics IP team, the bottleneck wasn’t talent or tooling — it was the time spent identifying the root causes of unexpected behavior during system verification. Here’s how AMD reduced schedule risk by accelerating root cause analysis.

The challenge: when debug time becomes schedule risk

In silicon design, verification is where schedules are made or broken. It accounts for close to 70% of project time — and within it, debug is the single biggest engineering drain. For AMD’s Graphics IP team, that pressure was compounded by the realities of working across a globally distributed Modeling Team.

C Simulation models are complex, and when something goes wrong, finding out why means regressions can easily stretch across days. Complex issues often required triage across multiple conference calls, teams, and time zones. Engineers rotating onto unfamiliar blocks were spending weeks building context before they could meaningfully contribute. Every one of those delays was a direct threat to delivery.

So AMD’s Modeling Team was looking for a way to optimize their verification workflow.

The solution: bringing runtime context to root cause analysis

The hardest part of debugging complex C Simulation regressions isn’t finding the symptom; it’s reconstructing what happened. Engineers often spend days reproducing issues, gathering logs, and building enough context to understand the root cause.

Undo solves that problem by capturing complete program execution — including state changes, execution paths, function calls, and data flows — into a recording that can be analyzed by both engineers and AI agents.

These recordings provide the runtime context needed to understand what actually happened during execution. Rather than starting with a symptom and working backward through fragmented evidence, engineers and AI systems can investigate failures using a complete record of the program’s behavior.

For AMD’s Graphics IP team, that meant less time spent gathering context and more time spent resolving issues.

The approach: prove it before you commit

Nalini Patel, Senior Manager of Design Engineering at AMD, had a clear principle when it came to new tooling: prove the value in practice before asking anyone to commit to it. So when she brought Undo into AMD, she introduced Undo through its Land and Prove Program — full, unlimited access for AMD’s engineering teams, with no large upfront commitment required.

Undo reinforced the program with something engineers found equally valuable: hands-on, in-person training from Undo’s own engineers. Not documentation. Not webinars. Real-time support from experts who could answer questions on the spot and help AMD’s teams realize value faster.

The results: from days to hours

The impact of Undo on AMD’s Graphics IP workflows was immediate and measurable across three areas that matter most to a distributed engineering team:

Regressions resolved in a morning, not two days. CSim regressions that had routinely consumed two days of engineering time are now resolved the same morning. That’s not an incremental improvement — it’s a fundamental change in how teams plan and pace their work.

Complex triage compressed from days to around two hours. Issues that previously required extended investigation across multiple stakeholders can now be investigated and resolved in a fraction of the time. On a fixed tape-out schedule, that delta is the difference between hitting quality targets and slipping.

Handoffs that don’t require a meeting. An engineer encountering an issue can record it with Undo, set bookmarks in the recording, and pass the recording to a colleague across time zones — no conference call, no lengthy written explanation, no loss of context across a language barrier. The recording carries the context that previously required a conversation.

Engineers can contribute to unfamiliar blocks without weeks of ramp-up. AMD’s teams are dynamic — engineers rotate between blocks and don’t always carry deep institutional knowledge of the code they’re debugging. Undo lets them go directly to the source of a discrepancy, rather than spending weeks building the mental model needed to get there.

Schedule is king. Thanks to Undo, we can hit the quality targets we need to hit on schedule.

Nalini Patel, Senior Manager of Design Engineering at AMD

Takeaways for engineering leaders

AMD’s experience with Undo points to a few principles worth internalizing:

  1. Debug time is schedule time. Anything that compresses the distance between “issue found” and “root cause identified” has a direct line to delivery dates.
  2. Distributed teams need async debugging tools. If your debug workflow depends on synchronous communication across timezones, you’re paying a tax that compounds at scale.
  3. Prove it before you commit. The Land and Prove model lets AMD de-risk the Undo investment and build internal confidence before scaling.

 

Curious whether Undo could help your team? Get in touch, we’d love to talk.

GET IN TOUCH

AMD partnered with Undo through the Land and Prove Program. Undo’s AI-powered root cause analysis platform is deployed across AMD’s C Simulation verification workflows in the Graphics IP division.

Stay informed. Get the latest in your inbox.