Skip to main content
Venn diagram: bug difficulty (deep traces across hierarchies, pipelines and large waveforms) overlaps with bug quantity (hundreds of failing sims across tests, seeds and root causes); the overlap is hard bugs at scale

Scaling Autonomous Root Cause Analysis for Regression Triage

Zackary Glazewski avatar
Zackary GlazewskiOctober 6, 2026

As semiconductor designs grow larger and more complex, debugging is becoming increasingly difficult to scale.

At DVCon earlier this year, NVIDIA’s Stuart Oberman described first-silicon success as approaching a “statistical impossibility.” Semiconductor Engineering cites that debugging accounts for 44% of verification effort; a number reported in 2018, before the scale and complexity of today’s designs.

As designs get larger and design spaces are expanding, engineers are working across systems containing millions of gates. If design complexity continues to scale, debugging needs a fundamentally different approach. In this article, we will explain an AI-native approach to root cause analysis (RCA) and regression triage.

Debugging Complexity scales along two axes

There are two important ways that debugging effort scales.

The first is bug difficulty.

Some bugs require deep tracing across design hierarchies, long pipelines, large waveforms, and multiple interacting components before the actual root cause becomes clear.

ChipAgents RCA was designed specifically for this axis. Rather than treating debugging as a simple question-answering problem, the system approaches RCA as a search problem. It explores high-probability traces, executes across different hypotheses, ranks the resulting explanations, and provides confidence scores to the engineer.

Consider one bug encountered in a real customer environment.

The design contained a bidirectional serialization interface between a manager and subordinate. During a transaction, a reset was triggered by the subordinate while the manager continued transmitting a frame beyond the reset grace period.

Part of the transmitted data happened to contain the same value used to represent the protocol’s start code. The subordinate interpreted that payload as the beginning of a new frame, eventually causing a downstream failure.

Several conditions had to coincide: a reset needed to occur during the transaction, the data needed to resemble the start code, and the transmission needed to outlive the reset grace period.

This is the type of bug where autonomous RCA becomes particularly valuable because identifying the root cause requires reasoning through a long chain of events.

But difficulty is only one dimension.

The second is bug quantity.

A regression may contain hundreds of simulations across many test cases and seeds. Multiple simulations can fail simultaneously, and those failures may represent different bugs—or many apparently different failures may actually originate from the same root cause.

The most challenging case lies at the intersection of these two axes:

hard bugs at scale.

Venn diagram: bug difficulty (deep traces across hierarchies, pipelines and large waveforms) overlaps with bug quantity (hundreds of failing sims across tests, seeds and root causes); the overlap is hard bugs at scale

Now the problem is no longer simply finding one root cause. The system must determine how many distinct problems actually exist, decide which failures are related, identify which simulations contain the information needed to debug them, and orchestrate RCA across those problems efficiently.

Why Regression Triage Is an Orchestration Problem

Regression debugging creates a fundamentally different scaling challenge.

There can be an arbitrary number of bugs with varying levels of difficulty. Fixing one problem may expose another. Hundreds of simulations may be running across different test cases and seeds.

And every simulation generates data.

Individual debug cases can already involve enormous logs and waveforms. At regression scale, that data multiplies across every failing simulation.

The naive approach would be to apply an agent to every failure, rerun simulations whenever more information is needed, generate waveforms for every failing test, attempt fixes, and then rerun the entire regression.

But this quickly becomes inefficient.

For autonomous regression triage to work at commercial scale, the system needs to minimize the amount of expensive work it performs.

That means being selective about which failures actually need independent investigation and which simulations actually need to be rerun.

It also means that the underlying RCA system needs to be highly reliable. If an autonomous system proposes an incorrect fix, reruns a simulation, fails again, tries another hypothesis, and repeats the process across dozens of failures, the compute cost can grow quickly.

Accuracy and efficiency therefore become tightly coupled.

Intelligent Binning: Finding the Bugs Behind the Failures

The first step is understanding how many distinct bugs actually exist.

Traditional regression systems already use binning techniques, often relying on deterministic scripts and pattern matching to group failures with similar signatures. These methods remain useful.

But there is an opportunity to go further.

Two simulations may produce different log signatures while still sharing the same underlying root cause. An intelligent system should therefore be able to collapse failures beyond what deterministic pattern matching alone can identify.

ChipAgents RCA Triage uses a secondary pass after conventional binning to identify failures that are most likely related by a common root cause.

Instead of treating every failing simulation as an independent debugging task, the system attempts to reduce the regression into a smaller set of meaningful failure groups.

That creates a much more manageable problem for the RCA system.

Rerun the Simulations That Matter

The same principle applies to waveform generation.

One straightforward approach would be to rerun every failing simulation and dump a waveform for each one.

But that is expensive and often unnecessary.

Instead, the system can identify a minimum set of witness simulations capable of explaining the failures represented by each bin.

Those witness simulations are then rerun with waveform dumping enabled.

The goal is not to collect every possible piece of debugging information. It is to collect the smallest amount of information necessary to explain the regression failures.

This is an important distinction because regression triage is a long-horizon task. A regression might take minutes or hours to execute. Every unnecessary simulation adds cost and delays the final result.

From Automated Triage to Autonomous Resolution

Binning and selective waveform reruns already represent forms of automation.

But automation does not have to stop there.

Once the regression has been reduced into meaningful failure groups and the necessary witness simulations have been identified, those debug packages can be handed directly into autonomous RCA.

The RCA system investigates the root cause for each bin, generates potential fixes, and performs verification.

If verification succeeds, the system gains another signal that the proposed fix is correct. Combined with RCA confidence scoring and the result of the witness simulation, the system can determine whether the fix is likely ready for a final regression run.

The final step is to rerun the regression and verify that the failures have actually been resolved.

This creates an end-to-end workflow:

Regression → intelligent binning → witness selection → waveform rerun → autonomous RCA → fix → verification → final regression.

Instead of an engineer waking up after an overnight regression, discovering a collection of failures, and beginning the triage process manually, the regression can become the beginning of an autonomous debugging workflow.

By the time the engineer returns, the system could have resolved some or all of those failures. If some remain unresolved, the engineer can receive a report showing what was fixed and what still requires investigation.

The paradigm shifts from starting regression triage to reviewing the results of regression triage.

Multi-Agent Systems Need to Scale Economically

There is another challenge that becomes increasingly important as these systems scale: cost.

Multi-agent systems naturally consume more resources than a single agent because they execute the agent primitive multiple times. That tradeoff exists for a reason: parallel exploration can improve performance and accuracy.

But simply spending more compute is not a sustainable path to better engineering results.

The objective should be to gain the accuracy advantages of multi-agent reasoning while improving token and cost efficiency.

In our internal benchmarks, ChipAgents' multi-agent systems achieve state-of-the-art performance while operating at costs comparable to traditional single-agent systems.

That is an important direction for autonomous semiconductor engineering.

As agents take on longer-horizon workflows, the question will increasingly move beyond whether an agent can perform a task.

The question becomes whether the system can perform that task reliably, efficiently, and economically across the scale of real semiconductor development.

Regression triage is a useful example of that transition.

Solving one difficult bug autonomously is valuable.

Solving difficult bugs across an entire regression while intelligently deciding what to investigate, what to rerun, and where to spend compute is the next step toward autonomous debugging at commercial scale.