DVCon Papers - Curated Collection

DVCon is where verification practice is reported by the people doing it, which makes its proceedings the best record of what actually changed on real projects, year by year. This collection is my reading list from those proceedings, going back to 2011, with a summary written for each paper that says what problem it solved and what a verification engineer can take from it today. It is not a mirror of the proceedings: papers are here because I read them and found something worth passing on.

The themes, in rough order of how much of the list they occupy. Verification of AI-assisted flows and AI applied to verification dominate 2025 and 2026, from agentic regression frameworks to LLM-generated assertions, and my separate analysis of the 2025 to 2026 AI papers reads across them. Open-source tooling, especially Verilator with UVM and SystemC, is the second thread and the one most useful to teams without a big EDA budget. The older papers hold up for methodology: UVM structure, sequences and phasing, coverage closure, formal in the mainstream flow, and the occasional paper on debug that is as relevant now as when it was written.

Editorial note: summaries are my own words and my own reading; where a summary says a paper claims something, the number is the authors', and the interpretation is mine.

53
Papers
15
Years
20
Companies

How to use: Browse by year, or use Ctrl+F to search for specific topics like "UVM", "Formal", "Coverage", etc.

DVCon 2026(15 papers)

Title / Company Summary Link
A 3-tiered agentic AI Framework for Verification Regression
Samsung • Jin Choi, Sangwoo Noh, Seonghee Yim, Youngsik Kim
Flat, monolithic regression pipelines rerun huge test sets with no awareness of design hierarchy or which block a change touched, so chip-level runs start on unstable blocks, engineer sanity runs go untracked, and tests silently never get executed. The authors split regression into three scopes (an engineer's own sanity runs, change-triggered block runs, and gated chip-level runs) and connect them through RabbitMQ publish-subscribe bridges; a gate agent only promotes a block to chip regression once its pass rate clears a threshold. Agents are typed as regression (decide what to run), bridge (route results between tiers) or operation (stateless workers for dispatch, failure classification and owner routing via filesystem, REST and JIRA MCP servers); no specific LLM is named and much of the logic is policy-driven.
Beyond Heuristics: AI/ML Driven Verification for Design Sign-off
Samsung SSIR • Gulshan Kumar Sharma, Sougata Bhattacharjee, Abhishek Raj, Avinash Bollu, Akshaya Kumar Jain
Static ranking of regression tests by hand-tuned weights (failure recency, coverage gain, runtime) goes stale as RTL and coverage goals change, wasting simulation cycles and delaying coverage closure. A supervised classifier (Random Forest chosen over logistic regression and gradient boosting, trained with scikit-learn) scores each test and seed on five features such as coverage novelty, failure likelihood, coverage-per-second and seed-outcome entropy, then feeds a prioritized list to the unchanged UVM regression scheduler. Retraining happens automatically every 2 to 3 regressions, low-utility tests are kept for occasional exploration, and Cadence SimAI is used for log extraction and generating the optimized test-set files.
ChipAgents in Practice: Lessons from One Year of Agentic AI EDA Deployment
ChipAgents • Mehir Arora, Luke Frey
Design teams face growing complexity and thick layers of legacy IP, and having engineers hand-craft prompts for each agent task does not scale. The tutorial argues for inverting the loop so the agent asks the human for what it needs, for agents that understand the project's workflow and workspace rather than only individual tools, for externalized knowledge banks that carry lessons between projects, for one general-purpose agent instead of many task-specific ones, and for deliberate human checkpoints such as plan review and interruptibility along a gradual autonomy spectrum.
DUET: Agentic Design Understanding via Experimentation and Testing
ChipStack (Cadence) • Gus Henry Smith, Sandesh Adhikary, Vineet Thumuluri, Karthik Suresh, Vivek Pandit, Kartik Hegde, Hamid Shojaei, Chandra Bhagavatula
LLM agents that only read SystemVerilog miss the time-dependent behavior hidden in RTL, so they produce shallow design descriptions and then fail at downstream jobs such as writing formal properties. DUET wraps an experimentation loop around the agent: the LLM proposes a hypothesis about the design, checks it by invoking EDA tools (Verilator simulation, Jasper formal runs, waveform dumps, Yosys analysis, and a sub-agent that turns a Jasper counterexample into a simulation testbench), then folds the observed behavior back into its understanding before retrying. The evaluated flow ran GPT-5.0 inside an agentic formal-verification pipeline and told the agent to experiment whenever a formal testbench failed.
AI Agent-based Error Resolution System for SoC RTL Verification
Samsung • Yonghyun Kwon, Hanna Jang, Seonghee Yim
After an overnight regression an engineer burns 30 to 60 minutes per failure just collecting logs, config files, spec data, and similar past cases before any real analysis begins. A LangGraph pipeline of five agents (Error Analyzer, Data Collector, Decision Maker, Auto Executor, Notification) running on GPT-OSS-120B, with Qwen3 embeddings retrieving one of 86 curated error-to-SOP patterns and MCP servers giving read access to MongoDB history, an Oracle spec database, and the file system. Only low-risk actions with backups are auto-executed; the human still makes the fix decision, and over 2,500 lines of hand-written domain prompts carry the verification know-how.
Etabot: Multi-Agent Verification Management with Chat-Accessible Midpoint Metrics and Finish-Date Forecasts
Silicon Labs • Zili Fang
Answering 'where are we and when will we finish' means hand-collecting logs, merging coverage, separating infrastructure noise from real design failures and writing up status, which eats engineering time every week. Etabot is a small on-prem Python framework: a planner turns a one-line intent into a strict JSON task list, a sequential orchestrator looks up replaceable agents in a registry (regression runner, coverage merge and parse, Jira bug stats via JQL, a Monte Carlo forecaster, a markdown reporter), and results are exposed to a CLI and an MCP server so a chatbot can query or trigger the flow. The forecaster fits recent coverage slope, detects resets from new coverpoints, and reports P50 and P80 finish dates; LLM planning is optional and rule-based by default, and the author notes the code itself was LLM-generated.
A Novel Fast Regression: An AI/ML Driven Automated Sanity Regression Flow
Samsung / Cadence • Joonho Chung, Daeseo Cha, Woojoo Space Kim, Youngsik Kim, Seonil Brian Choi, Nili Segal
After each RTL drop, sanity regression still depends on engineers hand-picking tests, reading hundreds of failure logs and hunting through waveforms, costing 4 to 8 engineer-hours per iteration and giving inconsistent results. Smart Sanity Regression (SSR) is a one-command orchestration layer over Cadence Verisium apps: CodeMiner plus AutoFocus pick tests correlated with the RTL diff against a golden baseline session, Manager plus AutoTriage cluster failures with unsupervised-then-supervised ML and rerun one representative per cluster with waveforms, and WaveMiner diffs those against golden waveforms to rank suspect signals. The contribution is the staged data handoff and rerun policy, not new algorithms.
FVDebug: An LLM-Driven Debugging Assistant for Automated Root Cause Analysis of Formal Verification Failures
NVIDIA • Yunsheng Bai, Ghaith Bany Hamad, Chia-Tung Ho, Syed Suhaib, Haoxing Ren
Working out why a formal property failed means hand-tracing a multi-cycle counterexample against the RTL and the spec, which can eat hours or days per bug. FVDebug turns a Jasper counterexample into a causal DAG by recursively calling the tool's visualize -why command, then runs an o3-mini pipeline over it: a Graph Scanner that scores every node with a forced pro-and-con argument prompt, an Insight Rover that agentically walks the graph to build and rank competing failure hypotheses, and a Fix Generator that produces RTL patches via five prompting strategies with consensus scoring. Output is a report with ranked hypotheses, a cycle timeline, and diff-ready fixes.
A novel ML-Driven Simulation Log Debugger
Samsung • Narasimha Rao Chinni, Sunil Shrirangrao Kashide, Garima Srivastava
SoC simulation logs run to a million lines, and the same test on two design labels prints messages differently and in a different order, so diffing a failing log against a passing one by hand is slow and unreliable. LogMiner is a Plotly Dash dashboard that buckets log lines into tabs (UVM, errors, tool, VIP, register writes, boot) and adds a semantic comparator: the all-MiniLM-L6-v2 sentence-transformer is fine-tuned with contrastive loss on Samsung log lines, and cosine similarity against a user-built reference log surfaces matching, missing, and out-of-order lines. The user picks a similarity threshold from a precision/accuracy sweep.
Multi-Agent Orchestration for Autonomous Regression Management
Samsung • Sangwoo Noh, Jin Choi, Seonghee Yim, Youngsik Kim
Regression automation can launch and watch jobs, but a human manager still has to read failures, decide who owns them, spot duplicates and replan, which makes triage the throughput bottleneck. They frame regression as a closed control loop and propose an Orchestrator coordinating six role-specific LLM agents (Planning, Execution, Monitoring, Resource, Debugging with root-cause and duplicate sub-agents, Notification) that pass schema-checked messages and must cite evidence or abstain. Only the triage slice (Ownership plus Debugging agents) is actually benchmarked, against a scripted CI heuristic and a single monolithic agent, using a deterministic mock tool registry (grep, git blame, database query) and fixed per-case budgets; the underlying LLM is not named.
Towards Self-Adaptive SoC Design Verification: KG-Enhanced Generative AI, RL and Backpropagation Debugging
Samsung • Insu Jang, Seonghee Yim, Hanna Jang, Youngsik Kim
When an IP changes, a static block-IP-test hierarchy is a poor guide to which tests will actually flip, because in practice many result changes show up in IPs that were never edited, so impact-based regression selection wastes runs and misses real effects. A Neo4j property graph holds blocks, IPs, features, testcases, and design elements with static edges from design and test-plan data; a PPO reinforcement-learning agent replays historical design versions, picks tests it expects to change, earns reward when its picks flip (+10) and pays when they do not (-10) or when unpicked tests flip (-5), and writes the learned influence back into weighted LEARNED_CORRELATION edges. Failures are then traced backward along those weighted edges to rank likely root-cause IPs.
SIGMA: Sign-off Intelligence with GenAI for Methodical Assurance in Formal Verification
Cadence • R C Sanjay Krushnan, Moola Jeevan Chaitanya Goud, Sakthivel Ramaiah, Erik Seligman
Three front-end formal tasks, pulling a verification plan out of the spec, standing up the formal environment, and writing routine SVAs, are slow, expert-dependent, and where checks get missed. SIGMA chains an in-house GenAI platform that reads architecture and micro-architecture specs and emits a classified verification plan (Gemini, GPT-OSS-120B and a Llama Nemotron variant were compared), a Python script that parses the DUT ports and generates the top, checker, file list and Makefile, and a Copilot-based assistant that drafts SVAs for basic checks with RAG over previously corrected assertions for harder ones. Every stage keeps an engineer review step before sign-off.
AI Driven Advanced Debugging in SoC Design Verification
Samsung • Shreya Jayateerth Joshi, Poonam M Shettar, Sanjoy Saha, Alok Kumar, S Shrinidhi Rao, Garima Srivastava
SoC testbench bring-up and post-simulation debug lean on manual log reading, waveform staring, and rerunning simulations, and ignore the reusable data piled up from earlier projects. A hybrid pipeline: before simulation, register init sequences are checked against a statistical baseline from past projects using rule checks, RapidFuzz and BERT-embedding name matching, and an Isolation Forest anomaly detector; after simulation, Python scripts use UVM checker messages to pull AHB/APB/AXI signals from the wave dump and feed them to a trained classifier (six architectures compared, Transformer encoder wins) that flags protocol violations, which rules then map to a testbench line or design update. A FAISS-indexed history of past errors plus a GPT-OSS-120B chat interface let engineers query the accumulated debug log.
Unified AI-Driven Verification: Combining Spec-RAG, Memory Networks, and Generative AI
Samsung • Seonghee Yim, Insoo Jang, Hanna Jang, Youngsik Kim
On a large SoC the spec, RTL, test plans, regression logs, and Jira tickets live in different systems and change at different rates, so tracing a failure back to its requirement or to a prior fix depends on tribal knowledge and slow manual digging. The team built a debug assistant with two retrieval memories: a short-term store over MongoDB regression logs and recent diffs, and a long-term store over Oracle tables plus a HippoRAG/Neo4j knowledge graph linking spec sections, block versions, test plans, results, and issues. A small decision model (MiniMax-M2) classifies which maturity stage (ML1 to ML4) a query belongs to and sets the STM/LTM retrieval mix, then an action model (GLM-4.6) reasons over the gathered evidence to write a root-cause explanation.
Passive Token Accounting For Cost-Aware LLM Assistance In Continuous SoC Verification Workflows
Dhenara Inc. • Ajith Jose
Once LLM triage and debug agents run inside nightly regressions, nobody in the EDA stack can say what a given test, agent, or heuristic change cost in tokens, so spend can balloon unnoticed or compete for on-prem GPU time. Passive Token Accounting reads the usage fields already present in provider responses through the open-source dhenara-ai client, converts them to USD with a per-model rate table, rolls them up per node and per agent inside the DAG-based agent runtime, and appends everything to a per-regression costs.json that a stateless gateway serves to a UI. Budget knobs let a deep-debug agent stop sub-loops at a cost or token ceiling and swap to a cheaper model when a soft cap is reached.

DVCon 2025(6 papers)

Title / Company Summary Link
AI Pair or Despair Programming: Using Aider to build a VIP with UVM-SV and PyUVM
Verilab • André Winkelmann, Damir Ahmetovic Ignjic
Nobody had measured, side by side, how well a terminal coding assistant can build a real verification IP in SystemVerilog UVM versus Python cocotb plus PyUVM. The authors scripted Aider with three frontier models (Claude Opus 4, Gemini 2.5 Pro, OpenAI o3) through an identical prompt sequence to generate an APB3 VIP and self-checking testbench in both ecosystems, with and without a coding-conventions file, then counted fix iterations, inspected waveforms, and logged API cost.
GAIL-V: Generative AI Leveraged Verification
Arm • Sriram Madavswamy, Daniel Larkin
Teams are adopting LLMs for verification one engineer and one script at a time, with no shared way to decide which parts of the flow actually suit a model or to measure how well it does there. The authors split the whole DV flow into a labelled tree of goals and leaf sub-goals (plan, implement, analyse, review), then for each leaf ask whether a model fits, sketch an ideal LLM-driven workflow, build a labelled dataset and scoring prompt, and regress many models and prompt styles against it. The worked example is a milestone sign-off audit: an LLM synthesises fake verification plans with deliberately seeded defects, a second LLM acts as auditor, and a third scoring prompt compares its verdicts to the seeded labels.
LLM-based Functional Coverage Generation and Auto-Evaluation Framework
Masaryk University • Ján Labuda, Marcela Zachariášová, Zdeněk Matěj
Functional coverage is a low-risk place to try LLM code generation, but there was no reproducible way to score what small, locally run models actually produce. The team defined a Python functional-coverage API for cocotb, fed open-weight models (DeepSeek-R1, Gemma 3, Qwen3, up to 14B parameters) natural-language verification requirements rather than the raw spec, and auto-compared the generated covergroups against an expert-written reference across repeated attempts.
Accelerating Coverage Closure with Reinforcement Learning: A Case Study on FSM Verification
Verification Consulting doo • Tijana Misic
Most seeds in a constrained-random regression add nothing to coverage yet still burn simulator licences and wall-clock time, and the sanity regressions gating every commit pay that cost repeatedly. A Python agent runs after a normal UVM regression, reads each test's coverage report, and assigns a reward when a run reaches a new FSM state and a penalty when it does not, keyed on the random selector value that chose that run's constraint set. The selector values with the best reward history are then replayed in a shorter follow-up regression, so the testbench itself is left untouched.
VerifLLMBench: An Open-Source Benchmark for Testbenches Generated with Large Language Models
University of Minnesota / Synopsys • Nishanth Somashekara Murthy, Eldon Nelson, Sachin S. Sapatnekar, John Sartori
There are public benchmarks for LLM-written RTL but none for LLM-written UVM environments, so there is no repeatable way to compare how good a model's testbench actually is. The authors take five already-verified DUTs from the RTLLM suite, hand the model a fixed interface, top module and a structured prompt listing every UVM class it must produce, then compile in VCS with up to four fix-the-syntax retries, simulate for coverage, and lint the generated classes with Euclide. They ran the loop five times each for ChatGPT-4o, Gemini 1.5 Pro and Llama 3.2, releasing the harness under an MIT licence.
Towards Automated Verification IP Instantiation via LLMs
NVIDIA • Ghaith Bany Hamad, Michael Marcotte, Syed Suhaib
Formal teams own libraries of reusable assertion VIPs for common structures like FIFOs and arbiters, but hooking them up to every instance in a large RTL codebase is slow, needs VIP expertise, and gets skipped under schedule pressure. A four-stage pipeline strips comments and preprocessor noise from the RTL, splits it into syntactically complete chunks under the model's token limit, filters chunks by FIFO and arbiter naming conventions, and asks an LLM to emit the VIP binding for each. The prompt is augmented either with a one-shot example (in-context learning on an internally fine-tuned Mixtral 8x7B) or with retrieved snippets from a vector store of VIP docs and prior instantiations (RAG on stock Llama-3-70B-instruct); a post-pass discards outputs with two or more unresolved signals as hallucinations and back-fills FIFO depth and width from source.

DVCon 2023(4 papers)

Title / Company Summary Link
Verilator + UVM-SystemC: A Match Made in Heaven
DVCon Europe • Luca Sasselli
Using Verilator with Accellera's UVM-SystemC library to build open-source verification environments. Case study verifying a RISC-V microprocessor.
Code-Test-Verify All for Free: Assertions + Verilator
DVCon India
Verification methodology for AMBA APB/AHB protocol checkers using SVA with open-source Verilator. Covers SVUnit test framework.
Demystifying Formal Testbenches: Tips, Tricks, and Recommendations
Practical guidance on formal verification testbenches. Covers input state space exploration, constraint assumptions, and assertion modeling.
Formal Verification Framework for Hardware Accelerator Designs
Novel framework for verifying hardware accelerators using formal verification techniques, addressing unique challenges of specialized computing engines.

DVCon 2022(2 papers)

Title / Company Summary Link
Challenges of Formal Verification on Deep Learning Hardware Accelerator
Formal verification challenges in verifying deep learning hardware accelerators. Achieving maximum formal coverage alongside constrained random verification.
Novel Approach for SoC Pipeline Latency and Connectivity Verification Using Formal
Formal verification of SoC pipeline connections for AXI, ACElite, and CHI protocols. Addresses latency and connectivity verification challenges.

DVCon 2021(1 papers)

Title / Company Summary Link
Portable Stimulus vs Formal vs UVM: A Comparative Analysis
Compares verification of AHB to APB Gasket IP using UVM, Portable Stimulus, and Formal across block-level DV, System DV, and board validation.

DVCon 2020(1 papers)

Title / Company Summary Link
System-Level Register Verification and Debug
Discusses 2020 UVM standard. EUVM testbenches compile to native binaries for embedded systems, enabling systems perspective to functional verification.

DVCon 2019(1 papers)

Title / Company Summary Link
Advancing System-Level Verification Using UVM in SystemC
Introduction to UVM-SystemC verification methodology and class library. Demonstrates resemblance with UVM standard for system-level verification.

DVCon 2018(2 papers)

Title / Company Summary Link
The Top Most Common SystemVerilog Constrained Random Gotchas
Illustrates the most common SystemVerilog constrained random gotchas in UVM testbenches. Helps avoid CR debugging pitfalls.
UVM Random Stability: Don't Leave It to Chance
Avidan Efody
Decoupling random parts for individual seeding. Changing specific parts keeps everything else constant for better debug reproducibility.

DVCon 2017(2 papers)

Title / Company Summary Link
Coverage Models for Formal Verification
Xiushan Feng, Xiaolin Chen, Abhishek Muchandikar
Coverage models specifically designed for formal verification flows. Bridges gap between simulation and formal coverage metrics.
On Verification Coverage Metrics in Formal Verification
Speeding verification closure with UCIS coverage interoperability standard. Unified coverage metrics across formal and simulation.

DVCon 2016(4 papers)

Title / Company Summary Link
UVM-Light: A Subset of UVM for Rapid Adoption
Identifies UVM subset (base classes, methods, macros) for faster learning. Enables engineers to become productive quickly with UVM.
Easier SystemVerilog with UVM: Taming the Beast
Doulos • John Aynsley
Practical approaches to simplify SystemVerilog and UVM adoption. Taming complexity for verification engineers.
Advanced UVM Register Modeling: There's More Than One Way to Skin A Reg
UVM register model in active and passive modes. Register operations via read/write methods converted to sequence items through adapters.
C Through UVM: Effectively Using C-Based Models with UVM VIP
Integrating C-based reference models with UVM-based verification IP. Bridges software and hardware verification methodologies.

DVCon 2015(3 papers)

Title / Company Summary Link
Design Guidelines for Formal Verification
Juniper Networks • Anamaya Sullerey
Design guidelines that facilitate application of formal verification on large blocks. Practical recommendations for formal-friendly RTL.
Automated Performance Verification to Maximize Your ARMv8 Pulling Power
Testbench automation for AMBA VIP configuration. Interconnect Validator for coherency and data consistency checks across Video, GPU, PCIe, DMA.
ACE'ing the Verification of a Coherent System Using UVM
Synopsys • Romondy Luo
AXI ACE VIP with System Environment and configurable ACE Interconnect components. Pure SystemVerilog architecture with UVM methodology.

DVCon 2014(3 papers)

Title / Company Summary Link
A Novel Processor Verification Methodology Based on UVM
UVM-based methodology for processor verification. Addresses unique challenges of CPU/processor design verification.
Highly Configurable UVM Environment for Parameterized IP Verification
Predictor takes real transaction packets from AMBA UVC TLM ports and outputs predicted transactions to scoreboard through TLM ports.
Smart Formal for Scalable Verification
Techniques for scaling formal verification to larger designs. Smart strategies for managing state space explosion.

DVCon 2013(4 papers)

Title / Company Summary Link
AMS Verification in a UVM Environment
Analog/Mixed-Signal verification within UVM testbenches. Bridging digital and analog verification methodologies.
A UVM SystemVerilog Testbench for Analog/Mixed-Signal Verification
Seamless integration with UVM capabilities: constrained randomization, functional coverage, TLM, and concurrent assertion of analog properties.
Off To The Races With Your Accelerated SystemVerilog Testbench
Hardware-assisted acceleration of SystemVerilog testbenches. Transaction-level testbenches for both simulation and acceleration.
Formal Verification in the Real World
Jonathan Bromley
Practical formal verification techniques. Achieving 100% toggle coverage and tracking achieved bound as function of tool runtime.

DVCon 2012(4 papers)

Title / Company Summary Link
Fabric Verification
Galen Blake
AMBA Fabric verification with PCIe, USB, ENET masters/slaves. OVM-native VIP for AMBA3 protocols (AXI, AHB, APB) with coverage collection.
Advanced Techniques for AXI Fabric Verification
AXI fabric verification in software-hardware OVM environment. APB for peripheral register access and slow peripheral data transport.
The Missing Link: The Testbench to DUT Connection
David Rich
Methodologies for connecting testbench to DUT. Most common approach using SystemVerilog's virtual interface.
SystemVerilog Checkers: Key Building Blocks for Verification IP
SystemVerilog checkers as reusable verification IP building blocks. Protocol checking and assertion-based verification.

DVCon 2011(1 papers)

Title / Company Summary Link
Discovering Deadlocks in a Memory Controller IP
Formal verification methodology to discover deadlocks in RTL. Memory controller between ARM AMBA AXI interface and memory device.

Last updated: 2026-09-06

Missing a paper? Leave a comment below with the details!

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment