Complete SystemVerilog Testbench for a Counter: Interface, Monitors, Reference Model, Scoreboard and SVA

A UVM testbench hides a lot of machinery behind base classes. Before you use that machinery it pays to build the same structure once by hand, in plain SystemVerilog, so that every mailbox, every clocking block and every handshake is something you wrote and can debug. This post does exactly that for a 4-bit loadable up/down counter: a layered, self-checking testbench with a generator, a driver, two monitors, a reference model, a scoreboard with functional coverage, and a bound SVA module. The original January 2026 version of this post was mostly code. This rewrite keeps every file, adds the two that were missing (the write monitor and the test class), explains the design decisions between the listings, and reports what the whole bench looks like under Verilator 5.

Key takeaway The structure here is the structure of every UVM agent and environment, minus the base classes. Generator maps to sequence, write BFM to driver, the two monitors to a monitor with an analysis port, mailboxes to TLM FIFOs, and the environment's build and run tasks to build_phase and run_phase. Learn it here where nothing is hidden, and UVM stops looking like magic.

The data flow

flowchart TB
    subgraph TB["Testbench"]
        GEN[Generator] -->|trans| WR_BFM[Write BFM]
        WR_BFM -->|drive| IF[Interface]
        IF -->|monitor| WR_MON[Write Monitor]
        IF -->|monitor| RD_MON[Read Monitor]
        WR_MON -->|trans| RM[Reference Model]
        RM -->|expected| SB[Scoreboard]
        RD_MON -->|actual| SB
    end
    
    IF <-->|signals| DUT[Counter DUT]
    DUT -.->|bind| ASSERT[SVA Assertions]
    
    style GEN fill:#d1fae5,stroke:#10b981
    style SB fill:#dbeafe,stroke:#3b82f6
    style DUT fill:#fef3c7,stroke:#f59e0b
    style ASSERT fill:#fee2e2,stroke:#ef4444

Stimulus flows left to right. The generator randomizes transactions and mails them to the write BFM, which drives the interface. The write monitor watches the same interface pins and mails what it saw to the reference model, not what the generator intended. That distinction matters: the reference model predicts from observed inputs, so a driver bug shows up as a mismatch instead of being hidden by a shared assumption. The read monitor captures the DUT output, and the scoreboard compares prediction against observation one transaction at a time. Assertions run alongside, bound to the DUT, and check the protocol cycle by cycle.

1. Design under test

The counter is deliberately small so the testbench is the subject. Synchronous reset, a load input that overrides counting, and an updown control.


module counter(clk, rst, data, updown, load, data_out);

  input clk, rst, load;
  input updown;
  input [3:0] data;
  output reg [3:0] data_out;

  always @(posedge clk) begin
    if (rst)
      data_out <= 4'b0;
    else if (load)
      data_out <= data;
    else
      data_out <= (updown) ? (data_out + 1'b1) : (data_out - 1'b1);
  end

endmodule

2. Interface with clocking blocks

The interface declares the pins once and then defines three clocking blocks, one for each component that touches the pins, with directions that match the component's role. Modports hand each component only its own clocking block.


interface counter_if(input logic clk);

  logic rst, updown, load;
  logic [3:0] data;
  logic [3:0] data_out;

  // Write BFM clocking block (driver)
  clocking wr_cb @(posedge clk);
    output load, updown, rst;
    output data;
  endclocking

  // Write Monitor clocking block
  clocking wrmon_cb @(posedge clk);
    input data;
    input load, rst, updown;
  endclocking

  // Read Monitor clocking block
  clocking rdmon_cb @(posedge clk);
    input data_out;
  endclocking

  // Modports for each component
  modport WR_BFM(clocking wr_cb);
  modport WR_MON(clocking wrmon_cb);
  modport RD_MON(clocking rdmon_cb);

endinterface: counter_if

Clocking blocks are the reason this testbench has no races. Every component samples and drives relative to posedge clk with the clocking block's default skew, so a value the BFM drives on one edge is seen by the monitors on the next, never in the same delta cycle. If you have ever debugged a testbench where the monitor sampled a signal a moment before the driver changed it, this is the cure.

3. Transaction class


class counter_trans;
  rand logic rst;
  rand logic load;
  rand logic updown;
  rand logic [3:0] data;
  logic [3:0] data_out;

  constraint c1 { rst dist {0:=95, 1:=5}; }
  constraint c2 { load dist {0:=70, 1:=30}; }

  function void display(string tag);
    $display("[%s] rst=%b load=%b updown=%b data=%h data_out=%h",
             tag, rst, load, updown, data, data_out);
  endfunction

  function bit compare(counter_trans t);
    return (this.data_out == t.data_out);
  endfunction
endclass

The distribution constraints shape the stimulus. Reset is rare (5 percent) so that long counting runs happen, and load is frequent enough (30 percent) to exercise the override path often. compare checks only data_out, because that is the only field the DUT produces; the inputs are what the scoreboard already knows.

4. Generator


class counter_gen;
  counter_trans trans_h;
  mailbox #(counter_trans) gen2wr;

  function new(mailbox #(counter_trans) gen2wr);
    this.gen2wr = gen2wr;
    trans_h = new();
  endfunction

  task start();
    fork
      for (int i = 0; i < no_of_transaction; i++) begin
        counter_trans t = new trans_h;
        assert(t.randomize());
        t.display("GENERATOR");
        gen2wr.put(t);
      end
    join_none
  endtask
endclass
new trans_h is a shallow copy. The generator randomizes one object and copies it before mailing, so every transaction in flight is a distinct object. Forgetting the copy is the classic beginner bug: every consumer ends up holding the same handle and sees the last randomization.

5. Write BFM (driver)


class counter_wr_bfm;
  virtual counter_if.WR_BFM wr_if;
  mailbox #(counter_trans) gen2wr;
  counter_trans trans_h;

  function new(virtual counter_if.WR_BFM wr_if,
               mailbox #(counter_trans) gen2wr);
    this.wr_if = wr_if;
    this.gen2wr = gen2wr;
  endfunction

  task drive();
    gen2wr.get(trans_h);
    @(wr_if.wr_cb);
    wr_if.wr_cb.rst    <= trans_h.rst;
    wr_if.wr_cb.load   <= trans_h.load;
    wr_if.wr_cb.updown <= trans_h.updown;
    wr_if.wr_cb.data   <= trans_h.data;
  endtask

  task start();
    fork
      forever drive();
    join_none
  endtask
endclass

The BFM is the only component allowed to drive the DUT inputs, and it does so exclusively through the wr_cb clocking block with nonblocking assignments. That combination guarantees the drive happens at the clock edge with the block's output skew and never races the DUT's own sampling.

6. Write monitor

This file was missing from the original post. It mirrors the read monitor: sample the input pins through wrmon_cb, package them into a transaction, and mail a copy to the reference model.


class counter_wr_mon;
  virtual counter_if.WR_MON wrmon_if;
  mailbox #(counter_trans) wrmon2rm;
  counter_trans trans_h;

  function new(virtual counter_if.WR_MON wrmon_if,
               mailbox #(counter_trans) wrmon2rm);
    this.wrmon_if = wrmon_if;
    this.wrmon2rm = wrmon2rm;
    trans_h = new();
  endfunction

  task monitor();
    @(wrmon_if.wrmon_cb);
    trans_h.rst    = wrmon_if.wrmon_cb.rst;
    trans_h.load   = wrmon_if.wrmon_cb.load;
    trans_h.updown = wrmon_if.wrmon_cb.updown;
    trans_h.data   = wrmon_if.wrmon_cb.data;
  endtask

  task start();
    fork
      forever begin
        counter_trans copy_h;
        monitor();
        trans_h.display("WR_MON");
        copy_h = new trans_h;      // shallow copy: the mailbox gets its own object
        wrmon2rm.put(copy_h);
      end
    join_none
  endtask
endclass

7. Read monitor


class counter_rd_mon;
  virtual counter_if.RD_MON rdmon_if;
  mailbox #(counter_trans) rdmon2sb;
  counter_trans trans_h;

  function new(virtual counter_if.RD_MON rdmon_if,
               mailbox #(counter_trans) rdmon2sb);
    this.rdmon_if = rdmon_if;
    this.rdmon2sb = rdmon2sb;
    trans_h = new;
  endfunction

  task monitor();
    @(rdmon_if.rdmon_cb);
    trans_h.data_out = rdmon_if.rdmon_cb.data_out;
    if ($isunknown(rdmon_if.rdmon_cb.data_out))
      trans_h.data_out = 0;
  endtask

  task start();
    fork
      forever begin
        counter_trans copy_h;
        monitor();
        trans_h.display("RD_MON");
        copy_h = new trans_h;      // shallow copy: the mailbox gets its own object
        rdmon2sb.put(copy_h);
      end
    join_none
  endtask
endclass

Both monitors make an explicit copy before put. The earlier version of this post wrote put(new trans_h) inline, which is legal SystemVerilog but which some tools, Verilator among them, do not parse as a method argument. A named copy is clearer anyway.

8. Reference model


class counter_rm;
  mailbox #(counter_trans) rm2sb, wrmon2rm;
  counter_trans wrmon2rm_h, temp_h;
  int count;

  function new(mailbox #(counter_trans) wrmon2rm,
               mailbox #(counter_trans) rm2sb);
    this.rm2sb = rm2sb;
    this.wrmon2rm = wrmon2rm;
    temp_h = new();
  endfunction

  task model();
    ++count;
    if (count > 1) begin
      temp_h.rst    = wrmon2rm_h.rst;
      temp_h.load   = wrmon2rm_h.load;
      temp_h.updown = wrmon2rm_h.updown;
      temp_h.data   = wrmon2rm_h.data;

      if (wrmon2rm_h.rst)
        temp_h.data_out = 0;
      else if (wrmon2rm_h.load)
        temp_h.data_out = wrmon2rm_h.data;
      else if (wrmon2rm_h.updown)
        temp_h.data_out = ++temp_h.data_out;
      else
        temp_h.data_out = --temp_h.data_out;
    end
  endtask

  task start();
    fork
      forever begin
        wrmon2rm.get(wrmon2rm_h);
        rm2sb.put(temp_h);
        temp_h = new temp_h;
        model();
      end
    join_none
  endtask
endclass

The model has a one-transaction offset built in, and the count > 1 guard is how it handles it. The DUT registers its output, so the value the read monitor captures on a given edge is the result of the inputs sampled on the previous edge. The reference model therefore mails its current prediction first, then computes the next one from the newly arrived inputs. Off-by-one errors between predicted and observed streams are the most common scoreboard failure in registered designs, and this is the place to reason about them.

9. Scoreboard with functional coverage


class counter_sb;
  mailbox #(counter_trans) rm2sb, rdmon2sb;
  event DONE;
  int count_transaction, data_verified;
  counter_trans cov_h, rcvd_h;

  covergroup counter_cov;
    option.per_instance = 1;
    RST:  coverpoint cov_h.rst     { bins r[] = {0, 1}; }
    LD:   coverpoint cov_h.load    { bins l[] = {0, 1}; }
    UD:   coverpoint cov_h.updown  { bins ud[] = {0, 1}; }
    DATA: coverpoint cov_h.data    { bins d[] = {[0:15]}; }
    DOUT: coverpoint cov_h.data_out { bins dout[] = {[0:15]}; }
    LDxDATA: cross LD, DATA;
    UDxDOUT: cross UD, DOUT;
  endgroup

  function new(mailbox #(counter_trans) rm2sb,
               mailbox #(counter_trans) rdmon2sb);
    this.rm2sb = rm2sb;
    this.rdmon2sb = rdmon2sb;
    counter_cov = new();
  endfunction

  task start;
    fork
      forever begin
        rm2sb.get(rcvd_h);
        cov_h = rcvd_h;
        counter_cov.sample();
        rdmon2sb.get(cov_h);
        check(rcvd_h);
      end
    join_none
  endtask

  task check(counter_trans rcvd_h);
    count_transaction++;
    if (cov_h.compare(rcvd_h)) begin
      counter_cov.sample();
      data_verified++;
    end
    if (count_transaction >= no_of_transaction)
      ->DONE;
  endtask

  function void report;
    $display("------------ SCOREBOARD REPORT ------------");
    $display("Transactions received : %0d", count_transaction);
    $display("Transactions verified : %0d", data_verified);
    $display("-------------------------------------------");
  endfunction
endclass

The scoreboard does three things: pulls a prediction and an observation, compares them, and samples coverage on the compared transaction. The covergroup crosses load with data and updown with data_out, which answers the questions "did we load every value" and "did we count both ways through every state". When count_transaction reaches the target the scoreboard fires DONE, and the environment's stop task uses that event to end the test cleanly instead of relying on a fixed delay.

10. Environment


class counter_env;
  virtual counter_if.WR_BFM wr_if;
  virtual counter_if.WR_MON wrmon_if;
  virtual counter_if.RD_MON rdmon_if;

  mailbox #(counter_trans) gen2wr   = new;
  mailbox #(counter_trans) wrmon2rm = new;
  mailbox #(counter_trans) rm2sb    = new;
  mailbox #(counter_trans) rdmon2sb = new;

  counter_gen    gen_h;
  counter_wr_bfm wr_h;
  counter_wr_mon wrmon_h;
  counter_rd_mon rdmon_h;
  counter_rm     rm_h;
  counter_sb     sb_h;

  function new(virtual counter_if.WR_BFM wr_if,
               virtual counter_if.WR_MON wrmon_if,
               virtual counter_if.RD_MON rdmon_if);
    this.wr_if    = wr_if;
    this.wrmon_if = wrmon_if;
    this.rdmon_if = rdmon_if;
  endfunction

  task build();
    gen_h   = new(gen2wr);
    wr_h    = new(wr_if, gen2wr);
    wrmon_h = new(wrmon_if, wrmon2rm);
    rdmon_h = new(rdmon_if, rdmon2sb);
    rm_h    = new(wrmon2rm, rm2sb);
    sb_h    = new(rm2sb, rdmon2sb);
  endtask

  task reset();
    @(wr_if.wr_cb);
    wr_if.wr_cb.rst <= 1;
    repeat(5) @(wr_if.wr_cb);
    wr_if.wr_cb.rst <= 0;
  endtask

  task start();
    gen_h.start(); wr_h.start(); wrmon_h.start();
    rdmon_h.start(); rm_h.start(); sb_h.start();
  endtask

  task stop();
    wait(sb_h.DONE.triggered);
  endtask

  task run();
    reset(); start(); stop(); sb_h.report();
  endtask
endclass

The environment owns the mailboxes and wires components together in build. run is the whole test lifecycle in one line: reset, start every component's forever loop, wait for the scoreboard to finish, print the report. In UVM this becomes build_phase, connect_phase, run_phase and report_phase, but the responsibilities are identical.

11. Test class

Also missing from the original. The test owns the environment and decides how many transactions to run, read from a plusarg so the same compile serves a smoke test and a long soak.


import counter_pkg::*;

class test;
  counter_env env_h;

  function new(virtual counter_if.WR_BFM wr_if,
               virtual counter_if.WR_MON wrmon_if,
               virtual counter_if.RD_MON rdmon_if);
    env_h = new(wr_if, wrmon_if, rdmon_if);
  endfunction

  task build_and_run();
    if (!$value$plusargs("TRANSACTIONS=%d", no_of_transaction))
      no_of_transaction = 100;
    env_h.build();
    env_h.run();
    $finish;
  endtask
endclass

12. Top module


`include "test.sv"

module top;
  reg clk;
  counter_if intf(clk);

  // DUT
  counter DUV(
    .clk(clk), .rst(intf.rst), .load(intf.load),
    .updown(intf.updown), .data(intf.data), .data_out(intf.data_out)
  );

  // Bind assertions
  bind DUV counter_assertion C_A(
    .clk(clk), .rst(intf.rst), .load(intf.load),
    .updown(intf.updown), .data(intf.data), .count(intf.data_out)
  );

  test test_h;

  initial begin
    test_h = new(intf, intf, intf);
    test_h.build_and_run();
  end

  initial begin
    clk = 0;
    forever #10 clk = ~clk;
  end
endmodule

The bind line attaches the assertion module to the DUT instance without editing the RTL. That is the standard way to keep protocol checks out of the design file while still having them fire in every simulation that instantiates the DUT.

13. SVA assertions


module counter_assertion(clk, rst, data, updown, load, count);
  input logic clk, rst, updown, load;
  input logic [3:0] data, count;

  // Reset clears counter
  property reset_prpty;
    @(posedge clk) rst |=> (count == 4'b0);
  endproperty

  // Up count
  property up_count_prpty;
    @(posedge clk) disable iff(rst)
    (!load && updown) |=> (count == ($past(count) + 1));
  endproperty

  // Down count
  property down_count_prpty;
    @(posedge clk) disable iff(rst)
    (!load && !updown) |=> (count == ($past(count) - 1));
  endproperty

  // Overflow: F -> 0
  property overflow_prpty;
    @(posedge clk) disable iff(rst)
    (!load && updown && count == 4'hF) |=> (count == 4'h0);
  endproperty

  // Underflow: 0 -> F
  property underflow_prpty;
    @(posedge clk) disable iff(rst)
    (!load && !updown && count == 4'h0) |=> (count == 4'hF);
  endproperty

  // Load data
  property load_prpty;
    @(posedge clk) disable iff(rst)
    load |=> (count == $past(data));
  endproperty

  RST:        assert property (reset_prpty);
  UP_COUNT:   assert property (up_count_prpty);
  DOWN_COUNT: assert property (down_count_prpty);
  OVERFLOW:   assert property (overflow_prpty);
  UNDERFLOW:  assert property (underflow_prpty);
  LOAD:       assert property (load_prpty);
endmodule

Each property is one line of the counter specification. disable iff (rst) keeps the counting properties quiet during reset, and $past compares against the previous cycle. The overflow and underflow properties are the ones most likely to catch a real bug, because wraparound is the case a hand-written directed test forgets. Assertions and the scoreboard are complementary: the scoreboard checks values end to end, the assertions check cycle-level behaviour and fail at the exact clock where it goes wrong.

14. Package and build


package counter_pkg;
  int no_of_transaction;

  `include "counter_trans.sv"
  `include "counter_gen.sv"
  `include "counter_wr_bfm.sv"
  `include "counter_wr_mon.sv"
  `include "counter_rd_mon.sv"
  `include "counter_rm.sv"
  `include "counter_sb.sv"
  `include "counter_env.sv"
endpackage

RTL     = ../rtl/counter.v
TB      = counter_if.sv counter_assertion.sv counter_pkg.sv top.sv
INC     = +incdir+. +incdir+../test
COVOPT  = -coveropt 3 +cover +acc

# Compile
compile:
	vlib work
	vlog $(COVOPT) $(RTL) $(TB) $(INC)

# Run test
run: compile
	vsim -c -coverage work.top +TEST1 \
	  -do "coverage save -onexit counter_cov; run -all; exit"
	vcover report -html counter_cov

# GUI mode
gui: compile
	vsim -coverage work.top +TEST1

clean:
	rm -rf work transcript *.wlf counter_cov* covhtmlreport

# Compile and run
make run

# View coverage report
firefox covhtmlreport/index.html

# Run in GUI
make gui

Compile order matters: the interface and assertion module first, then the package that includes every class, then top.sv which includes test.sv. Coverage is enabled at compile time and saved on exit, so make run leaves an HTML report behind.

What the bench looks like under Verilator

The full bench, all fourteen files, parses and elaborates under Verilator 5.052 with --lint-only --timing. The remaining warnings are style only: constructor arguments that shadow class members (VARHIDDEN), the assert(t.randomize()) idiom, which Verilator flags because randomize returns an int (WIDTHTRUNC), and the import counter_pkg::* in test.sv. None of them affect simulation on a commercial tool; fixing the shadowing by prefixing constructor arguments with i_ is a good exercise if you want a lint-clean bench.

Key takeaways

  • Clocking blocks in the interface remove driver and monitor races by construction.
  • Monitors feed the reference model from observed pins, not from the generator, so driver bugs are caught.
  • Copy transactions before mailing them; a shared handle is the most common layered-testbench bug.
  • A registered DUT means a one-transaction offset between prediction and observation. Handle it in the reference model deliberately.
  • End the test on a scoreboard event, not a fixed delay.
  • Bound SVA gives cycle-accurate checks without touching the RTL, and complements the scoreboard rather than replacing it.

Verified with Verilator 5.052 on macOS. Source files are listed in full above.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment