Complete SystemVerilog Testbench for a Counter: Interface, Monitors, Reference Model, Scoreboard and SVA
A UVM testbench hides a lot of machinery behind base classes. Before you use that machinery it pays to build the same structure once by hand, in plain SystemVerilog, so that every mailbox, every clocking block and every handshake is something you wrote and can debug. This post does exactly that for a 4-bit loadable up/down counter: a layered, self-checking testbench with a generator, a driver, two monitors, a reference model, a scoreboard with functional coverage, and a bound SVA module. The original January 2026 version of this post was mostly code. This rewrite keeps every file, adds the two that were missing (the write monitor and the test class), explains the design decisions between the listings, and reports what the whole bench looks like under Verilator 5.
build and run tasks to build_phase and run_phase. Learn it here where nothing is hidden, and UVM stops looking like magic.The data flow
flowchart TB
subgraph TB["Testbench"]
GEN[Generator] -->|trans| WR_BFM[Write BFM]
WR_BFM -->|drive| IF[Interface]
IF -->|monitor| WR_MON[Write Monitor]
IF -->|monitor| RD_MON[Read Monitor]
WR_MON -->|trans| RM[Reference Model]
RM -->|expected| SB[Scoreboard]
RD_MON -->|actual| SB
end
IF <-->|signals| DUT[Counter DUT]
DUT -.->|bind| ASSERT[SVA Assertions]
style GEN fill:#d1fae5,stroke:#10b981
style SB fill:#dbeafe,stroke:#3b82f6
style DUT fill:#fef3c7,stroke:#f59e0b
style ASSERT fill:#fee2e2,stroke:#ef4444
Stimulus flows left to right. The generator randomizes transactions and mails them to the write BFM, which drives the interface. The write monitor watches the same interface pins and mails what it saw to the reference model, not what the generator intended. That distinction matters: the reference model predicts from observed inputs, so a driver bug shows up as a mismatch instead of being hidden by a shared assumption. The read monitor captures the DUT output, and the scoreboard compares prediction against observation one transaction at a time. Assertions run alongside, bound to the DUT, and check the protocol cycle by cycle.
1. Design under test
The counter is deliberately small so the testbench is the subject. Synchronous reset, a load input that overrides counting, and an updown control.
module counter(clk, rst, data, updown, load, data_out);
input clk, rst, load;
input updown;
input [3:0] data;
output reg [3:0] data_out;
always @(posedge clk) begin
if (rst)
data_out <= 4'b0;
else if (load)
data_out <= data;
else
data_out <= (updown) ? (data_out + 1'b1) : (data_out - 1'b1);
end
endmodule
2. Interface with clocking blocks
The interface declares the pins once and then defines three clocking blocks, one for each component that touches the pins, with directions that match the component's role. Modports hand each component only its own clocking block.
interface counter_if(input logic clk);
logic rst, updown, load;
logic [3:0] data;
logic [3:0] data_out;
// Write BFM clocking block (driver)
clocking wr_cb @(posedge clk);
output load, updown, rst;
output data;
endclocking
// Write Monitor clocking block
clocking wrmon_cb @(posedge clk);
input data;
input load, rst, updown;
endclocking
// Read Monitor clocking block
clocking rdmon_cb @(posedge clk);
input data_out;
endclocking
// Modports for each component
modport WR_BFM(clocking wr_cb);
modport WR_MON(clocking wrmon_cb);
modport RD_MON(clocking rdmon_cb);
endinterface: counter_if
Clocking blocks are the reason this testbench has no races. Every component samples and drives relative to posedge clk with the clocking block's default skew, so a value the BFM drives on one edge is seen by the monitors on the next, never in the same delta cycle. If you have ever debugged a testbench where the monitor sampled a signal a moment before the driver changed it, this is the cure.
3. Transaction class
class counter_trans;
rand logic rst;
rand logic load;
rand logic updown;
rand logic [3:0] data;
logic [3:0] data_out;
constraint c1 { rst dist {0:=95, 1:=5}; }
constraint c2 { load dist {0:=70, 1:=30}; }
function void display(string tag);
$display("[%s] rst=%b load=%b updown=%b data=%h data_out=%h",
tag, rst, load, updown, data, data_out);
endfunction
function bit compare(counter_trans t);
return (this.data_out == t.data_out);
endfunction
endclass
The distribution constraints shape the stimulus. Reset is rare (5 percent) so that long counting runs happen, and load is frequent enough (30 percent) to exercise the override path often. compare checks only data_out, because that is the only field the DUT produces; the inputs are what the scoreboard already knows.
4. Generator
class counter_gen;
counter_trans trans_h;
mailbox #(counter_trans) gen2wr;
function new(mailbox #(counter_trans) gen2wr);
this.gen2wr = gen2wr;
trans_h = new();
endfunction
task start();
fork
for (int i = 0; i < no_of_transaction; i++) begin
counter_trans t = new trans_h;
assert(t.randomize());
t.display("GENERATOR");
gen2wr.put(t);
end
join_none
endtask
endclass
new trans_h is a shallow copy. The generator randomizes one object and copies it before mailing, so every transaction in flight is a distinct object. Forgetting the copy is the classic beginner bug: every consumer ends up holding the same handle and sees the last randomization.
5. Write BFM (driver)
class counter_wr_bfm;
virtual counter_if.WR_BFM wr_if;
mailbox #(counter_trans) gen2wr;
counter_trans trans_h;
function new(virtual counter_if.WR_BFM wr_if,
mailbox #(counter_trans) gen2wr);
this.wr_if = wr_if;
this.gen2wr = gen2wr;
endfunction
task drive();
gen2wr.get(trans_h);
@(wr_if.wr_cb);
wr_if.wr_cb.rst <= trans_h.rst;
wr_if.wr_cb.load <= trans_h.load;
wr_if.wr_cb.updown <= trans_h.updown;
wr_if.wr_cb.data <= trans_h.data;
endtask
task start();
fork
forever drive();
join_none
endtask
endclass
The BFM is the only component allowed to drive the DUT inputs, and it does so exclusively through the wr_cb clocking block with nonblocking assignments. That combination guarantees the drive happens at the clock edge with the block's output skew and never races the DUT's own sampling.
6. Write monitor
This file was missing from the original post. It mirrors the read monitor: sample the input pins through wrmon_cb, package them into a transaction, and mail a copy to the reference model.
class counter_wr_mon;
virtual counter_if.WR_MON wrmon_if;
mailbox #(counter_trans) wrmon2rm;
counter_trans trans_h;
function new(virtual counter_if.WR_MON wrmon_if,
mailbox #(counter_trans) wrmon2rm);
this.wrmon_if = wrmon_if;
this.wrmon2rm = wrmon2rm;
trans_h = new();
endfunction
task monitor();
@(wrmon_if.wrmon_cb);
trans_h.rst = wrmon_if.wrmon_cb.rst;
trans_h.load = wrmon_if.wrmon_cb.load;
trans_h.updown = wrmon_if.wrmon_cb.updown;
trans_h.data = wrmon_if.wrmon_cb.data;
endtask
task start();
fork
forever begin
counter_trans copy_h;
monitor();
trans_h.display("WR_MON");
copy_h = new trans_h; // shallow copy: the mailbox gets its own object
wrmon2rm.put(copy_h);
end
join_none
endtask
endclass
7. Read monitor
class counter_rd_mon;
virtual counter_if.RD_MON rdmon_if;
mailbox #(counter_trans) rdmon2sb;
counter_trans trans_h;
function new(virtual counter_if.RD_MON rdmon_if,
mailbox #(counter_trans) rdmon2sb);
this.rdmon_if = rdmon_if;
this.rdmon2sb = rdmon2sb;
trans_h = new;
endfunction
task monitor();
@(rdmon_if.rdmon_cb);
trans_h.data_out = rdmon_if.rdmon_cb.data_out;
if ($isunknown(rdmon_if.rdmon_cb.data_out))
trans_h.data_out = 0;
endtask
task start();
fork
forever begin
counter_trans copy_h;
monitor();
trans_h.display("RD_MON");
copy_h = new trans_h; // shallow copy: the mailbox gets its own object
rdmon2sb.put(copy_h);
end
join_none
endtask
endclass
Both monitors make an explicit copy before put. The earlier version of this post wrote put(new trans_h) inline, which is legal SystemVerilog but which some tools, Verilator among them, do not parse as a method argument. A named copy is clearer anyway.
8. Reference model
class counter_rm;
mailbox #(counter_trans) rm2sb, wrmon2rm;
counter_trans wrmon2rm_h, temp_h;
int count;
function new(mailbox #(counter_trans) wrmon2rm,
mailbox #(counter_trans) rm2sb);
this.rm2sb = rm2sb;
this.wrmon2rm = wrmon2rm;
temp_h = new();
endfunction
task model();
++count;
if (count > 1) begin
temp_h.rst = wrmon2rm_h.rst;
temp_h.load = wrmon2rm_h.load;
temp_h.updown = wrmon2rm_h.updown;
temp_h.data = wrmon2rm_h.data;
if (wrmon2rm_h.rst)
temp_h.data_out = 0;
else if (wrmon2rm_h.load)
temp_h.data_out = wrmon2rm_h.data;
else if (wrmon2rm_h.updown)
temp_h.data_out = ++temp_h.data_out;
else
temp_h.data_out = --temp_h.data_out;
end
endtask
task start();
fork
forever begin
wrmon2rm.get(wrmon2rm_h);
rm2sb.put(temp_h);
temp_h = new temp_h;
model();
end
join_none
endtask
endclass
The model has a one-transaction offset built in, and the count > 1 guard is how it handles it. The DUT registers its output, so the value the read monitor captures on a given edge is the result of the inputs sampled on the previous edge. The reference model therefore mails its current prediction first, then computes the next one from the newly arrived inputs. Off-by-one errors between predicted and observed streams are the most common scoreboard failure in registered designs, and this is the place to reason about them.
9. Scoreboard with functional coverage
class counter_sb;
mailbox #(counter_trans) rm2sb, rdmon2sb;
event DONE;
int count_transaction, data_verified;
counter_trans cov_h, rcvd_h;
covergroup counter_cov;
option.per_instance = 1;
RST: coverpoint cov_h.rst { bins r[] = {0, 1}; }
LD: coverpoint cov_h.load { bins l[] = {0, 1}; }
UD: coverpoint cov_h.updown { bins ud[] = {0, 1}; }
DATA: coverpoint cov_h.data { bins d[] = {[0:15]}; }
DOUT: coverpoint cov_h.data_out { bins dout[] = {[0:15]}; }
LDxDATA: cross LD, DATA;
UDxDOUT: cross UD, DOUT;
endgroup
function new(mailbox #(counter_trans) rm2sb,
mailbox #(counter_trans) rdmon2sb);
this.rm2sb = rm2sb;
this.rdmon2sb = rdmon2sb;
counter_cov = new();
endfunction
task start;
fork
forever begin
rm2sb.get(rcvd_h);
cov_h = rcvd_h;
counter_cov.sample();
rdmon2sb.get(cov_h);
check(rcvd_h);
end
join_none
endtask
task check(counter_trans rcvd_h);
count_transaction++;
if (cov_h.compare(rcvd_h)) begin
counter_cov.sample();
data_verified++;
end
if (count_transaction >= no_of_transaction)
->DONE;
endtask
function void report;
$display("------------ SCOREBOARD REPORT ------------");
$display("Transactions received : %0d", count_transaction);
$display("Transactions verified : %0d", data_verified);
$display("-------------------------------------------");
endfunction
endclass
The scoreboard does three things: pulls a prediction and an observation, compares them, and samples coverage on the compared transaction. The covergroup crosses load with data and updown with data_out, which answers the questions "did we load every value" and "did we count both ways through every state". When count_transaction reaches the target the scoreboard fires DONE, and the environment's stop task uses that event to end the test cleanly instead of relying on a fixed delay.
10. Environment
class counter_env;
virtual counter_if.WR_BFM wr_if;
virtual counter_if.WR_MON wrmon_if;
virtual counter_if.RD_MON rdmon_if;
mailbox #(counter_trans) gen2wr = new;
mailbox #(counter_trans) wrmon2rm = new;
mailbox #(counter_trans) rm2sb = new;
mailbox #(counter_trans) rdmon2sb = new;
counter_gen gen_h;
counter_wr_bfm wr_h;
counter_wr_mon wrmon_h;
counter_rd_mon rdmon_h;
counter_rm rm_h;
counter_sb sb_h;
function new(virtual counter_if.WR_BFM wr_if,
virtual counter_if.WR_MON wrmon_if,
virtual counter_if.RD_MON rdmon_if);
this.wr_if = wr_if;
this.wrmon_if = wrmon_if;
this.rdmon_if = rdmon_if;
endfunction
task build();
gen_h = new(gen2wr);
wr_h = new(wr_if, gen2wr);
wrmon_h = new(wrmon_if, wrmon2rm);
rdmon_h = new(rdmon_if, rdmon2sb);
rm_h = new(wrmon2rm, rm2sb);
sb_h = new(rm2sb, rdmon2sb);
endtask
task reset();
@(wr_if.wr_cb);
wr_if.wr_cb.rst <= 1;
repeat(5) @(wr_if.wr_cb);
wr_if.wr_cb.rst <= 0;
endtask
task start();
gen_h.start(); wr_h.start(); wrmon_h.start();
rdmon_h.start(); rm_h.start(); sb_h.start();
endtask
task stop();
wait(sb_h.DONE.triggered);
endtask
task run();
reset(); start(); stop(); sb_h.report();
endtask
endclass
The environment owns the mailboxes and wires components together in build. run is the whole test lifecycle in one line: reset, start every component's forever loop, wait for the scoreboard to finish, print the report. In UVM this becomes build_phase, connect_phase, run_phase and report_phase, but the responsibilities are identical.
11. Test class
Also missing from the original. The test owns the environment and decides how many transactions to run, read from a plusarg so the same compile serves a smoke test and a long soak.
import counter_pkg::*;
class test;
counter_env env_h;
function new(virtual counter_if.WR_BFM wr_if,
virtual counter_if.WR_MON wrmon_if,
virtual counter_if.RD_MON rdmon_if);
env_h = new(wr_if, wrmon_if, rdmon_if);
endfunction
task build_and_run();
if (!$value$plusargs("TRANSACTIONS=%d", no_of_transaction))
no_of_transaction = 100;
env_h.build();
env_h.run();
$finish;
endtask
endclass
12. Top module
`include "test.sv"
module top;
reg clk;
counter_if intf(clk);
// DUT
counter DUV(
.clk(clk), .rst(intf.rst), .load(intf.load),
.updown(intf.updown), .data(intf.data), .data_out(intf.data_out)
);
// Bind assertions
bind DUV counter_assertion C_A(
.clk(clk), .rst(intf.rst), .load(intf.load),
.updown(intf.updown), .data(intf.data), .count(intf.data_out)
);
test test_h;
initial begin
test_h = new(intf, intf, intf);
test_h.build_and_run();
end
initial begin
clk = 0;
forever #10 clk = ~clk;
end
endmodule
The bind line attaches the assertion module to the DUT instance without editing the RTL. That is the standard way to keep protocol checks out of the design file while still having them fire in every simulation that instantiates the DUT.
13. SVA assertions
module counter_assertion(clk, rst, data, updown, load, count);
input logic clk, rst, updown, load;
input logic [3:0] data, count;
// Reset clears counter
property reset_prpty;
@(posedge clk) rst |=> (count == 4'b0);
endproperty
// Up count
property up_count_prpty;
@(posedge clk) disable iff(rst)
(!load && updown) |=> (count == ($past(count) + 1));
endproperty
// Down count
property down_count_prpty;
@(posedge clk) disable iff(rst)
(!load && !updown) |=> (count == ($past(count) - 1));
endproperty
// Overflow: F -> 0
property overflow_prpty;
@(posedge clk) disable iff(rst)
(!load && updown && count == 4'hF) |=> (count == 4'h0);
endproperty
// Underflow: 0 -> F
property underflow_prpty;
@(posedge clk) disable iff(rst)
(!load && !updown && count == 4'h0) |=> (count == 4'hF);
endproperty
// Load data
property load_prpty;
@(posedge clk) disable iff(rst)
load |=> (count == $past(data));
endproperty
RST: assert property (reset_prpty);
UP_COUNT: assert property (up_count_prpty);
DOWN_COUNT: assert property (down_count_prpty);
OVERFLOW: assert property (overflow_prpty);
UNDERFLOW: assert property (underflow_prpty);
LOAD: assert property (load_prpty);
endmodule
Each property is one line of the counter specification. disable iff (rst) keeps the counting properties quiet during reset, and $past compares against the previous cycle. The overflow and underflow properties are the ones most likely to catch a real bug, because wraparound is the case a hand-written directed test forgets. Assertions and the scoreboard are complementary: the scoreboard checks values end to end, the assertions check cycle-level behaviour and fail at the exact clock where it goes wrong.
14. Package and build
package counter_pkg;
int no_of_transaction;
`include "counter_trans.sv"
`include "counter_gen.sv"
`include "counter_wr_bfm.sv"
`include "counter_wr_mon.sv"
`include "counter_rd_mon.sv"
`include "counter_rm.sv"
`include "counter_sb.sv"
`include "counter_env.sv"
endpackage
RTL = ../rtl/counter.v
TB = counter_if.sv counter_assertion.sv counter_pkg.sv top.sv
INC = +incdir+. +incdir+../test
COVOPT = -coveropt 3 +cover +acc
# Compile
compile:
vlib work
vlog $(COVOPT) $(RTL) $(TB) $(INC)
# Run test
run: compile
vsim -c -coverage work.top +TEST1 \
-do "coverage save -onexit counter_cov; run -all; exit"
vcover report -html counter_cov
# GUI mode
gui: compile
vsim -coverage work.top +TEST1
clean:
rm -rf work transcript *.wlf counter_cov* covhtmlreport
# Compile and run
make run
# View coverage report
firefox covhtmlreport/index.html
# Run in GUI
make gui
Compile order matters: the interface and assertion module first, then the package that includes every class, then top.sv which includes test.sv. Coverage is enabled at compile time and saved on exit, so make run leaves an HTML report behind.
What the bench looks like under Verilator
The full bench, all fourteen files, parses and elaborates under Verilator 5.052 with --lint-only --timing. The remaining warnings are style only: constructor arguments that shadow class members (VARHIDDEN), the assert(t.randomize()) idiom, which Verilator flags because randomize returns an int (WIDTHTRUNC), and the import counter_pkg::* in test.sv. None of them affect simulation on a commercial tool; fixing the shadowing by prefixing constructor arguments with i_ is a good exercise if you want a lint-clean bench.
Key takeaways
- Clocking blocks in the interface remove driver and monitor races by construction.
- Monitors feed the reference model from observed pins, not from the generator, so driver bugs are caught.
- Copy transactions before mailing them; a shared handle is the most common layered-testbench bug.
- A registered DUT means a one-transaction offset between prediction and observation. Handle it in the reference model deliberately.
- End the test on a scoreboard event, not a fixed delay.
- Bound SVA gives cycle-accurate checks without touching the RTL, and complements the scoreboard rather than replacing it.
Verified with Verilator 5.052 on macOS. Source files are listed in full above.
Comments (0)
Leave a Comment