6. SystemC Tutorial - Your First SystemC Testbench: Kernel-Level Stimulus
Why this matters
Rewritten 2026-05-23 with deeper first-principles material.
Almost every SystemC tutorial after this one will assume you can write a working testbench from scratch — that you can take a module, hook it up to stimulus, run a simulation, and inspect the result. UVM-SystemC, the SCV constrained-random library, and every framework built on top of SystemC ultimately boil down to "a testbench module with some processes drives signals into your DUT and another set of processes observes." If the framework layer makes the inner mechanism opaque, you cannot debug it when something goes wrong. The kernel-level testbench — no UVM, no SCV, no frameworks — is the irreducible unit. Once you can write one of those by hand, you can write any testbench, because everything else is library code on top of the same primitives.
This post builds three testbenches of increasing realism: a one-page sc_main-driven scaffold for "does it work at all" exploration, a separate-driver-and-monitor pattern that scales to non-trivial sequences, and a self-checking testbench with a scoreboard that calls sc_stop() when done. You will leave able to write any of the three; able to add VCD tracing in ten lines for waveform-based debugging; and able to read the testbench code that appears in every subsequent post in this series.
Prerequisites
- Part 1: Modules, Ports & Signals — for the DUT skeleton
- Part 2: Simulation Time & Clocks — for
sc_startandsc_clock - Part 3: Delta Cycles & Event-Driven Semantics — for understanding when writes become visible
- Part 4: Processes & Sensitivity — for choosing
SC_METHODvsSC_THREADin the testbench - SystemC 2.3.x installed (P1, P2)
- For VCD output viewing: GTKWave or any other VCD-aware waveform viewer
Mental model (first principles)
A SystemC testbench is just another module (or two, or several) — there is no separate "testbench language." Your stimulus modules connect to your DUT's input ports through sc_signal channels, and your observer modules connect to its output ports the same way. The kernel does not distinguish "DUT" from "TB"; both are equally citizens of the simulation. What makes a module a testbench is its role — drive, observe, check — not any structural difference.
Kernel-level testbench
──────────────────────
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ driver │ signals │ DUT │ signals │ monitor │
│ (SC_THREAD) │────────▶│ (your code) │────────▶│ (SC_METHOD) │
└─────────────┘ └─────────────┘ └─────────────┘
│ │
│ ▼
│ ┌─────────────┐
│ │ scoreboard │
│ ┌─────────────┐ │ compare │
└──────▶│ reference │──── expected ─────▶│ pass/fail │
│ model │ └─────┬───────┘
└─────────────┘ │
▼
sc_stop()
at end
sc_main constructs all of these, binds the signals,
calls sc_start, then returns. The kernel runs the
show in between.
Three roles to keep in mind. The driver issues stimulus — a sequence of input values, possibly synchronised to a clock. The monitor observes outputs — typically samples a signal on each clock edge or value change. The scoreboard (optional but recommended) compares what the monitor saw against what the reference model said should happen, accumulating pass/fail counts. Together they form a testbench. Any of the three roles can be omitted if the test is trivial enough — the simplest testbench has only sc_main itself as driver and monitor combined.
The kernel's job is the same as always: schedule processes per the 3-phase delta-cycle loop from Part 3, advance time when the runnable set is empty, return control to sc_main when sc_start's time limit is reached or sc_stop() is called. The testbench modules use exactly the same primitives as the DUT.
Beginner: First Principles
The simplest testbench: drive a value into a module, run the simulation, print the result. We use the adder8 module from Part 1 as the DUT:
// file: first_testbench.cpp
#include <systemc.h>
SC_MODULE(adder8) {
sc_in<sc_uint<8>> a, b;
sc_out<sc_uint<8>> y;
void eval() { y.write(a.read() + b.read()); }
SC_CTOR(adder8) {
SC_METHOD(eval);
sensitive << a << b;
}
};
int sc_main(int, char**) {
// Signals the testbench will use to talk to the DUT
sc_signal<sc_uint<8>> sa, sb, sy;
adder8 dut("dut");
dut.a(sa); dut.b(sb); dut.y(sy);
// Test vectors as a simple table
struct test_vec { int a, b, expected; };
test_vec vectors[] = {
{ 3, 4, 7 },
{ 10, 20, 30 },
{255, 2, 1 }, // 8-bit wrap
{ 0, 0, 0 },
{100, 100, 200 },
};
int passed = 0, failed = 0;
for (auto& v : vectors) {
sa.write(v.a);
sb.write(v.b);
sc_start(1, SC_NS);
unsigned got = sy.read();
bool ok = (got == (v.expected & 0xff));
std::cout << "[" << sc_time_stamp() << "] "
<< v.a << " + " << v.b << " = " << got
<< " (expected " << v.expected << ") "
<< (ok ? "PASS" : "FAIL") << "\n";
if (ok) ++passed; else ++failed;
}
std::cout << "\n========== TEST SUMMARY ==========\n";
std::cout << "Passed: " << passed << ", Failed: " << failed << "\n";
std::cout << "===================================\n";
return failed > 0 ? 1 : 0;
}
Build and run:
g++ -std=c++17 first_testbench.cpp -I$SYSTEMC_HOME/include \
-L$SYSTEMC_HOME/lib-linux64 -lsystemc -o first_testbench
./first_testbench
Expected output:
SystemC 2.3.4-Accellera --- ...
[1 ns] 3 + 4 = 7 (expected 7) PASS
[2 ns] 10 + 20 = 30 (expected 30) PASS
[3 ns] 255 + 2 = 1 (expected 1) PASS
[4 ns] 0 + 0 = 0 (expected 0) PASS
[5 ns] 100 + 100 = 200 (expected 200) PASS
========== TEST SUMMARY ==========
Passed: 5, Failed: 0
===================================
Five things to walk through.
First, sc_main is the testbench. No separate "test class," no SC_MODULE for stimulus. The sc_main function itself drives values into the signals and reads results back, advancing time with sc_start(1, SC_NS) between each test vector. This pattern is the simplest possible scaffold and the right starting point for any new module you want to exercise.
Second, stimulus is a plain C++ data structure. The vectors[] array is a regular C++ array of structs. There is no SystemC-specific stimulus type. You can build vectors from files, from random generation, from any C++ source you like. The framework imposes no constraints on how stimulus is represented.
Third, time advances under your control. sc_start(1, SC_NS) says "run the kernel for up to 1 ns of simulation time, then return to me." Between calls, simulation is paused — you can read sy, write to sa/sb, do any C++ computation. The kernel resumes from exactly the same state on the next sc_start call.
Fourth, comparison is inline. This testbench scores its own results inside the for loop. For a small module this is the right pattern; there is no need to introduce a separate scoreboard module. When testbenches grow, separating concerns into driver/monitor/scoreboard modules pays off — but for five test vectors, inline is fine.
Fifth, sc_main's return value is the simulation's exit code. 0 for success, non-zero for failure. This makes the testbench friendly to CI systems: a script can run the binary, check the exit code, and report pass/fail without parsing output.
A common confusion at this point: "Where does time advance happen, exactly?" The kernel does not toggle time forward smoothly; it jumps to the next interesting time. sc_start(1, SC_NS) says "advance up to 1 ns." Since the adder8::eval method runs as soon as inputs change (no clock involved), it completes in the same delta cycle as the input writes. The kernel finds nothing more to do, returns. The current time is 1 ns even though the work happened at 0+δ ns — see Part 2 Corner 9 for the precise guarantees about sc_time_stamp() between sc_start calls.
Intermediate: How It Really Works
Separating concerns: driver + monitor
The sc_main-as-everything pattern works for trivial tests but does not scale. The next step is to put the stimulus and observation into dedicated modules, each with their own process. The DUT in this example will be the synchronous counter from Part 2:
// file: driver_monitor_tb.cpp
#include <systemc.h>
SC_MODULE(counter) {
sc_in<bool> clk;
sc_in<bool> rst_n;
sc_out<sc_uint<8>> count;
sc_uint<8> value;
void tick() {
if (!rst_n.read()) value = 0;
else value = value + 1;
count.write(value);
}
SC_CTOR(counter) : value(0) {
SC_METHOD(tick);
sensitive << clk.pos();
dont_initialize();
}
};
SC_MODULE(driver) {
sc_in<bool> clk;
sc_out<bool> rst_n;
void run() {
std::cout << "[" << sc_time_stamp() << "] driver: asserting reset\n";
rst_n.write(false);
wait(2, SC_NS);
std::cout << "[" << sc_time_stamp() << "] driver: releasing reset\n";
rst_n.write(true);
// Let the counter run for 10 cycles
for (int i = 0; i < 10; ++i) {
wait(clk.posedge_event());
}
std::cout << "[" << sc_time_stamp() << "] driver: stopping simulation\n";
sc_stop();
}
SC_CTOR(driver) { SC_THREAD(run); }
};
SC_MODULE(monitor) {
sc_in<bool> clk;
sc_in<sc_uint<8>> count;
void sample() {
std::cout << "[" << sc_time_stamp() << "] monitor: count="
<< count.read() << "\n";
}
SC_CTOR(monitor) {
SC_METHOD(sample);
sensitive << clk.pos();
dont_initialize();
}
};
int sc_main(int, char**) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst_n;
sc_signal<sc_uint<8>> count_sig;
counter dut("dut");
driver drv("drv");
monitor mon("mon");
dut.clk(clk);
dut.rst_n(rst_n);
dut.count(count_sig);
drv.clk(clk);
drv.rst_n(rst_n);
mon.clk(clk);
mon.count(count_sig);
sc_start(); // run until driver calls sc_stop()
return 0;
}
Expected output (abbreviated):
[0 s] driver: asserting reset
[2 ns] driver: releasing reset
[10 ns] monitor: count=0 (first clk.pos; reset still asserted at 10 ns? no — released at 2 ns)
[10 ns] monitor: count=0 (wait — let's reason about this carefully below)
...
Actually let me trace this more carefully. The driver writes rst_n = false at time 0 (pending, committed in init update). Calls wait(2, SC_NS) and suspends. At 2 ns, driver resumes, writes rst_n = true (pending, committed in next update). Calls wait(clk.posedge_event()) and suspends. At 10 ns (first clk posedge per default posedge_first = true), the counter's tick fires — but rst_n is now true (committed at 2 ns + δ), so value = 0 + 1 = 1, count.write(1). Update phase commits count = 1. Monitor's sample also fires at 10 ns posedge, reads count = 1 (the just-committed value).
So the actual output is:
[0 s] driver: asserting reset
[2 ns] driver: releasing reset
[10 ns] monitor: count=1
[20 ns] monitor: count=2
[30 ns] monitor: count=3
...
[100 ns] monitor: count=10
[100 ns] driver: stopping simulation
(The relative order of monitor and driver prints at 100 ns is implementation-defined since both fire on the same clock edge; in the Accellera PoC, registration order determines it. The driver was registered first, so depending on which clock-edge consumer the kernel processes first, the output may differ slightly.)
The interesting design point: notice that this testbench produces the same simulation as sc_main-only driving — but the responsibilities are decomposed. The driver knows nothing about how the monitor checks; the monitor knows nothing about how the driver generates stimulus. Each module is independently testable, replaceable, and reusable in other testbenches.
Adding VCD tracing
Waveform-based debugging is invaluable for anything more complex than the smallest module. SystemC includes VCD support out of the box:
// In sc_main, before sc_start:
sc_trace_file* tf = sc_create_vcd_trace_file("waves");
tf->set_time_unit(1, SC_NS);
sc_trace(tf, clk, "clk");
sc_trace(tf, rst_n, "rst_n");
sc_trace(tf, count_sig, "count");
sc_start(); // call as usual; trace captures everything until close
// After sc_start returns:
sc_close_vcd_trace_file(tf);
This produces waves.vcd in the current directory. Open it with GTKWave (gtkwave waves.vcd) and you get a waveform viewer showing the clock, reset, and count value over the simulation. Every change is captured; the file is plain-text and human-readable if you want to inspect it directly.
Three things worth knowing about VCD tracing:
set_time_unit(1, SC_NS)controls the display unit in the VCD file, not the kernel's time resolution. The kernel still operates at whatever resolutionsc_set_time_resolutionset.- You must call
sc_trace(...)for every signal you want in the trace. There is no "trace all signals" mode in standard SystemC (some commercial simulators add it as a non-standard extension). sc_traceworks on any type SystemC knows how to trace:bool,sc_uint<W>,sc_int<W>,sc_logic,sc_bv<W>,sc_lv<W>, plainint/long, etc. For user-defined struct types, you must overloadsc_traceyourself (see the LRM §5.18.5 for the signature).
Self-checking testbench with a scoreboard
The driver_monitor pattern decomposes stimulus and observation. The next step is checking — a separate module that knows what the output should be and compares against what it actually is.
// file: scoreboard_tb.cpp
#include <systemc.h>
#include <queue>
SC_MODULE(driver) {
sc_in<bool> clk;
sc_out<sc_uint<8>> din;
sc_out<bool> valid;
std::queue<sc_uint<8>>* expected; // reference to scoreboard's queue
void run() {
valid.write(false);
din.write(0);
wait(20, SC_NS);
for (sc_uint<8> v = 1; v <= 10; ++v) {
wait(clk.posedge_event());
din.write(v);
valid.write(true);
expected->push(v); // tell the scoreboard what to expect
std::cout << "[" << sc_time_stamp() << "] driver: drove " << v << "\n";
wait(clk.posedge_event());
valid.write(false);
}
wait(20, SC_NS);
sc_stop();
}
SC_CTOR(driver) { SC_THREAD(run); }
};
SC_MODULE(scoreboard) {
sc_in<bool> clk;
sc_in<sc_uint<8>> observed;
sc_in<bool> observed_valid;
std::queue<sc_uint<8>> expected_queue;
int passes = 0, failures = 0;
void check() {
if (!observed_valid.read()) return;
if (expected_queue.empty()) {
std::cout << "[" << sc_time_stamp() << "] scoreboard: "
<< "FAIL — observed " << observed.read()
<< " but expected queue is empty\n";
++failures;
return;
}
sc_uint<8> e = expected_queue.front();
expected_queue.pop();
sc_uint<8> got = observed.read();
if (got == e) {
std::cout << "[" << sc_time_stamp() << "] scoreboard: "
<< "PASS — got " << got << "\n";
++passes;
} else {
std::cout << "[" << sc_time_stamp() << "] scoreboard: "
<< "FAIL — got " << got << " expected " << e << "\n";
++failures;
}
}
SC_CTOR(scoreboard) {
SC_METHOD(check);
sensitive << clk.pos();
dont_initialize();
}
~scoreboard() {
std::cout << "\n========== TEST SUMMARY ==========\n";
std::cout << "Passed: " << passes << ", Failed: " << failures
<< ", Outstanding: " << expected_queue.size() << "\n";
std::cout << "===================================\n";
}
};
// A trivial DUT that just passes input through with one cycle delay
SC_MODULE(simple_dut) {
sc_in<bool> clk;
sc_in<sc_uint<8>> in;
sc_in<bool> in_valid;
sc_out<sc_uint<8>> out;
sc_out<bool> out_valid;
void tick() {
out.write(in.read());
out_valid.write(in_valid.read());
}
SC_CTOR(simple_dut) {
SC_METHOD(tick);
sensitive << clk.pos();
dont_initialize();
}
};
int sc_main(int, char**) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<sc_uint<8>> din, dout;
sc_signal<bool> din_valid, dout_valid;
simple_dut dut("dut");
driver drv("drv");
scoreboard sb("sb");
dut.clk(clk); dut.in(din); dut.in_valid(din_valid);
dut.out(dout); dut.out_valid(dout_valid);
drv.clk(clk); drv.din(din); drv.valid(din_valid);
drv.expected = &sb.expected_queue;
sb.clk(clk); sb.observed(dout); sb.observed_valid(dout_valid);
sc_start(); // until driver calls sc_stop()
return sb.passes > 0 && sb.failures == 0 ? 0 : 1;
}
The scoreboard is just another module — SC_METHOD triggered on every clock edge, comparing the observed value against the expected queue. The driver pushes expected values onto the scoreboard's queue as it drives them. When sc_stop() is called, the scoreboard's destructor runs (when sc_main returns and modules go out of scope) and prints the summary.
Three patterns worth internalising. First, the scoreboard owns the expected queue, not the driver — the driver just pushes expected values. This keeps comparison logic in one place. Second, the destructor prints the summary — a slightly unusual pattern, but it works because module destruction happens after sc_main returns. Third, sc_main's return value reflects test outcome — non-zero on any failure or zero passes, making CI integration straightforward.
Bounded vs unbounded testbenches — a watchdog pattern
A testbench that runs forever if something goes wrong is a CI nightmare. Two patterns to bound the runtime:
// Pattern A: unbounded, terminated by sc_stop() from the driver
sc_start(); // runs until driver's internal logic decides "done"
// Pattern B: bounded by a hard time limit (watchdog)
sc_start(sc_time(10, SC_MS)); // never runs longer than 10 ms simulation time
// Pattern C: bounded with both — driver tries to sc_stop, but we cap regardless
sc_start(sc_time(10, SC_MS));
if (sc_get_status() == SC_RUNNING) {
std::cerr << "TIMEOUT — driver did not terminate within 10 ms\n";
}
sc_stop();
Pattern C is the safest production pattern. Even if your driver hangs (a wait that never gets its event, for example), the simulation returns after 10 ms simulated time and you get a clear "TIMEOUT" message instead of a hung CI job.
Performance note: trace file size
VCD files grow linearly with the number of traced signals and the simulation length. For long simulations with many signals, the trace file can exceed gigabytes. Two practices help:
- Trace only what you need to debug. Default to no tracing in CI runs; enable tracing only when investigating a failure.
- Use
sc_close_vcd_trace_fileearly if you only need a specific window of the simulation — open the trace, run the relevant window, close it, then continue without tracing.
For production-grade waveform capture, consider SST (the Simulation Snapshot Toolkit) or proprietary formats like fsdb (used by Verdi). These compress vastly better than VCD but require simulator-specific tooling.
Driver patterns: hand-rolled, table-driven, and file-fed
Three increasingly flexible stimulus-generation patterns. Pick the one that matches the test you're writing.
Pattern A — hand-rolled inline sequence. The driver's run() function spells out every transaction in C++ code:
void run() {
wait(20, SC_NS);
drive(0x10, 0xAA);
wait(50, SC_NS);
drive(0x20, 0xBB);
wait(50, SC_NS);
drive(0x30, 0xCC);
sc_stop();
}
Right for: small focused tests, regression tests for specific bugs, examples in documentation.
Pattern B — table-driven sequence. The driver reads from a std::vector of test vectors:
struct vec { sc_uint<8> addr; sc_uint<8> data; sc_time delay; };
std::vector<vec> sequence = {
{0x10, 0xAA, sc_time(50, SC_NS)},
{0x20, 0xBB, sc_time(50, SC_NS)},
{0x30, 0xCC, sc_time(100, SC_NS)},
};
void run() {
wait(20, SC_NS);
for (auto& v : sequence) {
drive(v.addr, v.data);
wait(v.delay);
}
sc_stop();
}
Right for: parameterised tests, easy expansion of test sets, separating "what to test" from "how to test".
Pattern C — file-fed sequence. The driver reads from a file (or stdin) at simulation start:
void run() {
std::ifstream in("stimulus.txt");
if (!in) { SC_REPORT_ERROR("driver", "stimulus.txt not found"); }
wait(20, SC_NS);
unsigned a, d;
while (in >> std::hex >> a >> d) {
drive(static_cast<sc_uint<8>>(a), static_cast<sc_uint<8>>(d));
wait(50, SC_NS);
}
sc_stop();
}
Right for: regression suites with many input files, fuzz testing with externally generated inputs, replaying a previously captured trace.
The patterns nest naturally: a hand-rolled driver might call into a table-driven helper which might in turn be initialized from a file. Pick the simplest pattern that meets your need.
A note on SC_REPORT_INFO / SC_REPORT_WARNING / SC_REPORT_ERROR
Instead of std::cout, kernel-level reporting via SC_REPORT_INFO("tag", "message") integrates with the SystemC report system:
SC_REPORT_INFO("driver", "starting stimulus sequence");
SC_REPORT_WARNING("monitor", "saw unexpected value but continuing");
SC_REPORT_ERROR("scoreboard","mismatch detected — aborting");
The report system has configurable severity-to-action mappings (sc_report_handler::set_actions(SC_ERROR, SC_DISPLAY | SC_LOG | SC_STOP)). Errors can be configured to abort, log to a file, or display and continue. The unit-of-output is tagged with the verbosity level, the tag string, the file/line, and the message. For testbenches that may run in CI, prefer SC_REPORT_* over std::cout because the report system gives you per-tag filtering and severity-based termination control out of the box.
Reusable testbench skeleton
A pattern worth copying for any new testbench you write — sc_main skeleton with driver/monitor/scoreboard and VCD tracing already wired:
// file: testbench_skeleton.cpp
#include <systemc.h>
// --- DUT goes here ---
SC_MODULE(my_dut) {
// ports + body
SC_CTOR(my_dut) { /* register processes */ }
};
// --- Driver, Monitor, Scoreboard skeletons ---
SC_MODULE(driver) {
sc_in<bool> clk;
void run() {
wait(10, SC_NS);
// stimulus here
sc_stop();
}
SC_CTOR(driver) { SC_THREAD(run); }
};
SC_MODULE(monitor) {
sc_in<bool> clk;
void sample() { /* observe + record */ }
SC_CTOR(monitor) {
SC_METHOD(sample);
sensitive << clk.pos();
dont_initialize();
}
};
SC_MODULE(scoreboard) {
sc_in<bool> clk;
int passes = 0, failures = 0;
void check() { /* compare */ }
SC_CTOR(scoreboard) {
SC_METHOD(check);
sensitive << clk.pos();
dont_initialize();
}
~scoreboard() {
std::cout << "PASS=" << passes << " FAIL=" << failures << "\n";
}
};
int sc_main(int, char**) {
sc_clock clk("clk", 10, SC_NS);
// signals
// instantiate dut, driver, monitor, scoreboard
// bind all ports
sc_trace_file* tf = sc_create_vcd_trace_file("waves");
tf->set_time_unit(1, SC_NS);
sc_trace(tf, clk, "clk");
// sc_trace each signal of interest
sc_start(sc_time(10, SC_MS)); // watchdog upper bound
sc_close_vcd_trace_file(tf);
return /* scoreboard exit code */ 0;
}
This is a 60-line starting point. Most production testbenches grow from exactly this skeleton.
End-to-end example: a complete FIFO testbench
To make all the patterns concrete, here is a complete testbench for a 4-entry synchronous FIFO. This is what the hands-on exercise builds toward; here it is fully worked.
// file: fifo_tb.cpp
#include <systemc.h>
#include <queue>
// --- DUT: a 4-entry synchronous FIFO ---
SC_MODULE(fifo4) {
sc_in<bool> clk;
sc_in<bool> rst_n;
sc_in<sc_uint<8>> wr_data;
sc_in<bool> push;
sc_out<sc_uint<8>> rd_data;
sc_in<bool> pop;
sc_out<bool> full;
sc_out<bool> empty;
sc_uint<8> mem[4];
int head = 0, tail = 0, count = 0;
void tick() {
if (!rst_n.read()) {
head = tail = count = 0;
full.write(false);
empty.write(true);
rd_data.write(0);
return;
}
bool do_push = push.read() && count < 4;
bool do_pop = pop.read() && count > 0;
if (do_push) {
mem[tail] = wr_data.read();
tail = (tail + 1) % 4;
count++;
}
if (do_pop) {
head = (head + 1) % 4;
count--;
}
rd_data.write(count > 0 ? mem[head] : sc_uint<8>(0));
full.write(count >= 4);
empty.write(count == 0);
}
SC_CTOR(fifo4) {
SC_METHOD(tick);
sensitive << clk.pos();
dont_initialize();
}
};
// --- Driver: fill then drain, drive push and pop appropriately ---
SC_MODULE(driver) {
sc_in<bool> clk;
sc_in<bool> full;
sc_in<bool> empty;
sc_out<bool> rst_n;
sc_out<sc_uint<8>> wr_data;
sc_out<bool> push;
sc_out<bool> pop;
std::queue<sc_uint<8>>* expected;
void run() {
// Reset
rst_n.write(false);
push.write(false);
pop.write(false);
wr_data.write(0);
wait(2, SC_NS);
rst_n.write(true);
wait(clk.posedge_event());
// Fill the FIFO with 4 values
for (sc_uint<8> v = 10; v <= 40; v += 10) {
wait(clk.posedge_event());
wr_data.write(v);
push.write(true);
expected->push(v);
std::cout << "[" << sc_time_stamp()
<< "] driver: pushing " << v << "\n";
}
wait(clk.posedge_event());
push.write(false);
// Verify FIFO is full
wait(SC_ZERO_TIME);
if (!full.read()) {
SC_REPORT_ERROR("driver", "FIFO should be full after 4 pushes");
}
// Drain the FIFO
for (int i = 0; i < 4; ++i) {
wait(clk.posedge_event());
pop.write(true);
}
wait(clk.posedge_event());
pop.write(false);
// Verify FIFO is empty
wait(clk.posedge_event());
if (!empty.read()) {
SC_REPORT_ERROR("driver", "FIFO should be empty after 4 pops");
}
wait(10, SC_NS);
sc_stop();
}
SC_CTOR(driver) { SC_THREAD(run); }
};
// --- Monitor: capture each popped value ---
SC_MODULE(monitor) {
sc_in<bool> clk;
sc_in<bool> pop;
sc_in<sc_uint<8>> rd_data;
std::queue<sc_uint<8>> observed;
void sample() {
// The DUT writes rd_data BEFORE we tick; on the cycle pop is asserted,
// rd_data is the value being popped.
if (pop.read()) {
observed.push(rd_data.read());
std::cout << "[" << sc_time_stamp()
<< "] monitor: observed pop of " << rd_data.read() << "\n";
}
}
SC_CTOR(monitor) {
SC_METHOD(sample);
sensitive << clk.pos();
dont_initialize();
}
};
// --- Scoreboard: compare expected vs observed ---
SC_MODULE(scoreboard) {
std::queue<sc_uint<8>> expected_queue;
monitor* mon;
int passes = 0, failures = 0;
void final_check() {
while (!expected_queue.empty() && !mon->observed.empty()) {
sc_uint<8> e = expected_queue.front();
sc_uint<8> o = mon->observed.front();
expected_queue.pop();
mon->observed.pop();
if (o == e) ++passes;
else {
std::cout << "[scoreboard] MISMATCH: expected " << e
<< " got " << o << "\n";
++failures;
}
}
if (!expected_queue.empty() || !mon->observed.empty()) {
++failures;
std::cout << "[scoreboard] count mismatch: "
<< expected_queue.size() << " expected, "
<< mon->observed.size() << " observed remaining\n";
}
}
SC_CTOR(scoreboard) {}
~scoreboard() {
final_check();
std::cout << "\n========== TEST SUMMARY ==========\n";
std::cout << "Passed: " << passes << ", Failed: " << failures << "\n";
std::cout << "===================================\n";
}
};
int sc_main(int, char**) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst_n;
sc_signal<sc_uint<8>> wr_data, rd_data;
sc_signal<bool> push, pop, full, empty;
fifo4 dut("dut");
driver drv("drv");
monitor mon("mon");
scoreboard sb("sb");
dut.clk(clk); dut.rst_n(rst_n);
dut.wr_data(wr_data); dut.push(push);
dut.rd_data(rd_data); dut.pop(pop);
dut.full(full); dut.empty(empty);
drv.clk(clk); drv.full(full); drv.empty(empty);
drv.rst_n(rst_n); drv.wr_data(wr_data); drv.push(push); drv.pop(pop);
drv.expected = &sb.expected_queue;
mon.clk(clk); mon.pop(pop); mon.rd_data(rd_data);
sb.mon = &mon;
sc_trace_file* tf = sc_create_vcd_trace_file("fifo_waves");
tf->set_time_unit(1, SC_NS);
sc_trace(tf, clk, "clk");
sc_trace(tf, rst_n, "rst_n");
sc_trace(tf, wr_data, "wr_data");
sc_trace(tf, rd_data, "rd_data");
sc_trace(tf, push, "push");
sc_trace(tf, pop, "pop");
sc_trace(tf, full, "full");
sc_trace(tf, empty, "empty");
sc_start(sc_time(1, SC_US)); // watchdog upper bound
sc_close_vcd_trace_file(tf);
return sb.failures > 0 ? 1 : 0;
}
Expected outcome: 4 values pushed (10, 20, 30, 40), 4 values popped (10, 20, 30, 40), scoreboard reports Passed: 4, Failed: 0. The VCD file fifo_waves.vcd lets you inspect the full/empty transitions, the push/pop ordering, and the data flow visually.
Three things worth noticing in this end-to-end example. First, the driver knows the expected sequence and pushes onto the scoreboard's queue as it drives — the scoreboard does not have to know the test's structure. Second, the monitor captures every observed pop without filtering — the scoreboard does the comparison in its destructor (after the simulation has fully terminated). Third, the scoreboard's final_check is robust to "more expected than observed" or "more observed than expected" — both are failures, but with different error messages, making debugging easier.
This pattern scales. Add more test phases to the driver, more scoreboard checks, more monitors for different DUT outputs — the structure stays the same. Every UVM-SystemC verification environment is ultimately built on top of this same skeleton, with more layers of abstraction.
Advanced: Edge Cases & LRM Corners
Corner 1: sc_main is main, with a wrinkle
The SystemC library provides a main that performs kernel initialization and then calls your sc_main. Per IEEE 1666-2011 §4.3, sc_main is the user-supplied entry point: int sc_main(int argc, char* argv[]). The library-provided main parses any SystemC-specific command-line flags (none in the standard, but Accellera reserves the namespace), strips them, and forwards the rest to sc_main.
A consequence: if you link your own main (e.g., from a non-SystemC framework), you must rename it or link against the variant of SystemC that lets you call its initialization manually. The standard library form requires sc_main and nothing else.
Corner 2: sc_get_status() and the simulation lifecycle
sc_get_status() returns one of:
SC_ELABORATION—sc_mainis constructing modules;sc_starthas not been called.SC_BEFORE_END_OF_ELABORATION—before_end_of_elaboration()callbacks are running.SC_END_OF_ELABORATION—end_of_elaboration()callbacks are running.SC_START_OF_SIMULATION—start_of_simulation()callbacks are running.SC_RUNNING—sc_startis active; processes are being scheduled.SC_PAUSED— betweensc_startcalls, with the simulation paused but not stopped.SC_STOPPED—sc_stop()has been called, kernel is shutting down.SC_END_OF_SIMULATION—end_of_simulation()callbacks are running.
The lifecycle callbacks (before_end_of_elaboration, end_of_elaboration, start_of_simulation, end_of_simulation) are overridable virtual functions on sc_module. They are the right place for late-bound initialization (e.g., wiring sockets dynamically) or for end-of-test reporting that needs access to the running simulation state.
Corner 3: Tracing time alignment and set_time_unit
VCD files store all timestamps as integer counts of a time unit declared in the file header. tf->set_time_unit(1, SC_NS) sets that unit to 1 ns. If the kernel's time resolution is finer than the trace's time unit (e.g., kernel at 1 ps, trace at 1 ns), the kernel's sc_time_stamp values are rounded down to the trace's resolution when written. Events that happened 500 ps apart at the kernel level may appear simultaneous in the VCD.
Choose the trace time unit to be at least as fine as the smallest interesting event in your simulation. For typical RTL-level work, 1 ps is the right trace unit if your kernel runs at 1 ps; 1 ns is right for purely architectural models that don't care about sub-ns timing.
Corner 4: Multiple sc_start calls and the trace file
You can call sc_start(t) multiple times between sc_create_vcd_trace_file and sc_close_vcd_trace_file. The trace accumulates across all calls. This is occasionally useful for "trace only this specific window" patterns:
sc_start(sc_time(100, SC_NS)); // not traced — no trace file open
sc_trace_file* tf = sc_create_vcd_trace_file("interesting_window");
tf->set_time_unit(1, SC_NS);
sc_trace(tf, signal_of_interest, "signal");
sc_start(sc_time(50, SC_NS)); // traced
sc_close_vcd_trace_file(tf);
sc_start(); // not traced again
The VCD file timestamps continue from wherever the trace was opened — so the file shows times 100 ns to 150 ns, not 0 ns to 50 ns.
Corner 5: User-defined types in sc_trace
For struct types, you must overload sc_trace:
struct transaction {
sc_uint<8> addr;
sc_uint<32> data;
bool write;
};
// Required overload:
void sc_trace(sc_trace_file* tf, const transaction& t,
const std::string& name) {
sc_trace(tf, t.addr, name + ".addr");
sc_trace(tf, t.data, name + ".data");
sc_trace(tf, t.write, name + ".write");
}
// Now sc_trace(tf, my_signal, "my_sig") works where my_signal is sc_signal<transaction>.
The function signature must be void sc_trace(sc_trace_file*, const T&, const std::string&) — the kernel finds it by argument-dependent lookup. Without the overload, you get a compile error mentioning incomplete type or no matching function.
Corner 6: The "phantom first sample" with dont_initialize and sc_trace
VCD traces capture the initial value of every traced signal at time 0 in the file header. If a signal's first real change happens at, say, 10 ns, the VCD shows it as constant from 0 to 10 ns then changing — fine.
But if you have a dont_initialize()'d monitor that never fires until 10 ns, and you also have sc_trace on the signal, the trace shows the signal at its default-constructed value (typically 0) for the entire 0–10 ns window — even if no process has actually evaluated it. This is correct (the signal's value really is 0 during that window per the default-construction rule), but it can confuse readers who expect the trace to only show "computed" values.
The takeaway: the trace shows the simulator's state, not the testbench's observation. Signals have values even when no process reads them.
Corner 7: sc_stop during elaboration
Calling sc_stop() before sc_start has any effect is technically legal but useless — it sets a flag that the kernel will check on entering the simulation loop, then immediately exits without running any processes. The pattern sc_stop() is sometimes used in elaboration callbacks to abort a test if some structural check fails (e.g., "this configuration is invalid; abort"). It's an unusual pattern; usually you'd just throw an exception or return non-zero from a callback.
Corner 8: sc_pause() and resumption from elsewhere
sc_pause() (when supported — added in 2.3.0) is the gentler sibling of sc_stop. It tells the kernel "finish the current delta cycle, then return from sc_start without terminating the simulation." After sc_pause, you can call sc_start again to resume. The simulation lifecycle moves to SC_PAUSED, not SC_STOPPED.
This is the right pattern when you want a co-simulation or interactive REPL style where sc_main (or some external driver) controls when the kernel runs and when it yields. Example:
SC_MODULE(pauseable_driver) {
sc_in<bool> clk;
void run() {
for (int i = 0; i < 100; ++i) {
wait(clk.posedge_event());
if (i == 50) {
std::cout << "[" << sc_time_stamp() << "] pausing at iteration 50\n";
sc_pause(); // returns control to sc_main
}
}
sc_stop();
}
SC_CTOR(pauseable_driver) { SC_THREAD(run); }
};
int sc_main(int, char**) {
sc_clock clk("clk", 10, SC_NS);
pauseable_driver d("d");
d.clk(clk);
sc_start(); // runs until sc_pause()
std::cout << "sc_main: simulation paused at " << sc_time_stamp() << "\n";
std::cout << "sc_main: doing host-side work...\n";
int host_result = some_host_computation(); // arbitrary host logic
std::cout << "sc_main: resuming\n";
sc_start(); // resumes until sc_stop()
return 0;
}
Most testbenches don't need sc_pause. It's the right tool for interactive debuggers, scripted simulations driven from Python via DPI/embedding, and co-simulation with non-SystemC components that need to step the simulator forward in chunks.
Corner 9: Watchdog timeout patterns done right
Pattern C from the Intermediate section uses sc_start(sc_time(10, SC_MS)) as a watchdog. The trap: if sc_start returns because the driver called sc_stop(), and the time limit was also reached, the simulation has terminated normally — but sc_get_status() may report SC_STOPPED either way. The reliable check is to track whether the driver explicitly signalled completion:
SC_MODULE(driver) {
bool finished_cleanly = false;
void run() {
for (int i = 0; i < 10; ++i) {
wait(clk.posedge_event());
data.write(i);
}
finished_cleanly = true;
sc_stop();
}
};
int sc_main(int, char**) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<sc_uint<8>> data;
driver d("d");
d.clk(clk); d.data(data);
sc_start(sc_time(10, SC_MS)); // watchdog
if (!d.finished_cleanly) {
SC_REPORT_ERROR("main", "watchdog timeout — driver did not complete");
return 2;
}
return scoreboard.failures > 0 ? 1 : 0;
}
The finished_cleanly flag is set before sc_stop() so even if the driver hangs partway through, sc_main can tell the difference between "test passed" and "test timed out."
Corner 10: The argc/argv parameters to sc_main
sc_main(int argc, char** argv) receives the command-line arguments after the SystemC library has stripped any of its own reserved flags. This lets your testbench accept its own command-line arguments — test selection, log verbosity, random seeds, file paths:
int sc_main(int argc, char** argv) {
std::string stimulus_file = "default_stimulus.txt";
int random_seed = 42;
bool enable_trace = false;
for (int i = 1; i < argc; ++i) {
std::string arg(argv[i]);
if (arg == "--stim" && i + 1 < argc) stimulus_file = argv[++i];
else if (arg == "--seed" && i + 1 < argc) random_seed = std::atoi(argv[++i]);
else if (arg == "--trace") enable_trace = true;
}
std::srand(random_seed);
if (enable_trace) std::cout << "tracing enabled\n";
// proceed to construct modules and call sc_start using stimulus_file and random_seed
return 0;
}
For complex argument parsing, link in a small library like cxxopts or CLI11. The simple loop above is enough for testbench-level needs.
Corner 11: Random number generation — rand() vs <random> vs SCV
C's std::rand() is fine for trivial tests, but its quality is poor and its state is global (no per-test isolation). Modern C++ <random> is the right default for new testbenches:
#include <random>
SC_MODULE(random_driver) {
sc_out<sc_uint<8>> data;
std::mt19937 rng; // Mersenne Twister
std::uniform_int_distribution<unsigned> dist;
random_driver(sc_module_name n, unsigned seed)
: sc_module(n), rng(seed), dist(0, 255) {
SC_HAS_PROCESS(random_driver);
SC_THREAD(run);
}
void run() {
for (int i = 0; i < 100; ++i) {
wait(10, SC_NS);
sc_uint<8> v = static_cast<sc_uint<8>>(dist(rng));
data.write(v);
}
}
};
The full SCV (SystemC Verification) library adds constrained-random capabilities (scv_constraint, scv_smart_ptr<T>) for cases where you need declarative randomization with constraints. SCV is a separate library (-lscv) and is beyond the Foundations section; it's covered properly in the UVM-SystemC Verification section (Section 4 of this series).
Corner 12a: The SC_REPORT_* system in depth
The SC_REPORT_* macros are not just convenience wrappers around std::cout. They participate in a structured reporting system with four severity levels (SC_INFO, SC_WARNING, SC_ERROR, SC_FATAL), per-severity action masks, per-tag filtering, and verbosity levels.
The action mask determines what happens when a report fires:
| Action | What happens |
|---|---|
SC_UNSPECIFIED |
Use default for this severity |
SC_DO_NOTHING |
Silent — useful to suppress noise during specific phases |
SC_THROW |
Throw an sc_report exception (caller handles) |
SC_LOG |
Append to the report log file |
SC_DISPLAY |
Print to stderr |
SC_CACHE_REPORT |
Buffer for later retrieval |
SC_INTERRUPT |
Call sc_interrupt_here() — useful for debuggers |
SC_STOP |
Call sc_stop() |
SC_ABORT |
Call C abort() — immediate termination, no cleanup |
You combine actions with bitwise OR: sc_report_handler::set_actions(SC_ERROR, SC_DISPLAY | SC_LOG | SC_STOP) means "when any SC_ERROR fires, display it, log it, and stop the simulation cleanly."
Defaults are reasonable: SC_INFO → SC_DISPLAY, SC_WARNING → SC_DISPLAY, SC_ERROR → SC_DISPLAY | SC_CACHE_REPORT | SC_THROW, SC_FATAL → SC_DISPLAY | SC_CACHE_REPORT | SC_ABORT. The defaults mean an SC_REPORT_ERROR in your testbench throws an exception and continues unless you catch — which is typically NOT what you want; configure it to SC_STOP instead so a single error terminates the test cleanly.
Per-tag filtering lets you silence specific tags without affecting others:
sc_report_handler::set_actions(
"noisy_module", // tag
SC_INFO, // severity
SC_DO_NOTHING); // suppress this tag at this severity
Per-verbosity filtering: SC_REPORT_INFO_VERB("tag", "message", SC_DEBUG) only fires if the global verbosity is at SC_DEBUG or higher. Set with sc_report_handler::set_verbosity_level(SC_MEDIUM). The levels are SC_NONE, SC_LOW, SC_MEDIUM, SC_HIGH, SC_FULL, SC_DEBUG.
For any non-trivial testbench, learning the report system pays off within a week of writing tests — it makes failure isolation, log analysis, and CI integration substantially easier.
Corner 12: Logging output to a file vs stdout
For long simulations, std::cout to a terminal can become a bottleneck (terminal scroll, redraw). Three patterns help:
- Redirect at the OS level:
./my_sim > out.log 2>&1. Simplest, no code change. - Use
SC_REPORT_*with file-logging actions:sc_report_handler::set_log_file_name("sim.log");and configureSC_INFOto useSC_LOG. - Open a
std::ofstreaminsc_mainand pass references around: more code, but lets modules write to different files for different concerns (driver log vs monitor log).
For most CI-friendly testbenches, option (1) is fine. Option (2) is the right answer when you want per-tag verbosity control.
Hands-on exercise
Build a bidirectional FIFO testbench:
- DUT: a 4-deep FIFO with
sc_in<sc_uint<8>> wr_data,sc_in<bool> push,sc_out<sc_uint<8>> rd_data,sc_in<bool> pop,sc_out<bool> full,sc_out<bool> empty. - Driver: an
SC_THREADthat fills the FIFO, then drains it, watchingfullandempty. - Monitor: an
SC_METHODonclk.pos()that records what comes out. - Scoreboard: an in-order queue of expected values, compared against what the monitor records on each pop.
Constraints:
- Predict before running: how many test vectors pass through? What does
sc_mainreturn? - Add VCD tracing to all signals. Open the resulting VCD in GTKWave and verify the FIFO depth visually.
- Include a watchdog (
sc_start(sc_time(1, SC_US))as upper bound) so a deadlock terminates cleanly.
Hint: you do not need to implement the FIFO yourself for the exercise — write a stub DUT that does the simplest possible thing (e.g., always reports empty=false, full=false, rd_data=wr_data) and exercise the testbench against the stub. The point is the testbench architecture, not the FIFO logic.
Common mistakes
- Forgetting
dont_initialize()on a monitor. Without it, the monitor's first invocation fires at simulation start with the default-constructed values of all observed signals, polluting the trace with a meaningless initial sample. - Calling
sc_start()without a time argument or asc_stop()mechanism. The kernel runs forever. CI jobs time out. Always have one of: a finitesc_start(t), an internalsc_stop()call, or both. - Forgetting to close the VCD trace file. Some kernels buffer trace data and the file is incomplete or empty if not closed. Always pair
sc_create_vcd_trace_filewithsc_close_vcd_trace_file. - Comparing observed values inside the monitor instead of in a separate scoreboard. Works for trivial tests but couples observation to checking, making it harder to reuse the monitor across tests. Prefer separation.
- Sharing a
std::queuebetween modules via raw pointers. Works for a toy example like the scoreboard above, but creates lifetime issues in larger testbenches. Prefer passing a reference to a single owner or using a dedicated channel. - Driving the DUT's input from
sc_mainand from a driver module simultaneously. Two drivers on a singlesc_signalis the multi-writer violation from Part 1 Corner 2. Pick one driver per signal. - Tracing too much. Every signal you trace adds to the VCD file size and slows simulation slightly. For long CI runs, trace nothing by default; enable trace for the failing tests only.
- Forgetting that
sc_mainreturns int. Returning nothing (relying on the implicit return) gives undefined behaviour. Return 0 on success, non-zero on failure — CI systems and test harnesses depend on this. - Using
SC_REPORT_ERRORand expecting the test to terminate cleanly. The default action forSC_ERRORisSC_THROW, which raises an exception. Unless you catch it insc_main, it propagates and terminates the program ungracefully (no scoreboard summary, no VCD close). Configure withsc_report_handler::set_actions(SC_ERROR, SC_DISPLAY | SC_STOP)at the top ofsc_mainto get clean termination instead. - Reading a signal in a monitor before the DUT has had a chance to write it. If the monitor's
SC_METHODfires in the same delta as a DUT write, the monitor might see the current value (pre-update) rather than the post-write value. The fix is to make the monitor sensitive to a different (later) event, or to add await(SC_ZERO_TIME)if it's a thread-style monitor. - Driving stimulus and observing in the same
sc_mainwithoutsc_startcalls between them. Writes fromsc_mainare pending until the nextsc_startruns the kernel. Callingsa.write(v); int got = sy.read();immediately insc_mainwill read the OLD value ofsy, not what the DUT will compute from the newsa. Insertsc_start(1, SC_NS)(or evensc_start(SC_ZERO_TIME)) between write and read.
Recap
After this post, you can:
- Write a testbench from
sc_mainalone for a small DUT, complete with stimulus, observation, and inline comparison. - Decompose a non-trivial testbench into driver, monitor, and scoreboard modules with clear responsibilities.
- Add VCD tracing in ten lines and view the results in GTKWave.
- Use
sc_stop()from inside a driver thread to terminate a simulation cleanly when the test sequence is done. - Use a time-bounded
sc_start(t)as a watchdog to prevent runaway simulations. - Return a meaningful exit code from
sc_mainso CI systems can detect failures. - Overload
sc_tracefor user-defined struct types. - Recognise the simulation lifecycle phases via
sc_get_status()and know when each lifecycle callback fires. - Build a watchdog'd, self-checking testbench combining all the above.
Further reading
- IEEE 1666-2011, §4.4 (sc_start, sc_stop), §5.18 (tracing). The tracing section is short and worth reading end-to-end.
- Accellera SystemC User's Guide — the testbench and tracing chapters. Plain-language treatment with worked examples.
- Accellera SystemC PoC source:
src/sysc/tracing/sc_vcd_trace.h(VCD writer implementation),src/sysc/kernel/sc_simcontext.h(simulation lifecycle). - Doulos Golden Reference Guide — the testbench chapter and the tracing chapter.
- GTKWave manual — for any non-trivial waveform-based debugging.
- Black, Donovan, Tahar, SystemC: From the Ground Up (2nd ed.) — the testbench chapter walks through the same pattern progression (single-file → driver/monitor → scoreboard) at greater length, with discussion of when each pattern starts to break down and motivate the next layer of abstraction.
- The Accellera SystemC reference distribution includes example testbenches under
examples/sysc/— particularlyexamples/sysc/fifo/,examples/sysc/simple_perf/, andexamples/sysc/2.1/. These are short, complete, and demonstrate the patterns described here in the simulator authors' own idiom.
Next in this section
→ Part 7: Elaboration & the Kernel Lifecycle — the final post in Foundations, covering the simulation lifecycle from sc_main entry through end_of_simulation(), the four override hooks (before_end_of_elaboration, end_of_elaboration, start_of_simulation, end_of_simulation) for late-bound initialization, and the kernel-level debugging techniques that let you trace what the simulator is actually doing at each phase boundary. After Part 7 the Foundations section is complete and you are ready for Section 2 (RTL-Level Modeling Patterns).
Comments (0)
Leave a Comment