14. SystemC Tutorial - Structural Composition & Hierarchy

Why this matters

Rewritten 2026-06-05 with deeper first-principles material.

Every interesting chip is bigger than one module. The ALU you built earlier is a few hundred lines; the CPU that contains it is a few thousand; the SoC that contains the CPU is a few million. None of that scale is achieved by writing one enormous module. It is achieved by composition — building small modules that each do one thing, then wiring them together into larger modules, then wiring those into still larger modules, until the top of the tree is the whole system. The skill of taking verified leaf blocks and assembling them into a working datapath-plus-control machine is the skill that turns a pile of components into a CPU. It is also, by a wide margin, where the most bugs live.

The reason is structural, not intellectual. Each leaf module, tested in isolation, behaves perfectly: the ALU adds, the register file stores, the decoder decodes. The defects appear at the seams — at the wires between instances. A signal that you named one thing in your head and another thing in the port declaration. An output port bound to the wrong input port of the same type, so the types match, elaboration succeeds, and the simulation runs and produces quietly wrong answers. A control signal that is active-high in the module that produces it and assumed active-low in the module that consumes it. A clock that you bound to five of your six leaf modules and forgot on the sixth, so that one block sits frozen while the rest of the design runs. A second driver accidentally bound to a signal that already has one, so the value that arrives downstream depends on which process the kernel happened to run last. Every one of those is a composition bug, invisible inside any single module, and every one of them is the daily bread of integration and verification engineers.

This post teaches structural composition in SystemC from first principles, so that you can assemble modules into hierarchies deliberately rather than by trial and error. We start with the smallest composition that is still interesting — two leaf modules, a counter and a comparator, wired together at a parent through a single signal — and we use it to extract every rule that matters: how a child module is declared and constructed, how a port is bound to a channel, who owns the wire between two instances, how the clock reaches a leaf, how the kernel checks bindings at elaboration and what its error messages mean, and how the hierarchical name of an object helps you debug. Only after those rules are concrete do we scale them up to the worked example: the complete single-cycle RV32I CPU, six leaf modules and four glue processes, wired into one Rv32iCpu module and smoke-tested with a three-instruction program. By the end you will be able to compose modules into a working machine, read an elaboration-time binding error and fix it on sight, and recognize the three composition bugs — unbound port, double-driven signal, unclocked leaf — that account for most integration failures. That is the bar.

Prerequisites

  • Part 1 — Modules, Ports & Signals. You must be comfortable declaring SC_MODULE, declaring sc_in / sc_out ports, declaring sc_signal channels, and writing an SC_CTOR. This post is entirely about connecting those pieces together, so the pieces themselves must be second nature.
  • Part 4 — Processes & Sensitivity (SC_METHOD vs SC_THREAD). The glue logic that lives between instances — muxes, next-PC selection — is written as SC_METHOD processes with explicit sensitivity lists. You need to know why a combinational glue process lists every signal it reads.
  • Part 8 — Combinational Logic / the RV32I ALU. The ALU is one of the six leaf modules we compose. You should remember its port list (operands, op-select, result, flags).
  • Part 10 — Sequential Logic / the Program Counter. The PC is the clocked heart of the fetch path; composition is where its next_pc input finally gets driven by real branch logic.
  • Part 12 — Memories & the Register File. The 32×32 register file is the storage leaf the datapath reads and writes.
  • SystemC 2.3.x installed. Every example compiles with a C++17 compiler (g++ 9+ or clang++ 10+) against any 2.3.x build.

Three ideas from earlier parts get heavy use here and are worth re-fixing in your mind before continuing: a module is an object with a name (Part 1), a port is bound to a channel by calling it like a function (Part 1), and an sc_signal is a wire whose value updates one delta after a write (the sequential-logic parts). Composition is nothing more than arranging many of each and connecting them correctly.

Mental model (first principles)

A composed SystemC design is a tree of modules. The leaves of the tree are the modules that contain actual behavior — processes that compute, registers that store. The internal nodes of the tree are container modules: they contain other modules and almost no behavior of their own beyond the wiring (and a little glue logic). The root of the tree is the top-level module of your design. Above the root sits sc_main, the ordinary C++ function where you create the top instance and call sc_start.

Three facts define how the tree is built, and everything else in this post follows from them.

Fact 1 — a child is a member of its parent. To put module B inside module A, you declare an instance of B as a data member of A, and you construct that member (giving it a name) in A's constructor. The instance lives exactly as long as the A that contains it. This is why a child instance can never be a local variable inside the constructor: a local variable is destroyed the instant the constructor returns, and the child would vanish before simulation even starts, taking its port bindings down with it. Member, not local. That single rule prevents a whole category of crashes.

Fact 2 — the parent owns the wires. When two children need to be connected — B's output feeding C's input — the wire between them is an sc_signal, and that signal is a member of the parent A, not of B or C. Neither child knows the other exists. The parent declares the signal, then binds B's output port to it and C's input port to it. The signal carries B's value to C exactly as a physical wire would, with the one-delta update semantics you already know from sequential modeling. Ownership flows downward: the parent reaches into each child and hands it the channel its port should see. A child never reaches up.

Fact 3 — binding names both endpoints. You connect a port to a channel by calling the port like a function: B.out_port(the_signal). This is the overloaded operator() on the port; B.out_port.bind(the_signal) is the identical, more explicit spelling. Both the port (B.out_port) and the channel (the_signal) are named in the call, so the order in which you write your binding statements does not matter to correctness — you can bind clocks first, or outputs first, or in declaration order; the connection graph is the same. What does matter is that you name the right pair. Bind the right port to the wrong (but same-typed) signal and SystemC will not complain: types match, elaboration succeeds, and your design simulates the wrong behavior silently. The type system is too permissive to catch a semantic miswire; only a testbench can.

Here is the smallest non-trivial instance of the tree. A Counter leaf and a Comparator leaf, composed at a CounterCompareTop parent that owns the single wire between them:

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
flowchart LR
    CLK([clk]) --> CNT
    RST([rst]) --> CNT
    CNT["Counter"] -->|count| CMP["Comparator"]
    CMP -->|at_limit| OUT([at_limit])

Read the diagram as ownership: CounterCompareTop is the box around everything; it declares one sc_signal<unsigned> count; it binds Counter.count_out to that signal (Counter writes it) and Comparator.value_in to the same signal (Comparator reads it). The clock and reset come into the parent from sc_main and the parent routes the clock down to the counter. The comparator is purely combinational and needs no clock. That is the entire pattern — and it is exactly the pattern, scaled up, that wires six leaf modules into a CPU.

There is no separate "netlist" object in SystemC the way there is a structural netlist in a synthesized design. The netlist is the set of operator() calls in your constructors. When the constructors finish running, elaboration walks every port and checks that each is bound to exactly the kind of thing it requires. Unbound ports of a binding-required kind abort the run with error (E109) complete binding failed. This check happens at sc_start() — at runtime, after all your C++ has executed — which is the single biggest mental adjustment for engineers coming from SystemVerilog, where the same class of error is caught at compile/elaborate time before any simulation begins. In SystemC, the design is assembled by running code, and the assembly is validated only when you ask the kernel to start.

One more load-bearing idea before we build: datapath versus control. A useful composed machine almost always splits into two cooperating halves. The datapath is the part that moves and transforms data — registers, ALUs, muxes, memories. The control is the part that decides what the datapath does this cycle — typically one or more FSMs (Part 11) whose outputs are the datapath's select and enable signals. Composition is where the two halves meet: the control module's output ports bind to the same parent-owned signals that the datapath module's control-input ports bind to. We will see the smallest version of this — a one-bit enable from a control FSM gating a counter — and then the full version, where a decoder's control outputs steer an ALU, a register file, and a data memory.

Beginner: First Principles

We will now build the counter-and-comparator composition end to end. Two leaf modules, one parent that owns the wire between them, one testbench. Read every binding call and ask yourself "which port, which signal, who owns the signal" — that question is the whole of structural composition.

The Counter leaf is a clocked module with a synchronous, active-high reset. It holds an internal count and exposes it on an output port. It is identical in spirit to the registered counters from the sequential-logic part.

// file: counter_compare.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            counter_compare.cpp -o counter_compare -lsystemc
// Run:   LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./counter_compare

#include <systemc.h>
#include <iostream>

// ---- Leaf 1: a clocked counter (synchronous active-high reset) ----
SC_MODULE(Counter) {
  sc_in<bool>      clk;
  sc_in<bool>      rst;
  sc_out<unsigned> count_out;   // the current count, exposed as a wire

  sc_signal<unsigned> count;    // internal state register

  void step() {                 // clocked: runs on each rising edge
    if (rst.read()) count.write(0);
    else            count.write(count.read() + 1);
  }

  void drive_out() {            // combinational: mirror state to the port
    count_out.write(count.read());
  }

  SC_CTOR(Counter) {
    SC_METHOD(step);
    sensitive << clk.pos();
    dont_initialize();          // clocked process: never run before first edge

    SC_METHOD(drive_out);
    sensitive << count;         // re-drive the port whenever count changes
  }
};

// ---- Leaf 2: a purely combinational comparator (no clock needed) ----
SC_MODULE(Comparator) {
  sc_in<unsigned> value_in;
  sc_out<bool>    at_limit;

  static constexpr unsigned LIMIT = 4;

  void compare() {
    at_limit.write(value_in.read() >= LIMIT);
  }

  SC_CTOR(Comparator) {
    SC_METHOD(compare);
    sensitive << value_in;      // combinational: list every input it reads
  }
};

// ---- Parent: owns both children AND the wire between them ----
SC_MODULE(CounterCompareTop) {
  sc_in<bool>  clk;
  sc_in<bool>  rst;
  sc_out<bool> at_limit;        // forwarded up from the comparator

  Counter    i_cnt;             // child instances are MEMBERS (Fact 1)
  Comparator i_cmp;

  sc_signal<unsigned> count;    // the wire between the children (Fact 2)

  SC_CTOR(CounterCompareTop)
    : i_cnt("i_cnt"), i_cmp("i_cmp")   // construct children with names
  {
    // Route clock + reset DOWN into the counter leaf.
    i_cnt.clk(clk);
    i_cnt.rst(rst);

    // The counter WRITES the interconnect signal...
    i_cnt.count_out(count);

    // ...and the comparator READS the same interconnect signal.
    i_cmp.value_in(count);

    // Forward the comparator's result up to the parent's own port.
    i_cmp.at_limit(at_limit);
  }
};

// ---- Testbench: a clocked monitor that prints each cycle ----
SC_MODULE(Monitor) {
  sc_in<bool> clk;
  sc_in<bool> at_limit;

  void watch() {
    std::cout << "[" << sc_time_stamp() << "] at_limit="
              << at_limit.read() << "\n";
  }

  SC_CTOR(Monitor) {
    SC_METHOD(watch);
    sensitive << clk.pos();
    dont_initialize();
  }
};

int sc_main(int /*argc*/, char* /*argv*/[]) {
  sc_clock        clk("clk", 10, SC_NS);   // 10 ns period, 50% duty
  sc_signal<bool> rst;
  sc_signal<bool> at_limit;

  CounterCompareTop top("top");
  top.clk(clk);
  top.rst(rst);
  top.at_limit(at_limit);

  Monitor mon("mon");
  mon.clk(clk);
  mon.at_limit(at_limit);

  rst.write(true);          // hold reset high for one cycle
  sc_start(10, SC_NS);
  rst.write(false);
  sc_start(90, SC_NS);

  sc_stop();
  return 0;
}

Compile and run. Before reading the walkthrough, try to predict the output: the counter resets to 0 at the first edge, then increments each cycle; the comparator asserts at_limit once the count reaches 4.

Expected output:

[10 ns] at_limit=0
[20 ns] at_limit=0
[30 ns] at_limit=0
[40 ns] at_limit=0
[50 ns] at_limit=1
[60 ns] at_limit=1
[70 ns] at_limit=1
[80 ns] at_limit=1
[90 ns] at_limit=1
[100 ns] at_limit=1

Walk it. The clock period is 10 ns, so rising edges land at 10, 20, 30, … ns. Reset is high from t=0 to t=10 ns. At the first edge (t=10 ns), Counter::step runs while rst is high and writes count = 0. The update phase commits that; the value-changed event on the counter's internal count fires; Counter::drive_out wakes and writes count_out = 0; that propagates onto the parent-owned count signal; Comparator::compare wakes, reads 0, and writes at_limit = (0 >= 4) = false. The monitor, clocked, reads the previous value of at_limit (still its default false) and prints at_limit=0. At t=20 ns reset is low; step writes count = 0 + 1 = 1; the chain re-evaluates; at_limit stays false. The count climbs 1, 2, 3 across t=20, t=30, t=40 ns — all below the limit of 4 — so the monitor prints 0 through t=40 ns.

At t=50 ns the counter reaches 4. 1 >= 4 was false, 2 >= 4 false, 3 >= 4 false, but at the edge that makes count = 4, the comparator computes 4 >= 4 = true, at_limit goes high, and the monitor prints at_limit=1. From there the count keeps climbing and at_limit stays high. The composition works exactly as the two leaves, wired through one parent-owned signal, should make it work.

Now look at what made this a composition rather than one big module, because every line of the connective tissue teaches a rule.

The two child instances i_cnt and i_cmp are declared as members of CounterCompareTop, and they are constructed in the member-initializer list with their names: : i_cnt("i_cnt"), i_cmp("i_cmp"). Those names are not decoration. They become the children's hierarchical names — top.i_cnt and top.i_cmp — which is what every error message and waveform trace uses to point at a specific instance. Forget to name a child and SystemC gives it a generated or empty name, and your binding errors read (unnamed): port not bound, which is far harder to chase.

The interconnect sc_signal<unsigned> count is a member of the parent, not of either child. The counter's count_out port and the comparator's value_in port are both bound to it: i_cnt.count_out(count) and i_cmp.value_in(count). One signal, one writer (the counter's output port), one reader (the comparator's input port). The parent is the only module that knows both children; the children are blissfully ignorant of each other. That is the ownership boundary working as designed.

The clock and reset enter the parent through its own clk and rst ports and are routed down into the counter with i_cnt.clk(clk) and i_cnt.rst(rst). Note what is and is not clocked: the counter is clocked (it has state), so it gets the clock; the comparator is purely combinational, so it has no clk port and receives no clock. This is deliberate. Binding a clock to a module that does not need one is harmless clutter; forgetting to bind a clock to a module that does need one is one of the three classic composition bugs, and we will reproduce it in the Advanced section.

Two leaf-module conventions are worth pausing on because they recur in the CPU. First, Counter separates its clocked step process from its combinational drive_out process. It would be tempting to have step write count_out directly, but the cleaner idiom — one process owns the state register, a second mirrors the state onto the output port — keeps the registered semantics crisp and matches the two-process discipline from the FSM and sequential parts. Second, the comparator's compare process lists sensitive << value_in: every signal a combinational process reads must appear in its sensitivity list. SystemC has no always @(*) auto-sensitivity. Omit value_in and the comparator never re-evaluates when the count changes — the output freezes at its initial value and the bug looks like "my comparator is stuck."

A final structural observation. The whole composition is three modules deep: sc_main creates CounterCompareTop, which creates Counter and Comparator. That is a tree of height two. The single-cycle CPU we build later is the same shape with a wider middle: one top module containing six leaves and four glue processes. Nothing about the mechanics changes as the tree grows — more members, more signals, more binding calls, same three facts.

Intermediate: How It Really Works

The counter-and-comparator composition was a pure datapath: data flowed left to right, no decisions. Real machines have a control half that decides what the datapath does each cycle. The smallest version of that pattern is one control bit gating the counter, and building it teaches three things at once: how to add a third leaf to a composition, how control and datapath share parent-owned signals, and what "named binding is by object, not by position" really protects you from.

Adding a control leaf: datapath plus control

We extend the design with a tiny ControlFsm leaf — a two-state machine (RUN, HOLD) that toggles an enable output. When enable is low, the counter holds; when high, it counts. The counter gains an enable input port. The parent now owns two interconnect signals: count (datapath, counter → comparator) and enable (control, FSM → counter). This is the canonical shape of every controlled datapath, in miniature.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
flowchart LR
    CLK([clk]) --> FSM
    CLK --> CNT
    FSM["ControlFsm"] -->|enable| CNT["Counter"]
    CNT -->|count| CMP["Comparator"]
    CMP -->|at_limit| OUT([at_limit])
// file: controlled_counter.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            controlled_counter.cpp -o controlled_counter -lsystemc

#include <systemc.h>
#include <iostream>

// ---- Control leaf: a 2-state FSM that pulses enable RUN/HOLD ----
SC_MODULE(ControlFsm) {
  sc_in<bool>  clk;
  sc_in<bool>  rst;
  sc_out<bool> enable;

  enum { ST_RUN = 0, ST_HOLD = 1 };
  sc_signal<int> state, next_state;

  void state_register() {
    if (rst.read()) state.write(ST_RUN);
    else            state.write(next_state.read());
  }

  void next_and_out() {
    int s = state.read();
    // Alternate RUN/HOLD every cycle: count on even cycles only.
    next_state.write(s == ST_RUN ? ST_HOLD : ST_RUN);
    enable.write(s == ST_RUN);     // Moore output: function of state only
  }

  SC_CTOR(ControlFsm) {
    SC_METHOD(state_register);
    sensitive << clk.pos();
    dont_initialize();

    SC_METHOD(next_and_out);
    sensitive << state;
  }
};

// ---- Datapath leaf: counter now gated by enable ----
SC_MODULE(Counter) {
  sc_in<bool>      clk;
  sc_in<bool>      rst;
  sc_in<bool>      enable;
  sc_out<unsigned> count_out;

  sc_signal<unsigned> count;

  void step() {
    if (rst.read())          count.write(0);
    else if (enable.read())  count.write(count.read() + 1);
    // else: hold (no write needed; sc_signal keeps its value)
  }

  void drive_out() { count_out.write(count.read()); }

  SC_CTOR(Counter) {
    SC_METHOD(step);
    sensitive << clk.pos();
    dont_initialize();

    SC_METHOD(drive_out);
    sensitive << count;
  }
};

SC_MODULE(Comparator) {
  sc_in<unsigned> value_in;
  sc_out<bool>    at_limit;
  static constexpr unsigned LIMIT = 4;
  void compare() { at_limit.write(value_in.read() >= LIMIT); }
  SC_CTOR(Comparator) { SC_METHOD(compare); sensitive << value_in; }
};

// ---- Parent: three children, two interconnect signals ----
SC_MODULE(ControlledTop) {
  sc_in<bool>      clk;
  sc_in<bool>      rst;
  sc_out<bool>     at_limit;
  sc_out<unsigned> count_dbg;     // observation port for the trace

  ControlFsm i_fsm;
  Counter    i_cnt;
  Comparator i_cmp;

  sc_signal<bool>     enable;     // control wire: FSM -> counter
  sc_signal<unsigned> count;      // datapath wire: counter -> comparator

  SC_CTOR(ControlledTop)
    : i_fsm("i_fsm"), i_cnt("i_cnt"), i_cmp("i_cmp")
  {
    // Control leaf
    i_fsm.clk(clk);
    i_fsm.rst(rst);
    i_fsm.enable(enable);

    // Datapath leaf
    i_cnt.clk(clk);
    i_cnt.rst(rst);
    i_cnt.enable(enable);          // SAME signal the FSM drives
    i_cnt.count_out(count);

    // Combinational leaf
    i_cmp.value_in(count);
    i_cmp.at_limit(at_limit);

    // Observation
    SC_METHOD(forward_count);
    sensitive << count;
  }

  void forward_count() { count_dbg.write(count.read()); }
};

SC_MODULE(Monitor) {
  sc_in<bool>      clk;
  sc_in<unsigned> count_dbg;
  sc_in<bool>      at_limit;
  void watch() {
    std::cout << "[" << sc_time_stamp() << "] count="
              << count_dbg.read() << " at_limit="
              << at_limit.read() << "\n";
  }
  SC_CTOR(Monitor) {
    SC_METHOD(watch);
    sensitive << clk.pos();
    dont_initialize();
  }
};

int sc_main(int, char*[]) {
  sc_clock        clk("clk", 10, SC_NS);
  sc_signal<bool> rst, at_limit;
  sc_signal<unsigned> count_dbg;

  ControlledTop top("top");
  top.clk(clk); top.rst(rst);
  top.at_limit(at_limit); top.count_dbg(count_dbg);

  Monitor mon("mon");
  mon.clk(clk); mon.count_dbg(count_dbg); mon.at_limit(at_limit);

  rst.write(true);
  sc_start(10, SC_NS);
  rst.write(false);
  sc_start(140, SC_NS);

  sc_stop();
  return 0;
}

The FSM toggles enable every cycle, so the counter advances on roughly half the cycles. Trace it: reset releases at t=10 ns with the FSM in ST_RUN (enable high) and count at 0; the counter increments only when it samples enable high.

Expected output:

[10 ns] count=0 at_limit=0
[20 ns] count=0 at_limit=0
[30 ns] count=1 at_limit=0
[40 ns] count=1 at_limit=0
[50 ns] count=2 at_limit=0
[60 ns] count=2 at_limit=0
[70 ns] count=3 at_limit=0
[80 ns] count=3 at_limit=0
[90 ns] count=4 at_limit=1
[100 ns] count=4 at_limit=1
[110 ns] count=5 at_limit=1
[120 ns] count=5 at_limit=1
[130 ns] count=6 at_limit=1
[140 ns] count=6 at_limit=1
[150 ns] count=7 at_limit=1

The count advances every other cycle because the FSM's enable is high on alternate cycles and the counter, sampling enable registered through the parent-owned signal, increments only then. It crosses the limit of 4 at t=90 ns and at_limit latches high. The detail that matters for composition is not the arithmetic — it is that the FSM's enable output and the counter's enable input meet on one parent-owned sc_signal<bool> enable, exactly as the counter's count_out and the comparator's value_in meet on count. Control and datapath are wired the same way; the only difference is which half writes the signal and which half reads it.

Member instances versus pointer instances

So far every child has been a plain data member, constructed in the initializer list. That is the right default: it is the simplest, has automatic lifetime tied to the parent, and cannot leak. But two situations force the alternative — a pointer (or smart pointer) to a heap-allocated child:

  1. The number of children is not known until run time. If a parent contains N identical lanes and N comes from a constructor argument or a config file, you cannot declare N members. You allocate them in a loop. (For arrays of homogeneous instances, prefer sc_vector<Child>, which is purpose-built for this and integrates with the hierarchy and tracing; raw new is the lower-level fallback.)
  2. A child needs constructor arguments computed in the parent's constructor body. Member-initializer lists run before the constructor body; if the argument depends on logic that has to run first, you defer the child to a pointer constructed in the body.
// Pointer instantiation: child allocated in the constructor BODY.
SC_MODULE(LaneArray) {
  sc_in<bool> clk;
  std::vector<Lane*> lanes;       // owned for the simulation lifetime

  SC_CTOR(LaneArray) {
    const int N = 4;
    for (int i = 0; i < N; ++i) {
      std::string nm = "lane_" + std::to_string(i);
      Lane* l = new Lane(nm.c_str());   // heap-allocated, named per lane
      l->clk(clk);
      lanes.push_back(l);               // keep the pointer so it survives
    }
  }
  ~LaneArray() { for (auto* l : lanes) delete l; }
};

The rules that change with pointers: you must keep the pointer somewhere that lives for the whole simulation (a member std::vector, not a local), and you are responsible for delete in the destructor (or use std::unique_ptr and let RAII handle it). The rule that does not change: the child is still constructed with a name, still bound by calling ->port(signal), and still owned by the parent. Pointer-vs-member is a lifetime-management choice, not a different composition model. For everything in this post the CPU included — plain members suffice, because the structure is fixed and known at compile time.

Why named binding saves you (and what still bites)

SystemC binding is by named object: every child.port(signal) call states both endpoints. This is unlike a Verilog positional instantiation alu u(clk, a, b, op) where the third argument lands on whatever the third port happens to be. So the famous Verilog hazard — insert one port in the module header and every positional instantiation downstream silently shifts by one — cannot happen in SystemC. There is no positional binding to shift.

What can still bite is the human-error cousin. When you write a block of twenty binding calls, nothing stops you from binding the right port to the wrong same-typed signal:

// Both are sc_signal<sc_uint<32>> — types match, so this COMPILES and ELABORATES:
i_dmem.addr(sig_rs2_data);      // BUG: address should be the ALU result...
i_dmem.wr_data(sig_alu_result); // ...and store-data should be rs2. They're swapped.

The kernel cannot catch this: both ports are bound, both to a signal of the correct type, so the (E109) complete-binding check passes and simulation runs — computing wrong effective addresses on every store. The only defenses are (a) discipline: bind in a consistent order and name signals so a swap is visually obvious (sig_alu_result next to addr reads wrong at a glance), and (b) a functional testbench that exercises the path and catches the wrong behavior. We will revisit this exact swap in the Advanced section, because it is the composition bug that the type system is powerless against.

Hierarchical names for debugging

Every module, port, and signal carries a hierarchical name, set from the sc_module_name (or signal name) you pass at construction. sc_object::name() returns the full path; basename() returns just the leaf. You can print them to orient yourself in a deep hierarchy:

void Counter::end_of_elaboration() {
  std::cout << "Counter full name: " << this->name()       // e.g. "top.i_cnt"
            << "  basename: "        << this->basename()    // e.g. "i_cnt"
            << "\n";
}

The payoff is in diagnostics. When the kernel reports Error: (E109) complete binding failed: port not bound: port 'top.i_cpu.i_alu.b', the dotted path tells you exactly which instance's b port is unbound — the b of the i_alu inside the i_cpu inside top. With well-named children that message is a map straight to the offending binding call. With anonymous children it reads (unnamed).(unnamed).b and you are reduced to bisecting. Name every instance and every interconnect signal; the few seconds it costs are repaid the first time an error fires.

This concludes the Intermediate section. You can now compose a datapath plus a control leaf, share signals between the two halves, choose member versus pointer instantiation deliberately, and use hierarchical names to read elaboration errors. The Advanced section drills into the LRM corners — the exact elaboration-phase semantics, and the three composition bugs reproduced with their real error output.

Advanced: Edge Cases & LRM Corners

The Beginner and Intermediate sections cover what composition looks like when it works. This section is the other side: the exact phase semantics the standard defines for assembly, and the three composition bugs that account for most real integration failures — each reproduced as a complete, runnable program with the actual kernel diagnostic it produces.

Corner 1: the elaboration phases, precisely

The standard (IEEE 1666-2011 §4.3) defines a strict ordering of what happens before your simulation's first delta cycle. Knowing it tells you when each class of error can fire.

  1. Construction. Every module constructor runs, top-down. Child instances are created (their constructors run, recursively). sc_signal members are created. Processes are registered (SC_METHOD/SC_THREAD). Port-binding calls (port(signal)) execute, recording edges in the connection graph. All of this is ordinary C++ running inside your sc_main.
  2. before_end_of_elaboration. A callback the kernel invokes on every object after construction but before binding is finalized. Legal place to add late port bindings or create additional elaboration-time structure.
  3. End-of-elaboration / complete-binding check. The kernel walks every port. Any port of a binding-required kind that has zero bindings triggers (E109) complete binding failed. This is the gate that catches unbound ports — and it is the first moment a missing wire is detected, which is after every constructor and every line of your sc_main setup has already run.
  4. end_of_elaboration. A callback the kernel invokes on every object once binding is complete and validated. The earliest legal point to do setup that requires all ports to be connected — e.g. reading a bound channel, or printing the resolved hierarchy.
  5. Start of simulation / start_of_simulation, then the first delta. sc_start() proceeds into timed simulation.

The practical consequence for composition: an unbound-port bug does not announce itself when you write the faulty constructor. It announces itself at sc_start(). Your testbench constructor will have run, your program-load loop will have executed, your std::cout banner may already have printed — then the kernel aborts. Engineers new to SystemC routinely misread this as "the error is in my testbench" when it is in a binding call three modules down. The fix is to internalize that binding is validated at the elaboration gate, not at the call site.

Note SystemVerilog catches an unconnected required port at compile/elaborate time, before any simulation. SystemC catches it at sc_start(). Same class of bug, much later detection — budget for it.

Corner 2: unbound port at elaboration

The first classic bug: a leaf port you simply forgot to bind. Here is a minimal two-module composition where the parent omits one binding.

// file: bug_unbound.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            bug_unbound.cpp -o bug_unbound -lsystemc

#include <systemc.h>

SC_MODULE(Adder) {
  sc_in<int>  a;
  sc_in<int>  b;
  sc_out<int> sum;
  void add() { sum.write(a.read() + b.read()); }
  SC_CTOR(Adder) { SC_METHOD(add); sensitive << a << b; }
};

SC_MODULE(Top) {
  Adder i_add;
  sc_signal<int> sa, sb, ssum;
  SC_CTOR(Top) : i_add("i_add") {
    i_add.a(sa);
    // BUG: i_add.b is never bound.
    i_add.sum(ssum);
  }
};

int sc_main(int, char*[]) {
  Top top("top");
  sc_start(10, SC_NS);   // <-- the (E109) error fires HERE, not above
  return 0;
}

Expected output:

Error: (E109) complete binding failed: port not bound: port 'top.i_add.b' (sc_in)
In file: ../../../src/sysc/communication/sc_port.cpp:231

The diagnostic is precise: it names top.i_add.b, the exact unbound port, by its full hierarchical path. (The file/line in the message comes from the SystemC library internals and varies by build; the (E109) code and the port path are the stable parts.) Note where it fires: at sc_start(10, SC_NS), not at the constructor that forgot the binding. The Top constructor completed without complaint; the gap surfaced only when the kernel ran the complete-binding check. The fix is one line — add i_add.b(sb); — and this is exactly why well-named instances pay off: the path top.i_add.b walks you straight to the missing call.

Corner 3: two drivers on one signal

The second classic bug: two writers on one sc_signal. The single-writer rule (§6.4) is not enforced as a hard error in most kernels — instead the kernel issues a multiple-writers warning at elaboration, and at run time the value is whichever writer's update committed last in the evaluation phase, which is implementation-defined ordering. The result is a value that looks plausible and is non-deterministically wrong.

// file: bug_two_drivers.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            bug_two_drivers.cpp -o bug_two_drivers -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(Top) {
  sc_signal<int> bus;     // ONE signal...

  void writer_a() { bus.write(10); }   // ...driven by TWO processes
  void writer_b() { bus.write(20); }

  void reader() {
    std::cout << "[" << sc_time_stamp() << "] bus=" << bus.read() << "\n";
  }

  SC_CTOR(Top) {
    SC_METHOD(writer_a); sensitive << bus;  // (re-trigger to keep it lively)
    SC_METHOD(writer_b); sensitive << bus;
    SC_METHOD(reader);   sensitive << bus;
  }
};

int sc_main(int, char*[]) {
  Top top("top");
  sc_start(10, SC_NS);
  return 0;
}

Expected output:

Warning: (W116) sc_signal<T> cannot have more than one driver:
 signal `top.bus' (sc_signal)
 first driver `top.writer_a' (sc_method_process)
 second driver `top.writer_b' (sc_method_process)
[0 s] bus=20

The kernel warns at elaboration that top.bus has two drivers and names both (top.writer_a and top.writer_b) — again, hierarchical names make the diagnostic actionable. At run time one writer's value wins (here 20, the later-committed write in this build), but you must never rely on which. In a composed design this bug usually appears when two output ports are bound to the same interconnect signal — e.g. you bind both the ALU's result and a mux's output to sig_alu_result. The fix is structural: a signal must have exactly one writer. If two sources genuinely need to drive one net, funnel them through a single combinational mux process (one writer, selecting between the two inputs) — which is precisely why the CPU uses SC_METHOD mux processes for the ALU-source and writeback paths rather than binding two ports to one wire.

Corner 4: clock not reaching a leaf

The third classic bug is the subtlest because it can present two ways. If you forget the clock binding entirely, you get the clean (E109) from Corner 2 — the clk port is simply unbound. But the nastier variant is binding the clock to the wrong channel: a plain sc_signal<bool> that nobody toggles, instead of the sc_clock. Then the port is bound, elaboration passes, simulation runs — and the leaf's clocked process never fires, because its "clock" never has a rising edge. The leaf sits frozen while the rest of the design advances.

// file: bug_dead_clock.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            bug_dead_clock.cpp -o bug_dead_clock -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(Counter) {
  sc_in<bool>      clk;
  sc_out<unsigned> count_out;
  sc_signal<unsigned> count;
  void step()      { count.write(count.read() + 1); }
  void drive_out() { count_out.write(count.read()); }
  SC_CTOR(Counter) {
    SC_METHOD(step);      sensitive << clk.pos(); dont_initialize();
    SC_METHOD(drive_out); sensitive << count;
  }
};

SC_MODULE(Top) {
  sc_in<bool>      clk;            // the REAL clock comes in here
  sc_out<unsigned> count_out;

  Counter i_cnt;
  sc_signal<bool>     dead_clk;    // BUG: a static signal nobody drives
  sc_signal<unsigned> count;

  SC_CTOR(Top) : i_cnt("i_cnt") {
    i_cnt.clk(dead_clk);           // bound to the dead signal, NOT to clk
    i_cnt.count_out(count);
    SC_METHOD(fwd); sensitive << count;
  }
  void fwd() { count_out.write(count.read()); }
};

SC_MODULE(Mon) {
  sc_in<bool>      clk;
  sc_in<unsigned> count_out;
  void watch() {
    std::cout << "[" << sc_time_stamp() << "] count=" << count_out.read() << "\n";
  }
  SC_CTOR(Mon) { SC_METHOD(watch); sensitive << clk.pos(); dont_initialize(); }
};

int sc_main(int, char*[]) {
  sc_clock clk("clk", 10, SC_NS);
  sc_signal<unsigned> count_out;
  Top top("top"); top.clk(clk); top.count_out(count_out);
  Mon mon("mon"); mon.clk(clk); mon.count_out(count_out);
  sc_start(50, SC_NS);
  sc_stop();
  return 0;
}

Expected output:

[10 ns] count=0
[20 ns] count=0
[30 ns] count=0
[40 ns] count=0
[50 ns] count=0

No error, no warning — and the counter never moves. Its step process is sensitive to dead_clk.pos(), and dead_clk is a default-constructed sc_signal<bool> that nobody ever writes, so it never has a positive edge, so step never runs. The count is stuck at its initial value forever while the monitor faithfully reports a healthy-looking (and entirely wrong) zero every cycle. This is the failure mode behind "one block in my design is frozen": its clock port is bound to something that does not toggle. The fix is i_cnt.clk(clk) — route the real clock down. The lesson generalizes: when one leaf in a composition is inexplicably inert, check what its clk port is actually bound to before suspecting its internal logic.

Corner 5: the same-type miswire the kernel cannot catch

For completeness, return to the swap from the Intermediate section, now as a standalone demonstration of why it is uncatchable. Two sc_signal<int> wires, two ports of type sc_in<int>, bound crosswise:

// Inside a parent constructor:
sc_signal<int> sig_x, sig_y;
i_block.in_a(sig_y);   // intended sig_x
i_block.in_b(sig_x);   // intended sig_y

Every port is bound. Every binding is type-correct (int to sc_in<int>). The (E109) check passes. There is no warning. Elaboration completes and the simulation runs the wrong computation with full confidence. The kernel has no notion of which signal should feed which port — it only knows the types line up. This is the structural reason composition correctness is a two-part claim: "every port bound" is checkable by the kernel; "every port bound to the right signal" is checkable only by a functional testbench. It is also why the CPU example below ends with a smoke test rather than just an elaboration pass: passing elaboration proves the wires exist, not that they go to the right places.

With the phase model and the three (plus one) classic bugs in hand, we can compose the real thing.

Worked example: composing the single-cycle RV32I CPU

Everything above scales, unchanged, to a real machine. The single-cycle RV32I CPU is six leaf modules — program counter, instruction memory, decoder, register file, ALU, data memory — wired into one Rv32iCpu container, plus four glue SC_METHOD processes for the muxes and branch logic that live between the leaves. The leaves are the ones you built in earlier parts: the ALU from Part 8, the program counter from Part 10, the FSM discipline from Part 11, the register file from Part 12, and the data memory from Part 13. This post does not rebuild them; it composes them.

The same three facts apply at this scale. Every leaf is a member of Rv32iCpu. Every wire between leaves is an sc_signal owned by Rv32iCpu. Every connection is a named operator() binding. The only thing that grows is the count: about thirty interconnect signals instead of one.

Here is the datapath, every arrow a parent-owned signal you must declare and bind:

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
flowchart LR
    PC["PC"] -->|pc| IMEM["IMEM"]
    IMEM -->|instr| DEC["Decoder"]
    DEC -->|rs1,rs2,rd,reg_write| RF["RegFile"]
    RF -->|rs1_data| ALU["ALU"]
    RF -->|rs2_data| BMUX["ALU-B Mux"]
    DEC -->|imm,alu_src| BMUX
    BMUX -->|alu_b| ALU
    DEC -->|alu_op| ALU
    ALU -->|result| DMEM["DMEM"]
    RF -->|rs2_data| DMEM
    ALU -->|result| WB["WB Mux"]
    DMEM -->|rd_data| WB
    PC -->|pc+4| WB
    WB -->|wr_data| RF
    ALU -->|zero| BR["Branch/Next-PC"]
    DEC -->|branch,imm| BR
    PC -->|pc| BR
    BR -->|next_pc| PC

Read it as datapath plus control. The datapath is PC → IMEM → Decoder → RegFile → ALU → DMEM → writeback, the left-to-right flow of an instruction's data. The control is the decoder's output bundle (alu_op, alu_src, reg_write, mem_read, mem_write, wb_sel, branch) steering the muxes and enables — exactly the FSM-gates-counter pattern from the Intermediate section, now with a decoder in the FSM's role and a full datapath in the counter's. The four glue processes (alu_b_mux, pc_plus4, branch_logic, writeback_mux) are the combinational seams where control selects among datapath values; each is the single writer of its output signal, so none of them violates the single-driver rule.

Rv32iCpu — the composed top module

// file: rv32i_cpu.h  (composition only — leaf modules from earlier parts)
#pragma once
#include <systemc.h>
#include "pc.h"
#include "imem.h"
#include "decoder.h"
#include "reg_file.h"
#include "alu.h"
#include "dmem.h"

SC_MODULE(Rv32iCpu) {
  // ---- External ports ----
  sc_in<bool>          clk;
  sc_in<bool>          rst;
  sc_out<sc_uint<32>>  dbg_pc;       // observation: current PC
  sc_out<sc_uint<32>>  dbg_instr;    // observation: current instruction
  sc_out<bool>         dbg_halt;     // observation: EBREAK seen

  // ---- Leaf instances (members) ----
  Pc       i_pc;
  Imem     i_imem;
  Decoder  i_dec;
  RegFile  i_rf;
  Alu      i_alu;
  Dmem     i_dmem;

  // ---- Interconnect signals (parent owns every wire) ----
  sc_signal<sc_uint<32>> sig_pc;          // current PC
  sc_signal<sc_uint<32>> sig_pc_plus4;    // PC + 4 (link value, fall-through)
  sc_signal<sc_uint<32>> sig_instr;       // fetched instruction word
  sc_signal<sc_uint<32>> sig_next_pc;     // next PC (from branch logic)

  sc_signal<sc_uint<5>>  sig_rs1_addr;
  sc_signal<sc_uint<5>>  sig_rs2_addr;
  sc_signal<sc_uint<5>>  sig_rd_addr;

  sc_signal<sc_uint<32>> sig_imm;
  sc_signal<sc_uint<4>>  sig_alu_op;
  sc_signal<bool>        sig_alu_src;      // 0 = rs2, 1 = imm
  sc_signal<bool>        sig_reg_write;
  sc_signal<bool>        sig_mem_read;
  sc_signal<bool>        sig_mem_write;
  sc_signal<sc_uint<3>>  sig_funct3;
  sc_signal<sc_uint<2>>  sig_wb_sel;       // 0 = alu, 1 = mem, 2 = pc+4
  sc_signal<bool>        sig_branch;
  sc_signal<bool>        sig_halt;

  sc_signal<sc_uint<32>> sig_rs1_data;
  sc_signal<sc_uint<32>> sig_rs2_data;

  sc_signal<sc_uint<32>> sig_alu_b;        // muxed ALU B input
  sc_signal<sc_uint<32>> sig_alu_result;
  sc_signal<bool>        sig_alu_zero;

  sc_signal<sc_uint<32>> sig_mem_rd_data;
  sc_signal<sc_uint<32>> sig_wr_data;      // writeback bus

  // ---- Glue processes (the seams between leaves) ----
  void alu_b_mux();
  void pc_plus4();
  void branch_logic();
  void writeback_mux();
  void drive_dbg();

  SC_CTOR(Rv32iCpu)
    : i_pc("i_pc"), i_imem("i_imem"), i_dec("i_dec"),
      i_rf("i_rf"), i_alu("i_alu"), i_dmem("i_dmem")
  {
    // --- PC ---
    i_pc.clk(clk);
    i_pc.rst(rst);
    i_pc.next_pc(sig_next_pc);
    i_pc.pc_out(sig_pc);

    // --- IMEM ---
    i_imem.addr(sig_pc);
    i_imem.instr(sig_instr);

    // --- Decoder ---
    i_dec.instr(sig_instr);
    i_dec.rs1_addr(sig_rs1_addr);
    i_dec.rs2_addr(sig_rs2_addr);
    i_dec.rd_addr(sig_rd_addr);
    i_dec.imm(sig_imm);
    i_dec.alu_op(sig_alu_op);
    i_dec.alu_src(sig_alu_src);
    i_dec.reg_write(sig_reg_write);
    i_dec.mem_read(sig_mem_read);
    i_dec.mem_write(sig_mem_write);
    i_dec.funct3(sig_funct3);
    i_dec.wb_sel(sig_wb_sel);
    i_dec.branch(sig_branch);
    i_dec.halt(sig_halt);

    // --- Register File ---
    i_rf.clk(clk);
    i_rf.rst(rst);
    i_rf.rs1_addr(sig_rs1_addr);
    i_rf.rs2_addr(sig_rs2_addr);
    i_rf.rd_addr(sig_rd_addr);
    i_rf.rd_data(sig_wr_data);
    i_rf.reg_write(sig_reg_write);
    i_rf.rs1_data(sig_rs1_data);
    i_rf.rs2_data(sig_rs2_data);

    // --- ALU ---
    i_alu.a(sig_rs1_data);
    i_alu.b(sig_alu_b);
    i_alu.op(sig_alu_op);
    i_alu.result(sig_alu_result);
    i_alu.zero(sig_alu_zero);

    // --- Data Memory ---
    i_dmem.clk(clk);
    i_dmem.rst(rst);
    i_dmem.addr(sig_alu_result);
    i_dmem.wr_data(sig_rs2_data);
    i_dmem.mem_read(sig_mem_read);
    i_dmem.mem_write(sig_mem_write);
    i_dmem.funct3(sig_funct3);
    i_dmem.rd_data(sig_mem_rd_data);

    // --- Glue: ALU-B mux (single writer of sig_alu_b) ---
    SC_METHOD(alu_b_mux);
    sensitive << sig_rs2_data << sig_imm << sig_alu_src;

    // --- Glue: PC + 4 ---
    SC_METHOD(pc_plus4);
    sensitive << sig_pc;

    // --- Glue: branch / next-PC (single writer of sig_next_pc) ---
    SC_METHOD(branch_logic);
    sensitive << sig_branch << sig_funct3 << sig_alu_zero
              << sig_pc << sig_imm << sig_halt;

    // --- Glue: writeback mux (single writer of sig_wr_data) ---
    SC_METHOD(writeback_mux);
    sensitive << sig_alu_result << sig_mem_rd_data
              << sig_pc_plus4 << sig_wb_sel;

    // --- Observation outputs ---
    SC_METHOD(drive_dbg);
    sensitive << sig_pc << sig_instr << sig_halt;
  }
};

The implementation file holds only the five glue processes. Each is the single writer of exactly one interconnect signal — the discipline that keeps every wire single-driven.

// file: rv32i_cpu.cpp
#include "rv32i_cpu.h"

// sig_alu_b = alu_src ? imm : rs2_data
void Rv32iCpu::alu_b_mux() {
  if (sig_alu_src.read()) sig_alu_b.write(sig_imm.read());
  else                    sig_alu_b.write(sig_rs2_data.read());
}

void Rv32iCpu::pc_plus4() {
  sig_pc_plus4.write(sig_pc.read() + 4);
}

// Next-PC selection: halt freezes; taken branch is PC-relative; else PC+4.
// This single program is BEQ-only for the smoke test, so funct3==0 (BEQ)
// is the only branch decoded; the default covers the rest.
void Rv32iCpu::branch_logic() {
  sc_uint<32> pc  = sig_pc.read();
  sc_uint<32> imm = sig_imm.read();
  bool taken = false;
  if (sig_branch.read()) {
    switch (sig_funct3.read()) {
      case 0: taken = sig_alu_zero.read();  break;  // BEQ
      default: taken = false;               break;
    }
  }
  sc_uint<32> next;
  if (sig_halt.read())   next = pc;          // EBREAK: hold
  else if (taken)        next = pc + imm;     // branch taken
  else                   next = pc + 4;       // sequential
  sig_next_pc.write(next);
}

// Writeback source select.
void Rv32iCpu::writeback_mux() {
  sc_uint<32> v;
  switch (sig_wb_sel.read()) {
    case 0:  v = sig_alu_result.read();  break;  // arithmetic
    case 1:  v = sig_mem_rd_data.read(); break;  // load
    case 2:  v = sig_pc_plus4.read();    break;  // link
    default: v = sig_alu_result.read();  break;
  }
  sig_wr_data.write(v);
}

void Rv32iCpu::drive_dbg() {
  dbg_pc.write(sig_pc.read());
  dbg_instr.write(sig_instr.read());
  dbg_halt.write(sig_halt.read());
}

Note how every line of that constructor is one of the three facts in action. The six leaves are members, constructed-and-named in the initializer list. Roughly thirty sig_* signals are members — the parent owns every wire. Each binding call names its port and its signal. The four glue muxes are single writers of sig_alu_b, sig_next_pc, sig_wr_data, and sig_pc_plus4 respectively — the same discipline that kept bus from having two drivers in the Advanced section, applied at scale.

Whole-CPU smoke test: a three-instruction program

Composition that elaborates is not composition that works — recall Corner 5: the kernel proves the wires exist, not that they go to the right leaves. The only proof is to run a program and check the result. We run three instructions:

0x00500093   addi x1, x0, 5     # x1 = 5
0x00300113   addi x2, x0, 3     # x2 = 3
0x002081b3   add  x3, x1, x2    # x3 = x1 + x2 = 8
0x00100073   ebreak             # halt

After execution, x1 must be 5, x2 must be 3, and x3 must be 8. The testbench loads the program into IMEM, pulses reset, runs until dbg_halt, traces each cycle, then checks the register file.

// file: tb_cpu.cpp
#include <systemc.h>
#include "rv32i_cpu.h"
#include <iostream>
#include <iomanip>

SC_MODULE(TbCpu) {
  sc_signal<bool>         rst;
  sc_signal<sc_uint<32>>  dbg_pc, dbg_instr;
  sc_signal<bool>         dbg_halt;

  sc_clock clk;          // 10 ns period, 50% duty
  Rv32iCpu u_cpu;

  static const uint32_t prog[4];

  void run_test() {
    // Load the program into instruction memory.
    for (int i = 0; i < 4; ++i)
      u_cpu.i_imem.load_word(i * 4, prog[i]);

    // Hold reset high for one cycle, then release.
    rst.write(true);
    wait(10, SC_NS);
    rst.write(false);

    std::cout << "=== single-cycle RV32I smoke test ===\n";
    int cycle = 0;
    while (!dbg_halt.read() && cycle < 20) {
      wait(clk.posedge_event());
      wait(SC_ZERO_TIME);                 // let combinational settle
      std::cout << "[cyc " << std::dec << std::setw(2) << cycle << "] "
                << "PC=0x" << std::hex << std::setw(2) << std::setfill('0')
                << dbg_pc.read()
                << " instr=0x" << std::setw(8) << dbg_instr.read()
                << std::setfill(' ') << "\n";
      ++cycle;
    }

    // Check architectural state.
    bool pass = true;
    auto chk = [&](int r, uint32_t want) {
      uint32_t got = u_cpu.i_rf.read_reg(r);
      if (got != want) {
        std::cout << "FAIL x" << std::dec << r << " = 0x" << std::hex << got
                  << " (want 0x" << want << ")\n";
        pass = false;
      }
    };
    chk(1, 5);
    chk(2, 3);
    chk(3, 8);
    std::cout << (pass ? "PASS: all registers correct\n"
                       : "FAIL: register mismatch\n");
    sc_stop();
  }

  SC_CTOR(TbCpu)
    : clk("clk", 10, SC_NS), u_cpu("u_cpu")
  {
    u_cpu.clk(clk);
    u_cpu.rst(rst);
    u_cpu.dbg_pc(dbg_pc);
    u_cpu.dbg_instr(dbg_instr);
    u_cpu.dbg_halt(dbg_halt);
    SC_THREAD(run_test);
  }
};

const uint32_t TbCpu::prog[4] = {
  0x00500093u,   // addi x1, x0, 5
  0x00300113u,   // addi x2, x0, 3
  0x002081b3u,   // add  x3, x1, x2
  0x00100073u,   // ebreak
};

int sc_main(int, char*[]) {
  TbCpu tb("tb");
  sc_start();
  return 0;
}

Expected output:

=== single-cycle RV32I smoke test ===
[cyc  0] PC=0x00 instr=0x00500093
[cyc  1] PC=0x04 instr=0x00300113
[cyc  2] PC=0x08 instr=0x002081b3
[cyc  3] PC=0x0c instr=0x00100073
PASS: all registers correct

Each cycle the PC advances by 4 and a new instruction is fetched, decoded, executed, and (for the three addi/add instructions) written back. At PC 0x0c the decoder raises halt for the ebreak, the branch logic freezes the PC, the loop sees dbg_halt and exits, and the register check confirms x1=5, x2=3, x3=8. The composition is correct: the wires exist and they go to the right leaves. Had the ALU-B mux been miswired, add x3, x1, x2 would have computed x1 + imm instead of x1 + x2, the register check would have caught x3 != 8, and the PASS line would have been a FAIL — which is exactly why the smoke test exists and why "it elaborated" is never sufficient evidence that a composition works.

One observation about the seam between datapath and control in this trace. The decoder's control bundle — alu_src, alu_op, reg_write, wb_sel — is purely combinational on the fetched instruction; it settles in the same cycle the instruction appears. The register file's write, however, is clocked: the writeback happens on the rising edge that ends the instruction's cycle. That is why this is a single-cycle machine — every instruction completes its datapath flow and commits its register write in one clock period. The composition did not create that timing; the leaves did. Composition merely connected them so the timing each leaf was built for lines up across the whole machine.

Hands-on exercise

Build a composed two-stage processing pipeline from scratch and prove it works, then break it deliberately to see each composition error fire.

Part A — compose. Build three leaf modules and one parent:

  • Producer — a clocked module (synchronous active-high reset) that outputs an incrementing 8-bit value on data_out each cycle, starting at 0.
  • Scaler — a combinational module with data_in and a one-bit mode input. When mode is 0 it outputs data_in; when mode is 1 it outputs data_in << 1 (doubled). Output is data_out.
  • Threshold — a combinational module that asserts over when its data_in exceeds 20.
  • PipeTop — owns all three leaves and the wires between them. The producer feeds the scaler; the scaler feeds the threshold. Route clk/rst to the producer. Drive mode from a parent port. Forward over and the scaler's output up to parent ports for observation.

Drive it from a testbench with mode = 1 (doubling on), a 10 ns clock, one reset cycle. Predict the cycle at which over first asserts (the producer counts 0, 1, 2, …; the scaler doubles; the threshold trips above 20 — so the doubled value first exceeds 20 when the raw count reaches 11, i.e. 11 << 1 = 22 > 20). Run and confirm.

Part B — add control. Add a ModeFsm leaf: a two-state Moore FSM that flips mode every four cycles (four cycles of pass-through, four cycles of doubling, repeating). Bind its mode output to the same parent-owned mode signal the scaler reads — removing the parent's mode port. Now the control half (the FSM) and the datapath half (producer → scaler → threshold) share one signal. Trace over and confirm it tracks the alternating mode.

Part C — break it three ways. Make three copies of the working Part B and introduce exactly one composition bug in each, then run and read the diagnostic:

  1. In copy 1, delete the i_scaler.data_in(...) binding. Confirm you get (E109) complete binding failed naming the exact port, and confirm it fires at sc_start, not at the constructor.
  2. In copy 2, bind a second writer to the mode signal (e.g. add a stray SC_METHOD in the parent that also writes mode). Confirm the (W116) multiple-driver warning names both drivers.
  3. In copy 3, bind the producer's clk to a fresh sc_signal<bool> that nobody drives, instead of the real clock. Confirm there is no error — and that over never asserts because the producer is frozen.

Part D — name and inspect. Add an end_of_elaboration() override to PipeTop that prints this->name() and the name() of each leaf. Confirm the hierarchical paths match what the error messages in Part C reported.

No solution is provided. The whole point is that composition is learned by wiring, mis-wiring, and reading the kernel's reaction.

Hints

  • Every leaf instance is a member of PipeTop and is constructed in the initializer list with a name string. Every interconnect is a member sc_signal of PipeTop.
  • The producer is the only clocked leaf in Part A; the scaler and threshold are combinational and have no clk port. In Part B the ModeFsm is also clocked.
  • Each combinational leaf's process must list every input it reads in sensitive << .... A frozen-output bug usually means a missing sensitivity entry, not a binding error.
  • For the multiple-driver bug in Part C, remember that two output ports bound to one signal, or two processes writing one signal, both trip (W116). Either reproduces it.
  • When the clock-to-a-dead-signal bug produces no diagnostic, that is the lesson: the type system and the binding check are both blind to it; only the missing output behavior reveals it.

Common mistakes

  • Declaring a child instance or interconnect signal as a constructor-local variable. A Child c("c"); or sc_signal<int> w; written inside the constructor body is destroyed when the constructor returns, leaving every port bound to it dangling. The simulation crashes — often with a confusing access-violation far from the actual cause. Fix: child instances and interconnect signals are always members of the parent (or heap-allocated with a member-held pointer that lives for the whole run). If it is a wire or a sub-block, it is a member.
  • Forgetting to bind a port, then blaming the testbench. An unbound port fires (E109) complete binding failed at sc_start() — after every constructor and all of sc_main has run. Engineers misread the timing as "the error is where execution was," but the error is at the binding call you omitted three modules down. Fix: read the hierarchical path in the (E109) message (top.i_cpu.i_alu.b); it names the exact unbound port. Bind it.
  • Binding the right port to the wrong same-typed signal. Swapping two sc_signal<sc_uint<32>> wires (e.g. address and write-data into the data memory) compiles, elaborates, and runs — silently wrong. The kernel cannot catch a semantic miswire; types match. Fix: name signals so a swap reads wrong at a glance, bind in a consistent order, and — the only real defense — run a functional test that exercises the path. Elaboration passing is not correctness.
  • Two drivers on one interconnect signal. Binding two output ports to one sc_signal, or having two processes write it, produces a (W116) multiple-driver warning and a value that depends on implementation-defined update ordering. Fix: every signal has exactly one writer. If two sources must drive one net, funnel them through a single combinational mux process (one writer selecting between inputs) — exactly what the CPU's ALU-B and writeback muxes do.
  • Clock bound to a channel that never toggles. Binding a clocked leaf's clk port to a plain sc_signal<bool> that nobody drives passes elaboration (the port is bound) but freezes the leaf — its clocked process never sees a rising edge. No error, no warning, just one inert block in an otherwise-running design. Fix: when a leaf is inexplicably frozen, check what its clk port is actually bound to before suspecting its logic. Route the real sc_clock down to every clocked leaf.
  • Leaving child instances unnamed. A child constructed without a name string (or omitted from the initializer list) gets an empty or generated name, and every diagnostic about it reads (unnamed).b instead of top.i_alu.b. Composition errors become a bisecting exercise. Fix: name every instance and every interconnect signal at construction; the cost is a few characters, the payoff is every future error message.

Recap

After working through this post you can now:

  • State the three facts of structural composition: a child is a member of its parent, the parent owns every wire (interconnect sc_signal) between its children, and binding names both endpoints by object so call-order is irrelevant to correctness.
  • Compose two leaf modules through a single parent-owned interconnect signal, routing the clock down to the clocked leaf and forwarding outputs up to parent ports.
  • Add a control leaf to a datapath and recognize the datapath-plus-control pattern — control outputs and datapath control-inputs meeting on shared parent-owned signals — at any scale.
  • Choose member instantiation by default and pointer/sc_vector instantiation when child count is dynamic or constructor arguments must be computed first.
  • Explain the elaboration phase model and why an unbound-port error fires at sc_start() rather than at the offending constructor.
  • Read and fix the three classic composition bugs on sight: the unbound port ((E109), names the port), two drivers on one signal ((W116), names both), and the clock bound to a dead channel (no diagnostic — a frozen leaf).
  • Use sc_object::name() / basename() and well-named instances to make every kernel diagnostic a map straight to the faulty binding.
  • Wire six leaf modules and four glue processes into a complete single-cycle RV32I CPU and prove the composition with a smoke test, understanding that elaboration proves the wires exist while only the test proves they go to the right leaves.

Further reading

Standards

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §5.2 (sc_module and hierarchy), §5.3 / §5.12–5.13 (ports and binding), §4.3.4 (elaboration and the complete-binding check), §6.4 (sc_signal and the single-writer rule), §5.15 (sc_object::name / basename), §5.2.13 (before_end_of_elaboration / end_of_elaboration).

Vendor and consortium documents

  • Accellera Systems Initiative, SystemC 2.3.x User Guide — sections on module hierarchy, port binding, and sc_vector.
  • Doulos, SystemC Golden Reference Guide, ch. 4–5 — sc_module, port binding, and sc_module_name conventions.

Textbooks

  • Bhasker, A SystemC Primer (2nd ed.), ch. 4 & 7 — module hierarchy and interconnect-signal idioms.
  • Grötker, Liao, Martin, and Swan, System Design with SystemC, ch. 3–4 — structural composition and the parent-owns-the-wire model.

Real-world references

  • CV32E40P top-level RTL (OpenHW Group) — a production RISC-V core whose datapath-plus-control composition mirrors, at industrial scale, the structure of the CPU built here.
  • Ibex top-level RTL (lowRISC) — independent corroboration of the same composition pattern in a different open-source core.

Next in this section

→ Part 8: Capstone — Build the Rest Yourself. You have composed a working single-cycle CPU from leaf modules. The capstone hands the work to you: extend the datapath to cover the remaining RV32I instructions, add the branch and load/store paths in full, and verify your machine against a software reference model running in lockstep — the exact methodology real CPU vendors use to sign off a core.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment