8. SystemC Tutorial - Combinational Modeling Patterns
Why this matters
Rewritten 2026-06-05 with deeper first-principles material.
Combinational logic is the soil that every other RTL pattern grows in. Before a chip has a single flip-flop, before it has a state machine or a pipeline or a memory, it has combinational cones — clouds of gates that take some inputs and produce some outputs with no memory and no clock. The carry-lookahead adder in the integer unit of the CPU you are reading this on is combinational. The address decoder that splits a physical address into a DRAM bank, row, and column is combinational. The priority encoder in the interrupt controller that picks the highest-pending IRQ is combinational. The byte-enable generator in a memory controller, the parity tree on an ECC-protected bus, the barrel shifter in a GPU's texture unit, the comparator that decides whether a TLB entry hit — all combinational. When a verification engineer finds that a design computes the wrong sum, the wrong comparison result, or the wrong mux selection, the bug is almost always in a combinational block whose inputs were not all accounted for, or whose output was not driven on every path.
Modeling combinational logic correctly in SystemC is therefore the first skill the RTL-patterns section teaches, and every later pattern depends on it. A finite state machine is a state register feeding combinational next-state logic. A datapath is combinational arithmetic between registers. A memory's read port is combinational address-to-data logic gated by an enable. If you cannot write a correct combinational block — sensitive to every input, driving every output on every path, settling deterministically across delta cycles — you cannot write any of the later patterns, because they are all built on top of it. Get the sensitivity list wrong and your block silently stops tracking one of its inputs while the simulation keeps running and the waveforms look plausible. Forget to assign an output on one branch and your "combinational" block quietly grows a memory it was never meant to have — the SystemC analogue of unintended latch inference, the single most-cited mistake in synthesizable RTL. This post fixes all of that from first principles. By the end you will be able to write a combinational block in SystemC from scratch, predict its output without running the simulator, reason about how a chain of combinational blocks settles over delta cycles, and recognize the two combinational bugs that account for most "the model is wrong but the simulation runs fine" reports on sight. We build the concept on a generic 4-bit adder first, then apply it to the RV32I ALU — the worked example that recurs through this whole section. That is the bar.
Prerequisites
- Part 1 — Modules, Ports & Signals. You need to be comfortable declaring an
SC_MODULE, declaringsc_in/sc_outports, bindingsc_signalchannels to those ports, and registering a process withSC_METHODinsideSC_CTOR. Every example in this post is a module with ports, a process, and a sensitivity list. - Part 4 — Processes & Sensitivity (SC_METHOD vs SC_THREAD). You need to understand why combinational logic uses
SC_METHOD(a non-blocking process that runs to completion) rather thanSC_THREAD, and how static sensitivity is built withsensitive << a << b << op. This post is, in one sense, an extended application of that part to the specific case of combinational hardware. - SystemC 2.3.x installed. All examples compile with a C++17 compiler (
g++9 or newer, orclang++10 or newer) against any 2.3.x SystemC build. If your environment is not set up yet, the installation tutorials earlier in this series cover Linux and macOS.
Two concepts from earlier parts get heavy use here: SC_METHOD (the process kind every combinational block uses) and static sensitivity (the sensitive << ... invocation that wires a process to its inputs). One more idea — the delta cycle and sc_signal's buffered update — appears in the Intermediate section when we look at how combinational chains settle. If any of those is hazy, re-read the relevant part before continuing.
Mental model (first principles)
A combinational function in hardware has a precise definition: its outputs are a pure function of its present inputs. There is no clock, no state, no memory. Give the same inputs and you get the same outputs, every time, with no dependence on history. A 2-input AND gate is combinational. A 32-bit adder is combinational. A multiplexer is combinational. The defining property is no feedback through storage — the output today does not depend on the output yesterday.
In hardware, a combinational block is drawn as a cloud of gates with inputs entering one side and outputs leaving the other:
┌───────────────────────────┐
a ──►│ │
b ──►│ combinational logic ├──► y
op ──►│ (no clock, no state) │
│ y = f(a, b, op) ├──► z
└───────────────────────────┘
The cloud has no clock input because nothing inside it is timed. The instant any input changes, the outputs recompute. In real silicon "the instant" is actually the propagation delay through the gates, but at the RTL modeling level we treat the recomputation as taking zero simulated time — it happens, conceptually, the moment an input settles.
In SystemC there is exactly one construct that captures this: an SC_METHOD process made statically sensitive to every input the block reads. The mapping is direct and worth memorizing because the rest of this post is variations on it:
- The combinational logic is the body of one
SC_METHOD. It reads its inputs with.read(), computes, and writes its outputs with.write(). It runs to completion every time it is invoked — nowait(), no suspension. That is the whole reason combinational logic usesSC_METHODand notSC_THREAD: a combinational block has no notion of "pause here and resume later," so the non-blocking, run-to-completionSC_METHODis the natural fit.
- The phrase "the instant any input changes, the outputs recompute" is expressed by the static sensitivity list:
sensitive << a << b << op. This wires the process to the value-changed events of every named input. When any ofa,b, oropchanges value, the kernel makes the process runnable, the body re-executes, and fresh outputs are written. This is the load-bearing line. If an input is read in the body but missing from the sensitivity list, the process will not wake when that input changes, and the output will silently go stale — it will keep showing the value computed from the last input that was in the list. SystemC has noalways @(*)auto-sensitivity. You build the list by hand, and you must list every input you read.
- The outputs are
sc_signal-backed ports written with.write(). Awriteis a request-update: the new value is buffered and committed in the kernel's update phase, after which the output's value-changed event fires. This matters the moment one combinational block feeds another — the downstream block does not see the new value until the update phase commits it and a fresh delta cycle begins. We will trace this precisely in the Intermediate section.
Two discipline rules turn this mapping into correct hardware. They are the entire substance of combinational modeling, and almost every combinational bug is a violation of one of them.
Rule 1 — sensitivity completeness: list every input you read. If the body reads cin but the sensitivity list is only sensitive << a << b, then a change on cin alone does not wake the process. The output reflects the old cin until some other input (a or b) changes and incidentally re-runs the body. The symptom is an output that intermittently ignores one of its inputs — maddening to debug because the block is "mostly right." In hardware terms you have modeled a block whose cin pin is disconnected from the recompute logic, which is not a thing real gates do.
Rule 2 — output completeness: assign every output on every path. A combinational output, by definition, is a function of the present inputs on all input combinations. In SystemC, if some path through your body fails to call .write() on an output, that output's sc_signal simply keeps its previous value. The block now remembers — it behaves like a latch, holding the last value it was given until a path that does assign the output executes. This is exactly the "latch inference" failure that synthesis tools warn about in Verilog and VHDL, except SystemC gives you no warning at all. The fix is mechanical: initialize every output to a default at the top of the body, or ensure every branch assigns it.
Here is the smallest example that exercises both rules — a generic 4-bit adder. Inputs a, b, and a carry-in cin; outputs a 4-bit sum and a carry-out cout. It is combinational: sum and cout are pure functions of a, b, cin. The sensitivity list names all three inputs (Rule 1). The body assigns both outputs unconditionally (Rule 2). We will build it in full in the Beginner section; here, hold the shape in your head:
sensitive << a << b << cin; // Rule 1: every input listed
...
sum.write(...); // Rule 2: both outputs
cout.write(...); // assigned every time
That two-rule, one-SC_METHOD pattern is the entire mental model. Everything else in this post is detail: the RV32I ALU is the same pattern with a wider input cone and signed-vs-unsigned type care; the delta-cycle discussion is about when the recomputed output becomes visible; the bug walkthroughs are violations of Rule 1 and Rule 2. Internalize the two rules and every combinational block you ever write is just filling in the function f.
flowchart LR
A[a] --> ADD[adder
SC_METHOD]
B[b] --> ADD
CIN[cin] --> ADD
ADD --> SUM[sum]
ADD --> COUT[cout]
style ADD fill:#d1fae5,stroke:#10b981
Read the diagram as: three inputs enter one combinational process, two outputs leave it, no clock anywhere. That visual — inputs in, outputs out, no clock pin — is the signature of combinational logic. The moment a clock arrow appears, you are looking at sequential logic, which is the next part of this section.
Beginner: First Principles
The simplest interesting combinational block is an adder. We will build a 4-bit adder with carry-in and carry-out, run it in our heads, predict the output, then verify on a real kernel. Nothing here is RISC-V-specific — this is generic combinational hardware, and the patterns transfer to every combinational block you will ever write.
// file: adder4.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// adder4.cpp -o adder4 -lsystemc
// Run: LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./adder4
#include <systemc.h>
#include <iostream>
// A 4-bit adder: sum and carry-out are pure functions of a, b, cin.
// One SC_METHOD, sensitive to ALL three inputs (Rule 1).
// Both outputs assigned on every invocation (Rule 2).
SC_MODULE(Adder4) {
sc_in<sc_uint<4>> a;
sc_in<sc_uint<4>> b;
sc_in<bool> cin;
sc_out<sc_uint<4>> sum;
sc_out<bool> cout;
void compute() {
// Widen to 5 bits so the carry-out is captured, then split.
sc_uint<5> full = a.read() + b.read() + (cin.read() ? 1 : 0);
sum.write(full.range(3, 0)); // low 4 bits = sum
cout.write(full[4]); // bit 4 = carry-out
}
SC_CTOR(Adder4) {
SC_METHOD(compute);
sensitive << a << b << cin; // Rule 1: every input read appears here
// No dont_initialize(): we WANT this to run at t=0 to establish outputs.
}
};
int sc_main(int /*argc*/, char* /*argv*/[]) {
sc_signal<sc_uint<4>> a, b, sum;
sc_signal<bool> cin, cout;
Adder4 add("add");
add.a(a); add.b(b); add.cin(cin); add.sum(sum); add.cout(cout);
// Helper lambda: apply inputs, settle one delta, print result.
auto apply = [&](unsigned av, unsigned bv, bool cv) {
a.write(av); b.write(bv); cin.write(cv);
sc_start(SC_ZERO_TIME); // let the SC_METHOD fire and outputs settle
std::cout << "a=" << av << " b=" << bv << " cin=" << cv
<< " -> sum=" << sum.read()
<< " cout=" << cout.read() << "\n";
};
apply(3, 4, false); // 3 + 4 + 0 = 7, no carry
apply(9, 6, false); // 9 + 6 + 0 = 15, no carry (fits in 4 bits)
apply(9, 7, false); // 9 + 7 + 0 = 16 -> sum=0, carry=1
apply(15, 15, true); // 15 + 15 + 1 = 31 -> sum=15, carry=1
apply(0, 0, false); // 0 + 0 + 0 = 0
return 0;
}
Compile and run, and you will see this output. Read it before the walkthrough — try to compute each line yourself first.
Expected output:
a=3 b=4 cin=0 -> sum=7 cout=0
a=9 b=6 cin=0 -> sum=15 cout=0
a=9 b=7 cin=0 -> sum=0 cout=1
a=15 b=15 cin=1 -> sum=15 cout=1
a=0 b=0 cin=0 -> sum=0 cout=0
Walk through it. The adder is purely combinational, so there is no clock and no sc_clock in sc_main at all — a fact worth pausing on, because it is the visible difference between this post and every later one in the section. We drive the inputs directly through sc_signals and call sc_start(SC_ZERO_TIME) to let the kernel run the combinational method and settle the outputs without advancing simulated time.
Line by line: 3 + 4 = 7, fits in 4 bits, no carry — sum=7, cout=0. Then 9 + 6 = 15, the largest value a 4-bit number holds, still no carry — sum=15, cout=0. Then 9 + 7 = 16, which is 0b10000; the low four bits are 0 and bit 4 is the carry, so sum=0, cout=1. Then 15 + 15 + 1 = 31 = 0b11111; low four bits 1111 = 15, bit 4 set, so sum=15, cout=1. Finally 0 + 0 = 0.
Notice the design choices that make Rule 1 and Rule 2 concrete.
Rule 1 in the code. The sensitivity list is sensitive << a << b << cin. All three inputs the body reads are named. If you change any of a, b, or cin, the kernel wakes compute() and recomputes both outputs. Drop one — say you wrote sensitive << a << b and forgot cin — and the adder would ignore the carry-in: toggling cin from 0 to 1 with a and b unchanged would not re-run the method, and cout/sum would keep their previous values. The simulation would not crash; it would simply give a wrong answer for any test that changes only cin. We demonstrate exactly this bug in the Common mistakes section.
Rule 2 in the code. Both sum.write(...) and cout.write(...) execute on every invocation, unconditionally — there is no if that could skip them. So both outputs are always a fresh function of the present inputs. There is no path through the body that leaves an output undriven, so neither output can "remember" a stale value. This is what makes the block genuinely combinational.
Three operational details are worth knowing before you continue.
First, the sc_start(SC_ZERO_TIME) after each input change. Writing to a signal does not immediately change what .read() returns — the write is buffered and committed in the update phase. The sc_start(SC_ZERO_TIME) call advances the scheduler by exactly the delta cycles needed to commit the input writes, fire the method, commit its output writes, and reach a stable state — all at the same simulated time. Without it, you would read the outputs before the method had a chance to react to the new inputs, and you would see the result from the previous apply call. The discipline for testing combinational logic is exactly this: write inputs, advance one zero-time step, read outputs.
Second, the absence of dont_initialize(). During the initialization phase the kernel runs every method process whose dont_initialize() was not called, exactly once. For a combinational block this is what we want: it establishes defined outputs at t=0, before any input has changed, so any downstream logic that samples the output at the start of simulation gets a real value rather than a default-constructed one. Calling dont_initialize() on a combinational process is a bug — it leaves the outputs at their default (0 for sc_uint, false for bool) until the first input change. We will see the consequence of that mistake later.
Third, the 5-bit intermediate full. A 4-bit sum plus a 4-bit b plus a carry can reach 15 + 15 + 1 = 31, which needs five bits. Computing into a 5-bit sc_uint<5> captures the carry in bit 4; we then split it: range(3,0) is the sum, [4] is the carry-out. If you computed into a 4-bit type you would lose the carry and cout would always be 0. Widening before adding, then slicing, is the canonical SystemC idiom for capturing carry/overflow out of fixed-width arithmetic.
A common question at this point: "Why one process for both outputs? Couldn't I write sum in one SC_METHOD and cout in another?" You can, and for this tiny adder it would even work — each would list the same three inputs in its sensitivity list. But there is no benefit and a real cost: two processes that read the same inputs duplicate the sensitivity list (two places to forget an input), and if both ever wrote the same output you would hit the implementation-defined "last writer wins" ordering. The convention across RTL modeling — Verilog, VHDL, and SystemC alike — is one combinational process per logical block, computing all of that block's outputs. Start there and you avoid a whole class of multi-driver problems.
Intermediate: How It Really Works
The 4-bit adder was deliberately tiny so the whole thing fit in your head. Real combinational blocks are wider — they take more inputs, select among more operations, and feed their outputs into other blocks. This section does two things: it scales the concept up to a real, recurring worked example (the RV32I ALU), and it looks precisely at when a combinational output becomes visible, because the answer — delta cycles — is what makes a chain of combinational blocks settle correctly.
Worked example: the RV32I ALU
An Arithmetic-Logic Unit is the combinational core of a CPU. It takes two 32-bit operands and an operation select, and produces a 32-bit result plus a zero flag. It is the natural next step up from the adder: same pattern (one SC_METHOD, sensitive to all inputs, every output assigned), wider cone (ten operations instead of one), and one new subtlety that the adder did not have — signed-versus-unsigned type discipline.
We are modeling RV32I, the 32-bit base integer instruction set of RISC-V: ten ALU operations, 32-bit operands, no multiply or divide. RISC-V defines some operations as signed (SLT, SRA) and some as unsigned (SLTU, SRL). Choosing the wrong SystemC type — sc_uint<32> where sc_int<32> is required — produces results that are silently wrong for negative numbers and large unsigned values: the simulation runs, easy tests pass, and the bug surfaces only on inputs with the most-significant bit set. This is the combinational-modeling lesson the ALU teaches that the adder could not: in a wide combinational block, the type of an intermediate value is part of the function f, and getting it wrong is just as much a combinational bug as a missing sensitivity entry.
Here is the complete, self-contained ALU with a directed test harness.
// file: alu.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// alu.cpp -o alu -lsystemc
// Run: LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./alu
#include <systemc.h>
#include <iostream>
#include <iomanip>
// RV32I ALU operation encoding (plain enum, used as the op select value).
enum alu_op_t {
ALU_ADD = 0,
ALU_SUB = 1,
ALU_AND = 2,
ALU_OR = 3,
ALU_XOR = 4,
ALU_SLT = 5, // signed less-than
ALU_SLTU = 6, // unsigned less-than
ALU_SLL = 7, // shift left logical
ALU_SRL = 8, // shift right logical (unsigned, zero-fill)
ALU_SRA = 9 // shift right arithmetic (signed, sign-fill)
};
// Pure combinational ALU. One SC_METHOD, sensitive to a, b, op (Rule 1).
// 'res' is initialized at the top and 'result'/'zero' assigned on every
// path, so neither output is ever left stale (Rule 2).
SC_MODULE(Alu) {
sc_in<sc_uint<32>> a;
sc_in<sc_uint<32>> b;
sc_in<sc_uint<4>> op;
sc_out<sc_uint<32>> result;
sc_out<bool> zero; // true when result == 0 (used by BEQ/BNE)
void compute() {
sc_uint<32> res = 0; // default value -> Rule 2: 'res' is always defined
switch ((int)op.read()) {
case ALU_ADD:
res = a.read() + b.read();
break;
case ALU_SUB:
res = a.read() - b.read();
break;
case ALU_AND:
res = a.read() & b.read();
break;
case ALU_OR:
res = a.read() | b.read();
break;
case ALU_XOR:
res = a.read() ^ b.read();
break;
case ALU_SLT:
// Signed less-than: cast both operands to sc_int<32> so the
// comparison interprets the MSB as a sign bit.
res = ((sc_int<32>)a.read() < (sc_int<32>)b.read()) ? 1 : 0;
break;
case ALU_SLTU:
// Unsigned less-than: default sc_uint<32> comparison.
res = (a.read() < b.read()) ? 1 : 0;
break;
case ALU_SLL:
// Only the low 5 bits of b are the shift amount in RV32I.
res = a.read() << b.read().range(4, 0);
break;
case ALU_SRL:
// Logical right shift: zero-fill from the left (sc_uint semantics).
res = a.read() >> b.read().range(4, 0);
break;
case ALU_SRA:
// Arithmetic right shift: cast to sc_int so >> sign-fills, then
// cast back to unsigned for the result port.
res = (sc_uint<32>)((sc_int<32>)a.read() >> b.read().range(4, 0));
break;
default:
res = 0; // undefined opcode -> defined output (zero), never stale
break;
}
result.write(res);
zero.write(res == 0);
}
SC_CTOR(Alu) {
SC_METHOD(compute);
sensitive << a << b << op; // Rule 1: all three inputs listed
}
};
// ---- Directed test harness ----
static int pass_count = 0;
static int fail_count = 0;
int sc_main(int /*argc*/, char* /*argv*/[]) {
sc_signal<sc_uint<32>> a, b, result;
sc_signal<sc_uint<4>> op;
sc_signal<bool> zero;
Alu dut("dut");
dut.a(a); dut.b(b); dut.op(op); dut.result(result); dut.zero(zero);
auto check = [&](const char* name, unsigned av, unsigned bv, alu_op_t ov,
unsigned exp_res, bool exp_zero) {
a.write(av); b.write(bv); op.write(ov);
sc_start(SC_ZERO_TIME); // settle the combinational method
bool ok = (result.read() == exp_res) && (zero.read() == exp_zero);
std::cout << (ok ? "PASS " : "FAIL ") << std::left << std::setw(5) << name
<< " result=0x" << std::hex << std::setw(8) << std::setfill('0')
<< (uint32_t)result.read()
<< " zero=" << std::dec << zero.read()
<< std::setfill(' ') << "\n";
if (ok) pass_count++; else fail_count++;
};
std::cout << "=== RV32I ALU directed test ===\n";
check("ADD", 5, 3, ALU_ADD, 8, false);
check("ADD", 0xFFFFFFFF, 1, ALU_ADD, 0, true); // wraps to 0
check("SUB", 10, 3, ALU_SUB, 7, false);
check("SUB", 5, 5, ALU_SUB, 0, true); // zero flag
check("AND", 0xFF00FF00, 0x0F0F0F0F, ALU_AND, 0x0F000F00, false);
check("OR", 0xFF000000, 0x00FF0000, ALU_OR, 0xFFFF0000, false);
check("XOR", 0xDEADBEEF, 0xDEADBEEF, ALU_XOR, 0, true); // x^x = 0
check("SLT", 0xFFFFFFFF, 1, ALU_SLT, 1, false); // -1 < 1
check("SLT", 1, 0xFFFFFFFF, ALU_SLT, 0, true); // 1 < -1? no
check("SLTU", 0xFFFFFFFF, 1, ALU_SLTU, 0, true); // big > 1
check("SLTU", 1, 0xFFFFFFFF, ALU_SLTU, 1, false);
check("SLL", 1, 4, ALU_SLL, 16, false); // 1 << 4
check("SLL", 1, 31, ALU_SLL, 0x80000000, false);
check("SRL", 0x80000000, 1, ALU_SRL, 0x40000000, false); // zero-fill
check("SRA", 0x80000000, 1, ALU_SRA, 0xC0000000, false); // sign-fill
std::cout << "=== Results: " << pass_count << " PASS, "
<< fail_count << " FAIL ===\n";
return (fail_count > 0) ? 1 : 0;
}
Expected output:
=== RV32I ALU directed test ===
PASS ADD result=0x00000008 zero=0
PASS ADD result=0x00000000 zero=1
PASS SUB result=0x00000007 zero=0
PASS SUB result=0x00000000 zero=1
PASS AND result=0x0f000f00 zero=0
PASS OR result=0xffff0000 zero=0
PASS XOR result=0x00000000 zero=1
PASS SLT result=0x00000001 zero=0
PASS SLT result=0x00000000 zero=1
PASS SLTU result=0x00000000 zero=1
PASS SLTU result=0x00000001 zero=0
PASS SLL result=0x00000010 zero=0
PASS SLL result=0x80000000 zero=0
PASS SRL result=0x40000000 zero=0
PASS SRA result=0xc0000000 zero=0
=== Results: 15 PASS, 0 FAIL ===
The ALU is the same pattern as the adder, scaled. One SC_METHOD, compute(), sensitive to all three inputs (a, b, op) — Rule 1. The local res is initialized to 0 at the top and the switch always lands on some case (every opcode, plus a default), so res is always defined; then result and zero are written unconditionally — Rule 2. There is no path through the body that leaves an output stale. Drop op from the sensitivity list and the ALU would not recompute when the opcode changes with operands held constant; drop the default arm and an undefined opcode would leave res at whatever the previous compute set it to — both are the same two bugs from the adder, just harder to spot in a wider block.
zero output is not a RISC-V instruction — it is a hardware convenience. The branch unit implements BEQ rs1, rs2 by computing SUB(rs1, rs2) and checking zero: if the difference is zero the registers are equal and the branch is taken. Adding zero to the ALU's combinational cone saves a dedicated comparator. It is still pure combinational logic: zero = (result == 0), a function of the present result.Three type subtleties carry the combinational lesson.
Signed vs unsigned comparison (SLT vs SLTU). sc_uint<32>(0xFFFFFFFF) is the unsigned value 4,294,967,295, which is not less than 1, so SLTU(0xFFFFFFFF, 1) is 0. But 0xFFFFFFFF interpreted as signed is -1, which is less than 1, so SLT(0xFFFFFFFF, 1) is 1. The only difference between the two cases in the code is the (sc_int<32>) cast. The cast changes nothing about the bits — it reinterprets the same 32-bit pattern as two's-complement signed — but it changes the meaning of <. This is the SystemC equivalent of SystemVerilog's $signed(). Forget the cast on SLT and your ALU silently computes SLTU semantics; every signed comparison of a negative number comes out wrong, and no tool warns you.
Logical vs arithmetic right shift (SRL vs SRA). sc_uint<32> >> n zero-fills from the left (logical); sc_int<32> >> n sign-fills from the left (arithmetic). 0x80000000 >> 1 is 0x40000000 logical but 0xC0000000 arithmetic — the difference is whether the vacated MSB is filled with 0 or with a copy of the sign bit. We get arithmetic behavior by casting to sc_int<32> before the shift, then back to sc_uint<32> for the output port.
Shift-amount masking. RV32I shifts use only the low 5 bits of the shift-amount operand (b.range(4, 0)), giving a range of 0–31. This is both a spec requirement and a safety measure: shifting a 32-bit value by 32 or more is undefined behavior in C++. The range(4, 0) slice enforces the hardware rule and removes the UB.
How the combinational output actually becomes visible: delta cycles
The 4-bit adder and the ALU both call sc_start(SC_ZERO_TIME) between writing inputs and reading outputs. That call is not decoration — it is what makes a combinational write visible. Understanding why is the difference between treating SystemC as a black box and reasoning about it precisely, which you need when one combinational block feeds another.
Recall from Part 4 (and the delta-cycle material it depends on) that sc_signal::write does not change the value read() returns immediately. The write is a request-update: the new value is buffered, and only the kernel's update phase commits it. The scheduler runs in alternating phases at a fixed simulated time:
one delta cycle, at a fixed simulated time:
┌──────────────┐ ┌──────────────┐
│ evaluate │ --> │ update │ --> (any value-changed events?)
│ (run │ │ (commit │ if yes: another delta;
│ methods) │ │ writes) │ if no: advance time
└──────────────┘ └──────────────┘
Trace one check() call on the ALU. You call a.write(...), b.write(...), op.write(...) — three buffered writes. Then sc_start(SC_ZERO_TIME):
- Delta 1, update phase: the three input writes commit.
a,b,opnow hold their new values, and their value-changed events fire. - Delta 2, evaluate phase: those events made
compute()runnable. It executes, reads the newa/b/op, computes, and callsresult.write(...)andzero.write(...)— two more buffered writes. - Delta 2, update phase: the output writes commit.
resultandzeronow hold their new values. - The scheduler checks for more runnable processes. Nothing in this design is sensitive to
resultorzero, so the runnable set is empty. No more deltas.sc_start(SC_ZERO_TIME)returns.
Now result.read() returns the settled value. Without the sc_start, you would still be at the moment after the input writes were requested but before they committed — compute() would not have run, and the outputs would carry the value from the previous call. This is the precise reason combinational testing follows "write inputs, advance one zero-time step, read outputs."
The same mechanism explains how a chain of combinational blocks settles. Suppose block X's output feeds block Y's input. Writing X's inputs takes one delta to commit; X runs and writes its output in the next delta; that output commit fires a value-changed event that wakes Y; Y runs and writes its output in the following delta. The whole chain settles across as many deltas as its combinational depth — all at the same simulated time, no clock involved. Each sc_signal hop in a combinational path costs one delta of settling, and sc_start(SC_ZERO_TIME) runs exactly as many deltas as needed to reach stability.
sc_start(SC_ZERO_TIME) (or, in a clocked testbench, sample on the next clock edge) and the stale read disappears. The value was never wrong; you just read it at the wrong delta.This concludes the Intermediate section. You now have a real combinational worked example (the RV32I ALU), the type discipline that wide combinational blocks demand, and a precise account of how combinational outputs become visible across delta cycles. The Advanced section drills into the LRM corners.
Advanced: Edge Cases & LRM Corners
The Beginner and Intermediate sections cover what nearly all combinational code does. This section is the rest — LRM corners that tutorials skip, bugs that take a senior engineer half a day to find, and constructs you should know exist even if you do not reach for them weekly.
Corner 1: the sensitivity-omission stale-output bug
This is Rule 1 violated, and it is the single most common combinational bug in SystemC because the language gives no warning. Take the 4-bit adder and drop cin from the sensitivity list:
// BUGGED: cin is read in the body but missing from the sensitivity list.
SC_CTOR(Adder4) {
SC_METHOD(compute);
sensitive << a << b; // <-- BUG: cin omitted
}
The body still reads cin.read(), so the logic is correct. But the process is no longer woken when cin changes alone. Consider this stimulus:
apply(2, 2, false); // a=2 b=2 cin=0 -> sum=4 cout=0 (correct)
// now change ONLY cin:
cin.write(true);
sc_start(SC_ZERO_TIME);
std::cout << "sum=" << sum.read() << "\n"; // expect 5; BUGGED prints 4
Expected output (bugged version):
a=2 b=2 cin=0 -> sum=4 cout=0
sum=4
The second line should read sum=5 — 2 + 2 + 1 = 5. But because cin is not in the sensitivity list, the write cin.write(true) fires cin's value-changed event, and nothing is listening. The method does not re-run. sum keeps its previous committed value, 4. The output is stale.
What makes this insidious is that it is intermittent. The next time a or b changes — which is in the sensitivity list — the method runs and reads the current cin, so the output snaps to the right value. So a test that always changes a or b alongside cin never sees the bug; only a test that changes cin in isolation exposes it. In a wide block like the ALU, an omitted op produces exactly this: change the opcode with operands held constant and the result does not update; change an operand and it suddenly corrects itself.
The fix is mechanical and absolute: every signal read inside a combinational process must appear in its sensitivity list. Linting tools (and synthesis front-ends, if the model is meant to be synthesizable) catch most omissions. The habit that prevents them is to write the sensitivity list and the reads together, treating << x and x.read() as a matched pair.
always @(*) and no always_comb auto-sensitivity. There is no construct that infers the sensitivity list from the reads in the body. If you are coming from SystemVerilog, this is the habit that will bite you first. Build the list by hand, every time, and list everything.Corner 2: output incompleteness — the latch-inference analogue
This is Rule 2 violated. A combinational output must be a function of the present inputs on all input combinations. If some path through the body fails to assign an output, that output's sc_signal retains its previous value — the block silently acquires memory. In synthesizable RTL this is "latch inference," and synthesis tools warn loudly about it; in SystemC simulation there is no warning at all.
Here is a mux-like block that gets it wrong:
// BUGGED: 'y' is not assigned on every path.
SC_MODULE(Selector) {
sc_in<sc_uint<8>> in0, in1;
sc_in<sc_uint<2>> sel;
sc_out<sc_uint<8>> y;
void compute() {
switch ((int)sel.read()) {
case 0: y.write(in0.read()); break;
case 1: y.write(in1.read()); break;
// case 2 and 3: y is NOT assigned -> y holds its previous value!
}
}
SC_CTOR(Selector) {
SC_METHOD(compute);
sensitive << in0 << in1 << sel;
}
};
When sel is 2 or 3, no y.write(...) executes. y keeps whatever it last held. If sel was 0 (so y == in0) and then sel becomes 2, y stays at the old in0 value — even if in0 later changes, y does not follow it until sel returns to 0 or 1. The block now remembers the operand it captured under sel==0. That is a latch, and it is not what "combinational selector" was supposed to mean.
The two canonical fixes both guarantee y is assigned on every path:
// FIX A: default-assign at the top, then override.
void compute() {
sc_uint<8> out = 0; // every path now has a defined value
switch ((int)sel.read()) {
case 0: out = in0.read(); break;
case 1: out = in1.read(); break;
default: out = 0; break; // explicit default
}
y.write(out); // single, unconditional write
}
// FIX B: explicit default arm in the switch.
void compute() {
switch ((int)sel.read()) {
case 0: y.write(in0.read()); break;
case 1: y.write(in1.read()); break;
default: y.write(0); break; // covers 2 and 3
}
}
Fix A — initialize a local to a default at the top, compute into it, write once at the bottom — is the pattern the ALU uses (sc_uint<32> res = 0; then a single result.write(res)). It is the most robust because it is impossible to leave an output unwritten: there is exactly one write, and it is unconditional. Prefer it for any block with more than a couple of branches.
Corner 3: dont_initialize() on a combinational process
In the Beginner section we noted that a combinational process should not call dont_initialize(). Here is what happens when you do.
// BUGGED:
SC_CTOR(Adder4) {
SC_METHOD(compute);
sensitive << a << b << cin;
dont_initialize(); // <-- BUG on a combinational process
}
During the initialization phase, the kernel normally runs every method whose dont_initialize() was not called, exactly once. That initial run is what gives a combinational block defined outputs at t=0. With dont_initialize(), the initial run is suppressed. So at the start of simulation sum and cout sit at their default-constructed values (0 and false) regardless of what the inputs are — until the first input change wakes the method.
Concretely: if you bind a=3, b=4, cin=0 before sc_start and then read the outputs before changing any input, the bugged version reports sum=0, cout=0 instead of sum=7. The block looks broken for exactly one evaluation, then corrects itself on the first input change. In a larger design this surfaces as "the datapath produces garbage for the first cycle after start, then settles" — a subtle, frustrating class of bug. The fix is one line: remove dont_initialize() from combinational processes. (It belongs on clocked processes, which is the subject of the next part in this section.)
Corner 4: multiple writers to one combinational signal — "last writer wins"
Suppose, against the one-process-per-block convention, you split a combinational output across two processes that both write the same signal in the same evaluate phase.
// BUGGED two-driver pattern: both write 'y' in the same evaluate phase.
void drive_from_a() { if (sel_a.read()) y.write(VAL_A); }
void drive_from_b() { if (sel_b.read()) y.write(VAL_B); }
When both processes run in the same evaluate phase and both call y.write(...), the kernel commits one of the two buffered values in the update phase — and per IEEE 1666-2011, which one is implementation-defined when independent processes write the same sc_signal in the same evaluation. The Accellera reference kernel resolves this by a particular ordering, but a different kernel may choose differently, and a randomized-for-stress kernel makes it nondeterministic. Worse, sc_signal is a single-writer channel by design; some configurations will flag a multiple-driver violation outright. Either way, the behavior is not something you should rely on.
The fix is the convention this post has stated from the start: one combinational process computes all of a block's outputs. Merge the two drivers into one method with a single, unconditional write:
// FIXED: one process, one writer, deterministic.
void compute() {
sc_uint<32> out = VAL_DEFAULT;
if (sel_a.read()) out = VAL_A;
else if (sel_b.read()) out = VAL_B;
y.write(out); // single writer
}
One process, one writer, one unconditional write — deterministic on every kernel, and it folds the priority between sel_a and sel_b into an explicit if/else chain instead of leaving it to scheduler ordering. Even when a block has a dozen inputs feeding one output, the single-process pattern scales: the if/else or switch grows, but there is still exactly one writer.
Corner 5: combinational loops — the one thing that is not allowed
Everything in this post assumes a combinational block has no feedback through itself. That assumption is load-bearing. If a combinational process's output feeds, directly or through other combinational blocks, back into its own input, you create a combinational loop — and the SystemC scheduler cannot settle it.
// PATHOLOGICAL: y depends on y.
void compute() {
y.write(y.read() ^ 1); // output is a function of itself
}
// with: sensitive << y;
Each time y changes, the method wakes, computes a different value, writes it, that write fires another value-changed event, the method wakes again... The runnable set never empties. sc_start(SC_ZERO_TIME) runs delta after delta at the same simulated time and never returns — the simulation hangs, spinning in delta cycles. This is the SystemC manifestation of a combinational loop in hardware (a ring oscillator, which real synthesis tools also reject). The lesson: combinational logic must be acyclic. Feedback belongs only around a storage element — a register — which breaks the loop into "this cycle's output depends on last cycle's value." That is sequential logic, and it is the subject of the next part. If your sc_start(SC_ZERO_TIME) hangs, suspect a combinational loop first.
Version differences
The kernel behavior described here — SC_METHOD semantics, static-sensitivity mechanics, sc_signal request-update and delta-cycle settling, and the initialization behavior governed by dont_initialize() — is identical across SystemC 2.3.1, 2.3.3, and 2.3.4. None of these mechanisms changed across those releases. Differences between them are confined to the build system, C++17 conformance, TLM-2.0 utilities, and sc_vector helpers. Every example in this post compiles and behaves the same on any 2.3.x release.
Hands-on exercise
Build a combinational 4-bit barrel shifter from first principles. Inputs: a 4-bit value data, a 2-bit shift amount shamt, and a 1-bit direction dir (0 = left, 1 = right). Output: a 4-bit result. The shifter shifts data by shamt positions in the direction dir, zero-filling the vacated bits (logical shift). Build it as one SC_METHOD sensitive to all three inputs, with result assigned on every path.
Drive it directly from sc_main with the write-settle-read pattern: for each test, write data, shamt, dir, call sc_start(SC_ZERO_TIME), and print result. Predict every line of output before you run it. Verify that shifting 0b0001 left by 2 gives 0b0100, shifting 0b1000 right by 3 gives 0b0001, and shifting by 0 in either direction returns data unchanged.
When that works, extend it. Add an arith input (1 bit): when dir is right and arith is 1, do an arithmetic right shift that sign-fills from bit 3 instead of zero-filling. Reuse the ALU's trick: cast data to sc_int<4> for the arithmetic case so >> sign-extends, then cast back. Predict what 0b1000 right-shifted arithmetically by 1 produces (sign bit is 1, so the vacated bit fills with 1: 0b1100) and verify.
Then deliberately break it two ways and watch the failure modes you learned to recognize. First, remove shamt from the sensitivity list, hold data and dir constant, change only shamt, and confirm the output goes stale (Corner 1). Second, add a case to the direction logic that forgets to assign result, drive dir into that case, and confirm result retains its previous value (Corner 2). Fix both, and you have internalized the two rules by feeling them fail.
No solution is provided. The lessons live in the building.
Hints
- One
SC_METHOD,compute(), withsensitive << data << shamt << dir << arith(listarithonly once you add it). - Initialize a local —
sc_uint<4> out = 0;— at the top, compute into it, and writeresultonce at the bottom. This guarantees Rule 2 by construction. - Use
data.read().range(...)and the<</>>operators; remember that for a 4-bit value the meaningful shift amounts are 0–3, which a 2-bitshamtcovers exactly. - For the arithmetic case, the cast pair is
(sc_uint<4>)((sc_int<4>)data.read() >> shamt.read())— exactly the SRA idiom from the ALU, narrowed to 4 bits. - There is no clock anywhere in this design. If you find yourself reaching for
sc_clockorclk.pos(), stop — combinational logic has no clock.
Common mistakes
- Omitting an input from the sensitivity list. The body reads a signal that is not in
sensitive << ..., so the process does not wake when that signal changes alone. The output goes stale until some other listed input changes and incidentally re-runs the body. The bug is intermittent and silent — the simulation runs, the waveforms look plausible, and only a test that changes the missing input in isolation exposes it. Fix: every signal you.read()must appear in the sensitivity list. Treat<< xandx.read()as a matched pair.
- Failing to assign an output on every path. A
switchorifthat leaves an output unassigned on some branch causes that output'ssc_signalto retain its previous value — the block silently behaves like a latch. This is the SystemC analogue of unintended latch inference, except there is no warning. Fix: initialize a local to a default at the top of the body, compute into it, and write the output once unconditionally (the ALU'ssc_uint<32> res = 0; ... result.write(res);pattern). Or give everyswitchan explicitdefaultarm.
- Calling
dont_initialize()on a combinational process. This suppresses the initialization run that establishes defined outputs at t=0, so the outputs sit at their default-constructed value until the first input change. The block produces garbage for the first evaluation, then corrects itself — a subtle "first-cycle wrong" bug. Fix: never calldont_initialize()on a combinational process; it belongs only on clocked processes.
- Using the wrong signedness type.
sc_uint<32>andsc_int<32>differ in comparison and right-shift semantics. Usingsc_uintwhere the spec wants signed behavior (RV32ISLT,SRA) silently produces wrong results for negative numbers and MSB-set values — the simulation runs, easy tests pass, and the bug hides until a negative operand appears. Fix: cast tosc_int<N>exactly where signed semantics are required (the SystemC equivalent of$signed()), and cast back tosc_uint<N>for the output port.
- Reading a combinational output before the deltas settle. Writing inputs does not immediately update what
.read()returns; the writes commit in the update phase, and the recompute happens in the next delta. Reading the output without first advancing the scheduler returns the previous value. Fix: in a directed test, callsc_start(SC_ZERO_TIME)between writing inputs and reading outputs; in a clocked testbench, sample on the next clock edge.
- Creating a combinational loop. If a combinational output feeds back into its own input (directly or through other combinational blocks), the scheduler never reaches a stable state and
sc_startspins in delta cycles forever. Fix: keep combinational logic acyclic. Feedback belongs only around a register — that is sequential logic, covered in the next part.
Recap
After working through this post you can now:
- State the definition of combinational logic (outputs are a pure function of present inputs; no clock, no state) and recognize it by its signature: inputs in, outputs out, no clock pin.
- Map a combinational block to exactly one
SC_METHODmade statically sensitive to all of its inputs, computing all of its outputs. - Apply Rule 1 (sensitivity completeness — list every input you read) and Rule 2 (output completeness — assign every output on every path), and explain why violating either produces a silent, simulation-passes-but-wrong bug.
- Build a combinational block from scratch — a 4-bit adder, an RV32I ALU, a barrel shifter — without consulting reference code.
- Reason about how a combinational output becomes visible across delta cycles, why
sc_start(SC_ZERO_TIME)is required between writing inputs and reading outputs, and how a chain of combinational blocks settles at a single simulated time. - Apply signed-versus-unsigned type discipline (
sc_intvssc_uint) for operations whose semantics depend on signedness, and the shift-amount masking idiom for RV32I shifts. - Diagnose and fix the combinational bugs that account for most "the model is wrong but the simulation runs fine" reports: sensitivity omission, output incompleteness, mis-placed
dont_initialize(), wrong signedness, and combinational loops.
Further reading
Standards
- IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §4.2 (the scheduler: evaluate and update phases, delta cycles), §5.2.11 (
SC_METHODsemantics), §5.2.18–§5.2.20 (static sensitivity,sensitive,dont_initialize), §6.4 (sc_signaland the request-update model). - RISC-V Foundation, The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA, RV32I base integer instruction set — operation definitions, signed/unsigned semantics for SLT/SLTU/SRA/SRL, and the 5-bit shift-amount rule used in the ALU worked example.
Vendor and consortium documents
- Accellera Systems Initiative, SystemC 2.3.x User Guide — process kinds and primitive-channel mechanics.
Textbooks
- Bhasker, A SystemC Primer (2nd ed.), ch. 4 —
SC_METHODfor combinational logic and sensitivity-list construction. - Grötker, Liao, Martin, and Swan, System Design with SystemC, ch. 3–4 — primitive channels and the request-update execution model.
- Doulos, SystemC Golden Reference Guide, ch. 4 —
SC_METHODandsensitiveconventions.
Training notes
- MIT 6.004 lecture notes — combinational versus sequential RTL pedagogy, and the latch-inference failure mode from incomplete output assignment.
Next in this section
→ Part 2: Decoders & Case Analysis — building combinational decoders with switch/case, the RV32I instruction decoder as the worked example after a generic opcode decoder is taught from first principles, and the output-completeness discipline applied to wide case structures. Read it here: 9. SystemC Tutorial — Instruction Decoder.
Comments (0)
Leave a Comment