10. SystemC Tutorial - Registered & Sequential Logic

Why this matters

Rewritten 2026-06-05 with deeper first-principles material.

Every digital system that does anything across time has memory inside it, and that memory is built from registers. A register is the single most important storage primitive in synchronous design — it is the thing that lets a circuit remember what happened on the last clock edge so it can do something different on the next one. The program counter that picks the next instruction in the CPU you are reading this on is a register. The accumulator in a DSP filter is a register. The credit counter in a PCIe flow-control engine, the sequence number in a TCP offload block, the retry count in an SD-card controller, the brightness level that fades your laptop's keyboard backlight after thirty idle seconds — all registers, all updated on a clock edge, all holding a value steady between edges so the rest of the logic has something stable to compute from. Without registers there is no state, and without state there is no finite state machine, no pipeline, no counter, no memory, no CPU. Combinational logic computes; registers remember. Everything sequential in hardware is some arrangement of combinational logic feeding registers feeding combinational logic.

Modeling registers correctly in SystemC is therefore the hinge on which the entire RTL-patterns section turns. The combinational post before this one taught you logic that has no memory — outputs are a pure function of present inputs. This post adds the one thing that makes hardware hardware: a storage element that captures a value on a clock edge and holds it until the next one. The post that follows this one — finite state machines — is nothing but a register (holding the state) wrapped in combinational logic (computing the next state and the outputs). If you do not have rock-solid intuition for what a register is, when it samples, what it holds between edges, and how sc_signal's deferred-update mechanic gives you flip-flop behavior for free, then every FSM, every pipeline, and every memory you build later will be guesswork. Get it right here, once, from first principles, and the rest of the section is filling in combinational detail around a storage skeleton you already understand cold. That is the bar for this post: by the end you will be able to write a clocked register in SystemC from scratch, predict the exact cycle on which it captures a value and the exact cycle on which that value becomes visible downstream, choose between synchronous and asynchronous reset and justify the choice, and read a VCD waveform and explain every edge in it. We will establish all of that on a four-bit counter first — the simplest interesting register — and only then apply it to the RISC-V program counter and instruction fetch, which is the same pattern at CPU scale.

Prerequisites

  • Part 1 — Modules, Ports & Signals. You need to be comfortable declaring SC_MODULE, binding sc_signal channels to sc_in / sc_out ports, and registering a process inside SC_CTOR.
  • Part 2 — Simulation Time & Clocks. You need to know how sc_clock generates a periodic bool waveform, what sc_time_stamp() reports, and how sc_start() advances simulated time. Registers in this post all clock off an sc_clock.
  • Part 3 — Delta Cycles & Event-Driven Semantics. You need to understand that an sc_signal::write is buffered, that the new value commits in the next update phase, and that the value-changed event then wakes any process statically sensitive to that signal. This deferred-update behavior is the register's storage; the whole post depends on it.
  • Part 4 — Processes & Sensitivity (SC_METHOD vs SC_THREAD). You need to know the difference between an SC_METHOD and an SC_THREAD, why a clocked register is naturally an SC_METHOD sensitive to clk.pos(), and how a static sensitivity list is built with sensitive << clk.pos() versus sensitive << count << reset.
  • SystemC 2.3.x installed. All examples compile with a C++17 compiler (g++ 9 or newer, or clang++ 10 or newer) against any 2.3.x SystemC build.

Three ideas from earlier parts do the heavy lifting here: sc_signal<T> (the buffered channel that becomes your register's storage), SC_METHOD (the process kind a clocked register uses), and the update phase (the kernel step that makes a buffered write look like a clock-to-Q delay). If any of those is fuzzy, re-read the relevant part before continuing.

Mental model (first principles)

Forget SystemC for a moment and think about hardware. A register (more precisely, a bank of edge-triggered D flip-flops) does exactly one thing: on the rising edge of its clock, it samples the value on its D input and copies it to its Q output. Between clock edges, Q holds steady no matter what D does — the register has captured a value and is now remembering it. That sampling-then-holding is the whole behavior. Three consequences follow, and they are the three things every sequential-logic bug ultimately traces back to:

  1. The register samples D only at the clock edge. Wiggle D all you like between edges; Q ignores it. Only the value present at the instant of the edge gets captured.
  2. Q holds between edges. The stored value is stable for the entire clock period. Downstream combinational logic gets a steady input to compute from.
  3. There is a clock-to-Q delay. The new Q does not appear at the exact edge instant; it appears a tiny moment after. This is what stops a chain of registers from racing — register B samples register A's old Q, because A's new Q has not propagated yet when the shared edge fires.

Now map that onto SystemC. There is no flip-flop primitive in the language. Instead, the storage element is an ordinary sc_signal<T>, and the clock-to-Q delay is the kernel's update-phase deferral that you met in Part 3. Here is the load-bearing insight of the entire post: when you write() an sc_signal during the evaluate phase, the new value does not become readable until the next update phase. A register is therefore just a clocked SC_METHOD whose body is essentially one line:

on each rising clock edge:
    state.write(  reset ? RESET_VALUE : next_state.read()  );

That single deferred write is the flip-flop. When the clock edge fires, the method reads next_state (the value to capture) and writes it into state. The write is buffered. The update phase commits it. The value-changed event fires. Every process sensitive to state wakes in the next delta and sees the new value. The one-delta gap between "the edge fired and the register wrote" and "downstream logic can read the new value" is precisely the clock-to-Q delay — and it is what makes registers behave correctly without any special handling. You get non-blocking-assignment semantics, the same thing SystemVerilog's <= gives you, completely for free from sc_signal's update mechanic.

The structure of every registered block is the same two-piece arrangement:

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
flowchart LR
    IN([inputs]) --> COMB[next-state logic\ncombinational]
    COMB -->|next_state| REG[state register\nsc_signal + clocked SC_METHOD]
    CLK([clk.pos]) --> REG
    REG -->|state| COMB
    REG --> OUT([outputs])

Read the diagram as a loop broken by the register. Combinational logic reads the current state (plus any inputs) and computes next_state. The register captures next_state on the clock edge and presents it as the new state. The new state feeds back into the combinational logic, which computes a fresh next_state for the following edge. The feedback path looks like a combinational loop but it is not — the register breaks it, because state only changes on a clock edge, never mid-cycle. This register-plus-combinational split is the exact skeleton of a finite state machine; the FSM post that follows simply names the register's value "the state" and makes the combinational block a switch over states. Internalize the split here and the FSM post is a short walk.

Note The single most common mental error for newcomers is to expect state.write(x) followed by state.read() in the same process to return x. It does not. The read returns the old value until the next update phase. That is not a bug — it is the exact behavior that makes a register hold its value across an edge. We will see this concretely in the counter trace below.

With the mental model in place, we build it.

Beginner: First Principles

The simplest register that does something visible is a counter. We will build a four-bit up-counter: it holds a value, and on every rising clock edge it increments by one, wrapping from 15 back to 0. It has an active-high synchronous reset that forces the count to zero. Four bits, sixteen states, one clocked process — small enough to run in your head, rich enough to show every register concept.

Here is the complete, self-contained program.

// file: counter4.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            counter4.cpp -o counter4 -lsystemc
// Run:   LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./counter4

#include <systemc.h>
#include <iostream>

SC_MODULE(Counter4) {
  sc_in<bool>        clk;     // clock — register samples on rising edge
  sc_in<bool>        reset;   // active-high SYNCHRONOUS reset
  sc_out<sc_uint<4>> count;   // current count value (0..15)

  // The state register's storage. This sc_signal IS the flip-flop bank.
  sc_signal<sc_uint<4>> count_reg;

  // --- Clocked register process: sensitive ONLY to clk.pos() ---
  // This is the entire register. One deferred write per rising edge.
  void reg_proc() {
    if (reset.read()) {
      count_reg.write(0);                       // synchronous reset to 0
    } else {
      count_reg.write(count_reg.read() + 1);    // increment, wraps at 16
    }
  }

  // --- Drive the output port from the register ---
  void output_proc() {
    count.write(count_reg.read());
  }

  SC_CTOR(Counter4) {
    SC_METHOD(reg_proc);
    sensitive << clk.pos();   // clocked: wake only on rising edge
    dont_initialize();        // do NOT run before the first real edge

    SC_METHOD(output_proc);
    sensitive << count_reg;   // combinational: follow the register
  }
};

SC_MODULE(Monitor) {
  sc_in<bool>        clk;
  sc_in<sc_uint<4>>  count;

  void watch() {
    std::cout << "[" << sc_time_stamp() << "] count = "
              << count.read() << "\n";
  }

  SC_CTOR(Monitor) {
    SC_METHOD(watch);
    sensitive << clk.pos();
    dont_initialize();
  }
};

int sc_main(int /*argc*/, char* /*argv*/[]) {
  // 10 ns period, 50% duty. The start_time=5 ns / posedge_first=false phasing
  // puts the FIRST rising edge at t=10 ns (clock starts high, falls at 5 ns,
  // rises at 10 ns), so reset can be held across a clean first edge.
  sc_clock           clk("clk", 10, SC_NS, 0.5, 5, SC_NS, false);
  sc_signal<bool>    reset;
  sc_signal<sc_uint<4>> count;

  Counter4 dut("dut");
  dut.clk(clk);
  dut.reset(reset);
  dut.count(count);

  Monitor mon("mon");
  mon.clk(clk);
  mon.count(count);

  // Hold reset high for one clock period, then release.
  reset.write(true);
  sc_start(10, SC_NS);
  reset.write(false);
  sc_start(180, SC_NS);

  sc_stop();
  return 0;
}

Run it and you will see the count climb from 0, increment once per clock edge, and wrap past 15 back to 0. Read the output before the walkthrough and try to account for every line yourself.

Expected output:

[10 ns] count = 0
[20 ns] count = 1
[30 ns] count = 2
[40 ns] count = 3
[50 ns] count = 4
[60 ns] count = 5
[70 ns] count = 6
[80 ns] count = 7
[90 ns] count = 8
[100 ns] count = 9
[110 ns] count = 10
[120 ns] count = 11
[130 ns] count = 12
[140 ns] count = 13
[150 ns] count = 14
[160 ns] count = 15
[170 ns] count = 0
[180 ns] count = 1

Now the walkthrough, edge by edge, because every sequential bug you will ever chase is a misunderstanding of one of these steps.

The clock period is 10 ns with 50% duty, so rising edges land at 10, 20, 30, … ns. From t=0 to t=10 ns the testbench holds reset high. At the first rising edge, t=10 ns, the register process reg_proc() wakes (it is sensitive only to clk.pos()). It reads reset — still high — so it writes count_reg = 0. That write is buffered; it does not take effect yet. The kernel finishes the evaluate phase, runs the update phase, and count_reg becomes 0, firing its value-changed event. In the next delta output_proc() wakes (sensitive to count_reg) and copies 0 to the count port. The monitor, also clocked at t=10 ns, prints count = 0.

At the second rising edge, t=20 ns, reset is now low. reg_proc() reads count_reg — which is 0 — and writes count_reg = 0 + 1 = 1. Here is the subtle and essential part: the read count_reg.read() returns the old value (0), not some value being computed during this edge. The write of 1 is buffered and commits in the update phase. The monitor prints count = 1. The register captured the increment of the value it held before the edge — exactly flip-flop behavior.

This continues every 10 ns: 2, 3, 4, … up to 15 at t=160 ns. At the edge at t=170 ns, reg_proc() reads count_reg = 15 and writes 15 + 1. Because sc_uint<4> is four bits wide, 15 + 1 wraps to 0. The counter rolls over. The monitor prints count = 0. The wrap is not special-cased anywhere in the code — it falls out of the fixed-width arithmetic of sc_uint<4>, exactly as a four-bit hardware counter wraps.

Three details in that program deserve a hard look, because they are the difference between a register that works and one that mysteriously does not.

Detail 1 — sensitive << clk.pos() makes it a register. The register process wakes only on the rising clock edge. That is what makes it sequential. If you instead wrote sensitive << count_reg, the process would wake whenever count_reg changed — and since it writes count_reg, it would wake itself in an infinite combinational loop, incrementing as fast as the kernel can run, with no relationship to the clock. The clock edge in the sensitivity list is the thing that says "sample once per cycle." Memorize the shape: a register is a method sensitive to clk.pos() and nothing else (except possibly an async reset, covered below).

Detail 2 — dont_initialize() is mandatory on the register. Every SC_METHOD is, by default, run once during the initialization phase — before any clock edge. For a combinational method that is exactly what you want (it establishes outputs at t=0). For a clocked register it is wrong: you do not want the register to "tick" before the first real clock edge. dont_initialize() suppresses that initial run. Omit it and the register fires once at t=0 with no edge, reading count_reg's default-constructed value and writing back — usually harmless for a counter that resets to 0, but a latent bug for any register whose reset value is not the default. Always call dont_initialize() on a clocked process. There is no exception.

Detail 3 — the deferred read is what makes it a register, not an adder. Look again at count_reg.write(count_reg.read() + 1). The read() returns the committed value from before this edge; the write() schedules the new value for after this edge. Read-before-write across the evaluate/update boundary is the entire mechanism. If sc_signal updated immediately on write(), this one line would be an infinite increment within a single delta. It does not, because §6.4 of the standard says writes are deferred to the update phase — so the line reads the old value once, computes one increment, and commits it once. That is one clock-to-Q transfer per edge. This is the same reason SystemVerilog uses <= (non-blocking) for register transfers and = (blocking) for combinational temporaries.

Note Notice we used a separate output_proc() to drive the count port from the internal count_reg signal, rather than writing the port directly inside reg_proc(). For this tiny counter you could write the port directly. We split them to foreshadow the register-plus-combinational structure that the rest of the section depends on: the register holds state in an internal sc_signal, and separate logic reads that state. Keeping the stored state in an internal signal (not a port) is the habit that scales to FSMs and pipelines.

Why the internal sc_signal and not the port?

You might ask why count_reg exists at all — why not register directly into the count output port? Two reasons. First, an sc_out port is a binding to an external sc_signal; reading a port you also write inside the same process is legal but muddies the "this signal is my state" intent. Second, and more importantly, real registered blocks compute their next value from the current state, and you want one named, internal signal that unambiguously is the state. When the block grows — a counter that also has a load input, an FSM with five states, a PC with a four-way next-value mux — having state/count_reg as a private sc_signal is what keeps the next-state logic readable. Start with the internal-signal habit on day one and you never refactor.

Intermediate: How It Really Works

The Beginner counter packed the register and its next-value logic into one process: count_reg.write(count_reg.read() + 1) does the "compute next" and the "store it" in one line. That is fine for a counter because the next value is a trivial function of the current value. Real registered blocks separate the two concerns — a combinational process computes the next value, a clocked process stores it — and they make a deliberate choice about reset behavior. This section covers both: the register-plus-combinational split, and synchronous versus asynchronous reset.

Splitting the register from its next-value logic

Here is the same four-bit counter, restructured into the two-process pattern that every FSM, pipeline stage, and datapath register uses. The combinational process computes next_count; the clocked process stores it.

// file: counter4_split.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            counter4_split.cpp -o counter4_split -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(Counter4Split) {
  sc_in<bool>        clk;
  sc_in<bool>        reset;     // active-high synchronous reset
  sc_in<bool>        enable;    // count only when high
  sc_out<sc_uint<4>> count;

  sc_signal<sc_uint<4>> count_reg;    // the state
  sc_signal<sc_uint<4>> next_count;   // combinational next value

  // --- Combinational: compute the next count ---
  void next_logic() {
    if (enable.read()) next_count.write(count_reg.read() + 1);
    else               next_count.write(count_reg.read());   // hold
  }

  // --- Clocked register: store next_count (or reset) on the edge ---
  void reg_proc() {
    if (reset.read()) count_reg.write(0);
    else              count_reg.write(next_count.read());
  }

  // --- Drive output ---
  void out_proc() { count.write(count_reg.read()); }

  SC_CTOR(Counter4Split) {
    SC_METHOD(next_logic);
    sensitive << count_reg << enable;   // every input it READS is listed

    SC_METHOD(reg_proc);
    sensitive << clk.pos();
    dont_initialize();

    SC_METHOD(out_proc);
    sensitive << count_reg;
  }
};

SC_MODULE(Monitor) {
  sc_in<bool>       clk;
  sc_in<bool>       enable;
  sc_in<sc_uint<4>> count;
  void watch() {
    std::cout << "[" << sc_time_stamp() << "] en=" << enable.read()
              << " count = " << count.read() << "\n";
  }
  SC_CTOR(Monitor) {
    SC_METHOD(watch);
    sensitive << clk.pos();
    dont_initialize();
  }
};

int sc_main(int, char*[]) {
  sc_clock              clk("clk", 10, SC_NS, 0.5, 5, SC_NS, false);  // first posedge at 10 ns
  sc_signal<bool>       reset, enable;
  sc_signal<sc_uint<4>> count;

  Counter4Split dut("dut");
  dut.clk(clk); dut.reset(reset); dut.enable(enable); dut.count(count);

  Monitor mon("mon");
  mon.clk(clk); mon.enable(enable); mon.count(count);

  // All input changes happen MID-CYCLE (5 ns off the edges) to avoid races
  // with the clock edge. The monitor is clocked, so it prints the count value
  // that was registered on the PREVIOUS edge — a one-cycle read lag.
  reset.write(true);  enable.write(true);
  sc_start(15, SC_NS);           // through the t=10 ns edge (reset), now at t=15
  reset.write(false);            // deassert reset mid-cycle
  sc_start(50, SC_NS);           // edges 20..60: count climbs ; now at t=65
  enable.write(false);           // disable mid-cycle
  sc_start(30, SC_NS);           // edges 70,80,90: hold ; now at t=95
  enable.write(true);            // re-enable mid-cycle
  sc_start(30, SC_NS);           // edges 100,110,120: resume ; now at t=125

  sc_stop();
  return 0;
}

Expected output:

[10 ns] en=1 count = 0
[20 ns] en=1 count = 0
[30 ns] en=1 count = 1
[40 ns] en=1 count = 2
[50 ns] en=1 count = 3
[60 ns] en=1 count = 4
[70 ns] en=0 count = 5
[80 ns] en=0 count = 5
[90 ns] en=0 count = 5
[100 ns] en=1 count = 5
[110 ns] en=1 count = 6
[120 ns] en=1 count = 7

The split changes nothing about the register's timing — it still ticks once per edge — but it adds an enable input and makes the structure explicit. Walk it: the combinational next_logic() is sensitive to count_reg and enable; whenever either changes it recomputes next_count. When enable is high, next_count = count_reg + 1; when low, next_count = count_reg (hold). The clocked reg_proc() ignores all of that between edges and, on each rising edge, copies next_count into count_reg (unless reset). Reset is held across the t=10 ns edge (count register forced to 0) and deasserted at t=15 ns, so the first increment lands on the t=20 ns edge. From t=70 to t=90 ns enable is low, so next_count equals count_reg, the stored value does not change, and the count holds. At t=100 ns onward enable is high again and counting resumes.

One subtlety in the printed numbers: the monitor is itself clocked, and it reads count (a registered output exposed through a port) on the same edge the register updates. Because the port's new value commits one delta after the edge, the monitor sees the value registered on the previous edge — a consistent one-cycle read lag. That is why the printed count appears to trail the internal count_reg by one line. It is not a bug; it is the same clock-to-Q deferral seen from the consumer's side, and it is exactly how a downstream clocked block would sample this counter's output in a real pipeline.

Two things to lock in from this structure:

The combinational process lists every signal it reads. sensitive << count_reg << enable. SystemC has no always @(*) auto-sensitivity. If you forget << enable, the process never re-evaluates when enable changes, the hold logic silently breaks, and the counter keeps counting through the disabled window. Every read in a combinational method must correspond to one entry in its sensitivity list. This rule returns, with teeth, in the FSM post.

The clocked process lists only clk.pos(). It reads reset and next_count inside the body, but those are not in its sensitivity list — because the register must sample them only at the edge, not react to them mid-cycle. This is the synchronous-reset signature: reset is read inside a clk.pos()-sensitive process, so it only takes effect on an edge. That is the topic of the next subsection.

Synchronous reset: the section default

In the counter above, reset is read inside reg_proc(), whose sensitivity list is clk.pos() only. That makes the reset synchronous: asserting reset does nothing until the next rising clock edge, at which point the register samples reset high and writes 0. Synchronous reset is the default for this section, and for most modern ASIC flows, because the reset path goes through the same flip-flop sampling as ordinary data — it is easy to time, easy to reason about, and immune to glitches on the reset line between edges.

The waveform signature of synchronous reset: assert reset at any time between edges and the register's output does not change until the next edge. In the Beginner counter, reset was held high across the t=10 ns edge, so the register sampled it high at t=10 ns and produced 0. If you had asserted reset at t=12 ns and deasserted it at t=18 ns — entirely between the t=10 and t=20 ns edges — the register would never have sampled it high, and the reset would have had no effect at all. A synchronous reset that is not present at a clock edge is invisible. That is a feature: glitches on the reset line cannot disturb the register.

Asynchronous reset: the labeled variant

Sometimes you need a reset that takes effect immediately, the instant it asserts, without waiting for a clock edge — for example, a power-on reset that must force every register to a known state before the clock is even running. That is an asynchronous reset. It changes the sensitivity list: the register process becomes sensitive to both the clock edge and the reset signal, and checks reset first.

Here is the four-bit counter with asynchronous active-high reset, clearly labeled as the non-default variant.

// file: counter4_async.cpp  --- ASYNC RESET VARIANT (not the section default) ---
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            counter4_async.cpp -o counter4_async -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(Counter4Async) {
  sc_in<bool>        clk;
  sc_in<bool>        reset;    // active-high ASYNCHRONOUS reset
  sc_out<sc_uint<4>> count;

  sc_signal<sc_uint<4>> count_reg;

  void reg_proc() {
    if (reset.read()) {
      count_reg.write(0);                     // takes effect the moment reset asserts
    } else {
      count_reg.write(count_reg.read() + 1);  // otherwise increment on the edge
    }
  }

  void out_proc() { count.write(count_reg.read()); }

  SC_CTOR(Counter4Async) {
    SC_METHOD(reg_proc);
    sensitive << clk.pos() << reset;   // <-- reset IS in the sensitivity list now
    dont_initialize();

    SC_METHOD(out_proc);
    sensitive << count_reg;
  }
};

SC_MODULE(Monitor) {
  sc_in<bool>       clk;
  sc_in<sc_uint<4>> count;
  void watch() {
    std::cout << "[" << sc_time_stamp() << "] count = " << count.read() << "\n";
  }
  SC_CTOR(Monitor) {
    SC_METHOD(watch);
    sensitive << clk.pos();
    dont_initialize();
  }
};

int sc_main(int, char*[]) {
  sc_clock              clk("clk", 10, SC_NS, 0.5, 5, SC_NS, false);  // first posedge at 10 ns
  sc_signal<bool>       reset;
  sc_signal<sc_uint<4>> count;

  Counter4Async dut("dut");
  dut.clk(clk); dut.reset(reset); dut.count(count);

  Monitor mon("mon");
  mon.clk(clk); mon.count(count);

  reset.write(true);
  sc_start(10, SC_NS);
  reset.write(false);
  sc_start(35, SC_NS);          // count up: 1,2,3
  // Assert reset asynchronously BETWEEN clock edges, at t=45 ns.
  reset.write(true);
  sc_start(2, SC_NS);           // reset fires immediately, not on an edge
  reset.write(false);
  sc_start(35, SC_NS);          // resume counting

  sc_stop();
  return 0;
}

Expected output:

[10 ns] count = 0
[20 ns] count = 1
[30 ns] count = 2
[40 ns] count = 3
[50 ns] count = 1
[60 ns] count = 2
[70 ns] count = 3
[80 ns] count = 4

The difference shows at the reset pulse. The counter reaches 3 by t=40 ns. At t=45 ns — between the t=40 and t=50 ns edges — the testbench asserts reset. Because the register is sensitive to reset, reg_proc() wakes immediately at t=45 ns (no clock edge involved), sees reset high, and writes count_reg = 0. The reset took effect mid-cycle. Then at the t=50 ns edge, reset has already deasserted, so the register increments the freshly-reset 0 to 1, and the monitor prints count = 1 rather than the 4 it would have printed without the async reset. Compare to the synchronous version, where a reset pulse entirely between edges would have been invisible.

Note Asynchronous reset is powerful but has real costs in synthesis: the reset becomes part of the timing on a different path, reset deassertion must be synchronized to avoid metastability, and glitches on the reset line corrupt state immediately. Modern ASIC methodology overwhelmingly prefers synchronous reset, reserving async reset for the global power-on case. That is why this section defaults to synchronous reset and treats async as the labeled exception. Use async only when you have a specific reason and you have handled deassertion synchronization.

sc_trace: getting waveform evidence

Printed monitor lines are fine for small examples, but the moment a design has more than a couple of signals you want a waveform. SystemC writes VCD (Value Change Dump) files through sc_trace, and any VCD viewer — GTKWave is the common free one — renders them. Here is the split counter again, instrumented to dump a VCD instead of (or alongside) printing.

// file: counter4_trace.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            counter4_trace.cpp -o counter4_trace -lsystemc
// After running, open counter4.vcd in GTKWave.

#include <systemc.h>

SC_MODULE(Counter4) {
  sc_in<bool>        clk;
  sc_in<bool>        reset;
  sc_out<sc_uint<4>> count;
  sc_signal<sc_uint<4>> count_reg;

  void reg_proc() {
    if (reset.read()) count_reg.write(0);
    else              count_reg.write(count_reg.read() + 1);
  }
  void out_proc() { count.write(count_reg.read()); }

  SC_CTOR(Counter4) {
    SC_METHOD(reg_proc);
    sensitive << clk.pos();
    dont_initialize();
    SC_METHOD(out_proc);
    sensitive << count_reg;
  }
};

int sc_main(int, char*[]) {
  sc_clock              clk("clk", 10, SC_NS, 0.5, 5, SC_NS, false);  // first posedge at 10 ns
  sc_signal<bool>       reset;
  sc_signal<sc_uint<4>> count;

  Counter4 dut("dut");
  dut.clk(clk); dut.reset(reset); dut.count(count);

  // Open a VCD file and register the signals we want to see.
  sc_trace_file* vcd = sc_create_vcd_trace_file("counter4");
  vcd->set_time_unit(1, SC_NS);
  sc_trace(vcd, clk,   "clk");
  sc_trace(vcd, reset, "reset");
  sc_trace(vcd, count, "count");

  reset.write(true);
  sc_start(10, SC_NS);
  reset.write(false);
  sc_start(180, SC_NS);

  sc_close_vcd_trace_file(vcd);
  sc_stop();
  return 0;
}

Running this produces counter4.vcd. Open it in GTKWave and you will see three traces: clk toggling every 5 ns, reset high for the first 10 ns, and count stepping 0 → 1 → 2 → … → 15 → 0, each step landing one delta after a rising clock edge. The waveform makes the clock-to-Q delay visible: count changes just after each rising clk edge, never exactly on it. That tiny gap is the update-phase deferral — the same mechanism we have been leaning on all post, now drawn on a timeline.

The key sc_trace facts: you create the file once with sc_create_vcd_trace_file (it appends .vcd automatically), register each signal with sc_trace(file, signal, "name") before sc_start, and close it with sc_close_vcd_trace_file at the end. Tracing is the primary debugging tool for sequential logic, because a register bug almost always shows up as "the value changed on the wrong edge" or "the value did not change when it should have" — both instantly visible in a waveform and nearly invisible in printed text.

Advanced: Edge Cases & LRM Corners

You now have the register pattern from first principles, the register-plus-combinational split, and both reset flavors. This section does two things. First it applies all of it to a real, CPU-scale register — the RISC-V program counter with instruction fetch, the worked example for this part of the section. Then it drills into the LRM corners that bite people: the deferred-read trap, multiple writers to one signal, and the dont_initialize() placement bug.

Worked example: the RV32I program counter + instruction fetch

The program counter is the simplest register in a CPU and the most important. One word wide, it holds the address of the instruction to execute. On most cycles it just advances by four (the next sequential instruction, since RV32I instructions are four bytes). But on a taken branch or a jump, the next PC is not PC+4 — it is a computed target. That makes the PC a textbook register-plus-combinational block: a clocked register holding the current PC, and a combinational next-PC mux choosing among PC+4, a branch target, a JAL target, and a JALR target.

This is exactly the structure from the Intermediate section, scaled up. The register is pc_reg (clocked, synchronous reset to the reset vector). The combinational next_pc_logic() is a four-way mux. The fetch side is a combinational ROM (imem) that converts the byte-addressed PC into a word index and reads the instruction. The four next-PC sources:

Condition Next PC RISC-V instructions
Sequential (default) pc + 4 all non-branch, non-jump
Branch taken pc + sign_extend(offset) BEQ, BNE, BLT, BGE, BLTU, BGEU
JAL pc + sign_extend(offset) JAL
JALR (rs1 + sign_extend(imm)) & ~1 JALR

The & ~1 on JALR is mandatory per the RISC-V ISA manual (§2.5): the target is formed by adding the immediate to rs1 and then forcing the least-significant bit to zero, so the jump always lands on a two-byte boundary. Omit it and an odd JALR target fetches from a misaligned address — a trap on real hardware, silently wrong behavior in a naive model.

Here is the complete, self-contained program: pc register + next-PC mux, imem combinational ROM, a small driver that walks through sequential fetch and a taken branch, and a monitor.

// file: pc_fetch.cpp
// RV32I program counter + instruction fetch, self-contained.
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            pc_fetch.cpp -o pc_fetch -lsystemc

#include <systemc.h>
#include <cstdint>
#include <iostream>
#include <iomanip>

// ---- Program counter: clocked register + combinational next-PC mux ----
SC_MODULE(Pc) {
  sc_in<bool>         clk;
  sc_in<bool>         reset;          // active-high synchronous reset
  sc_in<bool>         branch_taken;
  sc_in<sc_uint<32>>  branch_offset;  // sign-extended B-type immediate
  sc_in<bool>         jump;           // JAL
  sc_in<sc_uint<32>>  jump_offset;    // sign-extended J-type immediate
  sc_in<bool>         jalr;           // JALR
  sc_in<sc_uint<32>>  jalr_target;    // rs1 + imm, computed elsewhere
  sc_out<sc_uint<32>> pc_out;

  sc_signal<sc_uint<32>> pc_reg;      // the PC register (state)
  sc_signal<sc_uint<32>> next_pc;     // combinational next value

  static const uint32_t RESET_VECTOR = 0x00000000u;

  // --- Combinational: next-PC mux ---
  void next_pc_logic() {
    sc_uint<32> cur = pc_reg.read();
    sc_uint<32> nxt;
    if (jalr.read()) {
      sc_uint<32> t = jalr_target.read();
      t[0] = 0;                                   // force LSB to 0 (ISA §2.5)
      nxt = t;
    } else if (branch_taken.read()) {
      sc_int<32> s = (sc_int<32>)cur + (sc_int<32>)branch_offset.read();
      nxt = (sc_uint<32>)s;
    } else if (jump.read()) {
      sc_int<32> s = (sc_int<32>)cur + (sc_int<32>)jump_offset.read();
      nxt = (sc_uint<32>)s;
    } else {
      nxt = cur + 4;                              // sequential
    }
    next_pc.write(nxt);
  }

  // --- Clocked register: store next_pc (or reset vector) on the edge ---
  void pc_reg_proc() {
    if (reset.read()) pc_reg.write(RESET_VECTOR);
    else              pc_reg.write(next_pc.read());
  }

  void out_proc() { pc_out.write(pc_reg.read()); }

  SC_CTOR(Pc) {
    SC_METHOD(next_pc_logic);
    sensitive << pc_reg << branch_taken << branch_offset
              << jump << jump_offset << jalr << jalr_target;

    SC_METHOD(pc_reg_proc);
    sensitive << clk.pos();
    dont_initialize();

    SC_METHOD(out_proc);
    sensitive << pc_reg;
  }
};

// ---- Instruction memory: combinational ROM (no clock) ----
SC_MODULE(Imem) {
  sc_in<sc_uint<32>>  addr;    // byte address (from pc_out)
  sc_out<sc_uint<32>> instr;

  static const int MEM_WORDS = 256;
  uint32_t mem[MEM_WORDS];

  void read_proc() {
    sc_uint<32> a = addr.read();
    sc_uint<30> word_idx = a.range(31, 2);    // a >> 2
    if (word_idx >= (sc_uint<30>)MEM_WORDS) {
      instr.write(0x00000013u);               // NOP for out-of-range
      return;
    }
    instr.write(mem[(uint32_t)word_idx]);
  }

  void load(const uint32_t* prog, int n) {
    for (int i = 0; i < n && i < MEM_WORDS; i++) mem[i] = prog[i];
  }

  SC_CTOR(Imem) {
    for (int i = 0; i < MEM_WORDS; i++) mem[i] = 0x00000013u;  // NOP fill
    SC_METHOD(read_proc);
    sensitive << addr;
  }
};

// ---- Driver: sequential fetch, then a taken branch at PC=0x10 ----
SC_MODULE(Driver) {
  sc_in<bool>         clk;
  sc_in<sc_uint<32>>  pc_out;
  sc_out<bool>        branch_taken;
  sc_out<sc_uint<32>> branch_offset;

  void step() {
    // Take a backward branch (offset -16) once we reach PC = 0x10.
    if (pc_out.read() == 0x10) {
      branch_taken.write(true);
      branch_offset.write((sc_uint<32>)(sc_int<32>)(-16));   // 0x10 - 16 = 0x00
    } else {
      branch_taken.write(false);
      branch_offset.write(0);
    }
  }

  SC_CTOR(Driver) {
    SC_METHOD(step);
    sensitive << pc_out;     // react combinationally to the current PC
  }
};

SC_MODULE(Monitor) {
  sc_in<bool>        clk;
  sc_in<sc_uint<32>> pc_out;
  sc_in<sc_uint<32>> instr;

  void watch() {
    std::cout << "[" << sc_time_stamp() << "] pc=0x"
              << std::hex << std::setw(8) << std::setfill('0')
              << (uint32_t)pc_out.read()
              << " instr=0x" << std::setw(8) << std::setfill('0')
              << (uint32_t)instr.read() << std::dec << "\n";
  }

  SC_CTOR(Monitor) {
    SC_METHOD(watch);
    sensitive << clk.pos();
    dont_initialize();
  }
};

int sc_main(int, char*[]) {
  sc_clock              clk("clk", 10, SC_NS, 0.5, 5, SC_NS, false);  // first posedge at 10 ns
  sc_signal<bool>       reset;
  sc_signal<bool>       branch_taken, jump, jalr;
  sc_signal<sc_uint<32>> branch_offset, jump_offset, jalr_target;
  sc_signal<sc_uint<32>> pc_out, instr;

  Pc pc("pc");
  pc.clk(clk); pc.reset(reset);
  pc.branch_taken(branch_taken); pc.branch_offset(branch_offset);
  pc.jump(jump); pc.jump_offset(jump_offset);
  pc.jalr(jalr); pc.jalr_target(jalr_target);
  pc.pc_out(pc_out);

  Imem imem("imem");
  imem.addr(pc_out); imem.instr(instr);

  Driver drv("drv");
  drv.clk(clk); drv.pc_out(pc_out);
  drv.branch_taken(branch_taken); drv.branch_offset(branch_offset);

  Monitor mon("mon");
  mon.clk(clk); mon.pc_out(pc_out); mon.instr(instr);

  // Unused control inputs held inactive.
  jump.write(false); jump_offset.write(0);
  jalr.write(false); jalr_target.write(0);

  // A tiny program at word indices 0..5 (byte addresses 0x00..0x14).
  static const uint32_t prog[] = {
    0x00500093u,  // 0x00: addi x1, x0, 5
    0x00300113u,  // 0x04: addi x2, x0, 3
    0x002080b3u,  // 0x08: add  x1, x1, x2
    0x40208133u,  // 0x0C: sub  x2, x1, x2
    0x00000013u,  // 0x10: nop  (branch is taken here, back to 0x00)
    0x00000013u   // 0x14: nop  (not reached)
  };
  imem.load(prog, 6);

  reset.write(true);
  sc_start(10, SC_NS);     // one cycle of reset -> PC = 0x00
  reset.write(false);
  sc_start(80, SC_NS);     // fetch 0x00,04,08,0C,10 then branch back to 0x00

  sc_stop();
  return 0;
}

Expected output:

[10 ns] pc=0x00000000 instr=0x00500093
[20 ns] pc=0x00000004 instr=0x00300113
[30 ns] pc=0x00000008 instr=0x002080b3
[40 ns] pc=0x0000000c instr=0x40208133
[50 ns] pc=0x00000010 instr=0x00000013
[60 ns] pc=0x00000000 instr=0x00500093
[70 ns] pc=0x00000004 instr=0x00300113
[80 ns] pc=0x00000008 instr=0x002080b3

Walk the trace and notice it is the counter pattern in disguise. Reset is held across t=10 ns, so the PC register samples reset high and writes the reset vector 0x00000000; the combinational imem reads word index 0 and produces 0x00500093. The monitor prints PC 0x00, instruction 0x00500093. On each subsequent edge with no branch, next_pc_logic() computes pc_reg + 4, the register captures it, and the PC marches 0x04, 0x08, 0x0C, 0x10 — incrementing by four, exactly like the four-bit counter incremented by one. The imem fetch is purely combinational: whenever pc_out changes, read_proc() re-runs and the instruction follows one delta later, with no clock of its own.

The branch is where the next-PC mux earns its keep. At PC 0x10 the Driver asserts branch_taken with offset -16. The combinational next_pc_logic() (sensitive to pc_reg, branch_taken, and branch_offset) computes 0x10 + (-16) = 0x00 and writes it to next_pc. At the next clock edge (t=60 ns) the PC register captures 0x00 instead of 0x14, and the program loops back to the top. The monitor at t=60 ns prints PC 0x00 again. The register did not change behavior — it still just stores next_pc on each edge. All that changed is what the combinational mux fed it. That separation — a dumb register that stores whatever the smart combinational logic computes — is the entire point of the register-plus-combinational pattern, and it is exactly what the FSM post formalizes.

Note The imem here is combinational (no clock), modeling an asynchronous-read ROM whose output is ready the same cycle the address is presented. Real instruction memories are usually synchronous-read SRAMs whose output appears one cycle after the address is registered. The combinational model is a fine pedagogical simplification: it keeps fetch to one stage so the trace stays readable. When you later build a pipeline (Section 5), the fetch register is exactly the kind of register this post teaches, inserted between the PC and the decoder.

LRM corner 1: the deferred-read trap

The single most common sequential-logic bug in SystemC comes from forgetting that sc_signal::read() returns the committed value, not a value you just wrote in the same evaluate phase. Consider this broken attempt to "increment twice per cycle":

// BUGGED: expects count to advance by 2 per edge. It advances by 1.
void reg_proc_buggy() {
  count_reg.write(count_reg.read() + 1);   // schedules count+1
  count_reg.write(count_reg.read() + 1);   // reads OLD count again -> count+1
}

The author imagines the first write updates count_reg, so the second read sees the incremented value and the net effect is +2. It is not. Per §6.4, both read() calls return the same committed value (the value from before this edge), because neither write has been committed yet — commits happen in the update phase, after the method returns. The two writes both schedule a request-update; the last write wins; the net effect is count + 1. The fix, if you genuinely want +2, is to use a local C++ variable for within-cycle sequencing and write the signal once:

void reg_proc_fixed() {
  sc_uint<4> v = count_reg.read();
  v = v + 1;
  v = v + 1;                 // plain C++ variable: ordinary sequencing
  count_reg.write(v);        // single write of the final value
}

Local variables follow normal C++ semantics — read-after-write works as you expect. sc_signal values do not, within an evaluate phase. The rule: use locals for intermediate computation, write the signal once with the final value. This is the same discipline as SystemVerilog, where you use blocking = on combinational temporaries and a single non-blocking <= to the register.

LRM corner 2: multiple writers to one signal

Suppose two processes both write next_pc in the same evaluate phase — say one handles branches and another handles jumps, and you forgot to merge them:

// BUGGED: two processes drive next_pc in the same delta.
void branch_writer() { if (branch_taken.read()) next_pc.write(branch_tgt); }
void jump_writer()   { if (jump.read())         next_pc.write(jump_tgt);   }

If both conditions are true in the same evaluate phase, both processes write next_pc, and which value the update phase commits is implementation-defined (§6.4 — the standard does not guarantee an order across independent processes driving the same signal). The Accellera reference kernel happens to let the last-registered writer win, but you must not rely on that; a different or randomized kernel makes it nondeterministic. The fix is the one the worked example already uses: compute the next value in a single combinational process with one prioritized if/else if chain (or one switch), so there is exactly one writer of next_pc. One signal, one driver, deterministic on every kernel. The next-PC mux above is precisely this single-writer pattern, and it is why JALR, branch, jump, and sequential are arms of one if chain rather than four independent processes.

LRM corner 3: dont_initialize() placement

dont_initialize() belongs on the clocked register process, never on the combinational one. Put it on the wrong process and you get one of two failures.

If you forget it on the clocked register, the register runs once during the initialization phase — before any clock edge — reading whatever the default-constructed next signal holds (0 for arithmetic types) and writing it into the state. For a counter or PC that resets to 0 this is invisibly harmless. For a register whose reset value is non-zero, the spurious initialization write corrupts the starting state before the first edge.

If you add it to the combinational process by mistake, the combinational logic does not run at t=0, so the design's outputs sit at their default-constructed values until the first input change wakes the process. If the state register also happens to start at its default value (so it produces no value-changed event on the first reset-released edge), the combinational process never wakes and the outputs are stuck at default forever. This is a genuinely nasty bug because it only manifests when the default state value coincides with the reset value — exactly the common case — making it intermittent across designs.

The rule has no exceptions: dont_initialize() on every clocked (clk.pos()-sensitive) process; never on a combinational (sensitive << inputs) process. The worked example follows it — dont_initialize() is on pc_reg_proc and reg_proc, and absent from every combinational method.

Version differences

Everything in this post — SC_METHOD semantics, static sensitivity, dont_initialize(), sc_signal's update-phase deferral, sc_clock mechanics, and sc_trace/VCD output — behaves identically across SystemC 2.3.1, 2.3.3, and 2.3.4. No changes to these mechanisms exist across those releases; differences are confined to build tooling, C++17 conformance, and TLM-2.0/sc_vector helpers that none of these examples touch. Anything here compiles and behaves the same on any 2.3.x build.

Hands-on exercise

Build a register with more behavior than a plain counter, to cement the register-plus-combinational pattern before the FSM post leans on it.

Part A — a loadable up/down counter. Build an eight-bit counter (sc_uint<8>) with four inputs: clk, reset (active-high synchronous), load, up, plus a load_value data input. Behavior, in priority order, sampled on each rising edge: if reset, go to 0; else if load, capture load_value; else if up, increment; else decrement. Use the two-process split: a combinational next_logic() that lists every input it reads, and a clocked reg_proc() sensitive to clk.pos() with dont_initialize(). Drive it with a clocked stimulus that loads 100, counts up three, switches to down for five, then resets. Predict the printed value on each cycle before you run, then verify. The trap to watch for: if you omit any input from next_logic's sensitivity list, that input's changes will be ignored until the next time a listed signal changes — find out which one breaks first.

Part B — add VCD tracing. Instrument Part A with sc_trace: dump clk, reset, load, up, the load_value, and the count to a VCD file. Open it in GTKWave (or any VCD viewer). Confirm visually that the count changes one delta after each rising clock edge, never on it — that gap is the clock-to-Q delay you have been reading about. Confirm that asserting load mid-cycle does nothing until the next edge (synchronous behavior).

Part C — make the reset asynchronous, then compare. Take your Part A counter and convert the reset to asynchronous: add reset to the clocked process's sensitivity list and check it first. Assert reset between two clock edges (use a stimulus thread that writes reset high at a non-edge time, then low a couple of nanoseconds later). Trace both the synchronous and asynchronous versions and put the two VCDs side by side. The synchronous reset pulse that falls entirely between edges should be invisible in the count; the asynchronous one should zero the count immediately. Seeing the two waveforms differ on exactly that pulse is the payoff — it is the whole synchronous-vs-asynchronous distinction in one picture.

Part D — stretch. Extend Part A so the counter saturates instead of wrapping: at 255 with up asserted, hold at 255; at 0 with down asserted, hold at 0. This is pure combinational change in next_logic — the register does not change at all. That is the lesson: behavior lives in the combinational next-value logic; the register is a dumb storage element. Internalizing that is exactly the mindset the FSM post requires.

No solutions are provided. The understanding is in the building and in reading your own waveforms.

Common mistakes

  • Forgetting dont_initialize() on the clocked register. Without it the register fires once during initialization, before any clock edge, reading the default-constructed next value (0) and writing it into the state. For a register whose reset value is also 0 this is invisible; for any other reset value it corrupts the starting state silently. The bug never shows in waveforms because the register still runs — it just begins in the wrong state. Fix: dont_initialize() on every clk.pos()-sensitive process.
  • Putting the clock in a combinational process's sensitivity list, or vice versa. A register is sensitive to clk.pos() only. A combinational next-value process is sensitive to its data inputs and the state signal, never to the clock. Mix these up and you either get a "register" that reacts to data mid-cycle (a latch, not a flip-flop) or "combinational" logic that only updates on clock edges (a pipeline you did not intend). Fix: clocked process gets clk.pos(); combinational process gets every signal it reads.
  • Expecting read() to see a same-cycle write(). count_reg.write(x) followed by count_reg.read() in the same process returns the old value, not x, because the write commits in the update phase after the method returns. Code that increments a signal "twice per cycle" with two writes advances by one, not two (last writer wins). Fix: use a local C++ variable for within-cycle sequencing and write the signal once with the final value.
  • Two processes writing the same signal in one delta. When two independent processes drive the same sc_signal in the same evaluate phase, which value commits is implementation-defined — your design becomes kernel-dependent. Fix: every signal must have exactly one driver. Merge all next-value logic for a given register into a single combinational process with one prioritized if/else chain or switch.
  • Writing the state register from the combinational process. The state signal must be written only inside the clocked process. If the combinational process writes state (instead of next_state), you bypass the clock, the register loses its storage, and the design degenerates into a transparent latch network whose value changes whenever any input wiggles. Fix: combinational logic writes only next_state; the single state.write(...) lives in the clocked register process.
  • Forgetting a signal in the combinational sensitivity list. SystemC has no always @(*). If next_logic reads enable but the list says only sensitive << count_reg, the process never re-evaluates when enable changes, and the behavior silently lags or freezes. Fix: every read inside a combinational method corresponds to exactly one entry in its sensitivity list — check them against each other every time.
  • Forgetting the JALR LSB clear in the PC. A subtle worked-example bug: JALR must force bit 0 of its target to zero (t[0] = 0). Omit it and an odd target fetches from a misaligned address — a trap on real hardware, silent corruption in a model that does not bounds-check alignment. Fix: clear the LSB on the JALR path; it is an ISA requirement (§2.5), not an option.

Recap

After working through this post you can now:

  • State what a register is from first principles — a storage element that samples its input on the clock edge, holds the value between edges, and exposes it after a clock-to-Q delay — and explain why combinational logic computes while registers remember.
  • Build a clocked register in SystemC as an SC_METHOD sensitive to clk.pos() with dont_initialize(), and explain why sc_signal's deferred update gives you flip-flop (non-blocking) semantics for free, with no flip-flop primitive in the language.
  • Predict the exact cycle on which a register captures a value and the exact delta on which that value becomes visible downstream, by reasoning about the evaluate/update phases.
  • Split a registered block into a combinational next-value process and a clocked register process, list the sensitivity for each correctly, and recognize this as the same skeleton an FSM uses.
  • Choose between synchronous reset (sampled on the edge, the section default) and asynchronous reset (takes effect immediately, sensitivity-list change, the labeled exception), and justify the choice on synthesis and metastability grounds.
  • Use sc_trace to dump a VCD and read the clock-to-Q delay directly off the waveform.
  • Apply the pattern at CPU scale: a program counter as a clocked register fed by a combinational next-PC mux (PC+4 vs branch target vs JAL vs JALR), with a combinational instruction-memory fetch.
  • Diagnose and fix the recurring sequential bugs: missing dont_initialize(), clock-in-combinational or data-in-clocked sensitivity errors, the deferred-read trap, multiple writers to one signal, and a combinational write to the state register.

The one sentence to carry into the next post: a register is a clocked SC_METHOD that copies next_state into state once per edge, and the value becomes visible one delta later — which is exactly the storage element a finite state machine wraps in combinational next-state logic.

Further reading

Standards

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC Language Reference Manual, §5.2.16 (SC_METHOD semantics), §5.2.18 (static sensitivity and dont_initialize()), §6.4 (sc_signal write/update semantics), §4.2.1.3 (the update phase and delta cycles), §6.7 (sc_clock).
  • RISC-V ISA Manual, Volume I: Unprivileged ISA, §2.5 (unconditional jumps; the JALR target LSB-clear requirement).

Vendor and consortium documents

  • Accellera Systems Initiative, SystemC 2.3.x User Guide — process kinds, channel mechanics, and tracing.

Textbooks

  • Bhasker, A SystemC Primer (2nd ed.), ch. 4–5 — register modeling and the clocked-method idiom.
  • Grötker, Liao, Martin, and Swan, System Design with SystemC, ch. 3–4 — signal update semantics and synchronous design.

Training notes

  • MIT 6.004 lecture notes — registered vs combinational logic, the register-plus-logic split.
  • Berkeley CS152 lecture notes — RV32I single-cycle PC and the next-PC mux derivation.

Real-world references

  • CV32E40P and Ibex PC/prefetch logic (public sources, OpenHW Group and lowRISC) — production next-PC mux and reset-vector handling that mirror the worked example here.

Next in this section

→ Part 4: Finite State Machines (Moore vs Mealy) — the register you built in this post becomes the state register of an FSM, wrapped in combinational next-state and output logic. You will learn the two-process FSM pattern, the Moore-versus-Mealy output-timing distinction cycle-by-cycle, and how to predict and debug FSM behavior without running the simulator. Read it here: 11. SystemC Tutorial — Finite State Machines.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment