11. SystemC Tutorial - Finite State Machines (Moore vs Mealy)
Why this matters
Rewritten 2026-05-25 with deeper first-principles material.
Every non-trivial RTL design you will ever model in SystemC has at least one finite state machine inside it, and most have dozens. The controller in CV32E40P that sequences fetch, decode, execute, and writeback is an FSM. The PCIe LTSSM that walks a link through Detect, Polling, Configuration, L0, Recovery, and the low-power states is an FSM. The USB 2.0 device-side enumeration that the bus controller in your laptop's southbridge ran when you plugged this article's USB-C dock in this morning is an FSM. The arbitration logic at the heart of every AXI interconnect, the cache controller that decides whether a miss escalates to a fill or a forward, the DRAM scheduler that picks the next bank to precharge — all FSMs. When the chip you are validating hangs in post-silicon bring-up, the first place verification engineers look is at the FSM that should have advanced and did not. When a pre-silicon model boots Linux but the kernel reports a spurious interrupt at second 47, the cause is almost always an FSM whose output was sampled in the wrong cycle.
Modeling FSMs correctly in SystemC is therefore not a "nice to have" — it is the single most load-bearing skill in the RTL-patterns section of this series. Get the Moore-vs-Mealy distinction wrong by one cycle and your virtual platform mis-models the latency of every protocol your CPU touches. Get the sensitivity list of the combinational next-state process wrong and your kernel produces correct waveforms while running 50× slower than necessary. Write the state register from a combinational process by accident and your FSM has no memory, your simulation hangs on the first input pattern that requires the state to remember anything, and you spend two days bisecting your own code looking for the bug. This post fixes all of that. By the end you will be able to write a two-process FSM in SystemC from scratch, predict its output cycle-by-cycle without running the simulator, recognize whether a given design wants Moore or Mealy or registered-Mealy semantics, and debug the three most common FSM bugs in production code on sight. That is the bar.
Prerequisites
- Part 1 — Modules, Ports & Signals. You need to be comfortable declaring
SC_MODULE, bindingsc_signalchannels tosc_in/sc_outports, and registering a process withSC_METHODinsideSC_CTOR. - Part 3 — Delta Cycles & Event-Driven Semantics. You need to know that an
sc_signal::writeis buffered, that the new value commits in the next update phase, and that a value-changed event wakes any process statically sensitive to the signal. The whole register-as-update-delay idea of FSMs in SystemC builds on this. - Part 4 — Processes & Sensitivity (SC_METHOD, SC_THREAD, SC_CTHREAD). You need to understand the difference between an
SC_METHODand anSC_THREAD, why FSMs use methods (not threads), and how a sensitivity list is built usingsensitive << clk.pos()versussensitive << state << input_a << input_b. - SystemC 2.3.x installed. The installer posts (P1 and P2 in this series, the original installation tutorials for Linux and macOS) cover environment setup. All examples in this post compile with a C++17 compiler (
g++9 or newer, orclang++10 or newer) against any 2.3.x SystemC build.
Three concepts from earlier parts will get heavy use here: sc_signal<T> (the buffered channel that becomes your state register), SC_METHOD (the non-blocking process kind both halves of the FSM use), and static sensitivity (the sensitive << ... invocation that wires a process to its inputs). If any of those is hazy, re-read the relevant section before continuing — the rest of this post assumes them.
Mental model (first principles)
A finite state machine in hardware is two pieces of logic separated by one storage element. The storage element holds the current state. The logic on one side of it — the next-state logic — computes what state to enter on the next clock edge as a function of the current state and the inputs. The logic on the other side — the output logic — computes the FSM's outputs. In a Moore machine, the output logic is a function of the current state only. In a Mealy machine, the output logic is a function of the current state and the inputs.
Hardware engineers draw this in three boxes:
┌────────────────────────┐
inputs ────► │ next-state logic ├──► next_state
┌─► │ (combinational) │ │
│ └────────────────────────┘ │
│ ▼
│ ┌─────────────────┐
│ clk │ state register │
│ ───►│ (storage) │
│ └────────┬────────┘
│ │ state
│ │
├───────────────────────────────────────┤
│ │
│ ┌────────────────────────┐ │
inputs ──┴──►│ output logic │◄─────────┤ (Mealy: inputs feed
│ (combinational; Moore │ the output logic;
│ ignores inputs, Mealy │ Moore: only state)
│ reads them) │
└────────────┬───────────┘
▼
outputs
In SystemC there is no separate flip-flop primitive. The sc_signal<state_t> channel is the flip-flop, and the kernel's update-phase delay (covered in Part 3) is the clock-to-Q delay. This is the load-bearing insight of the whole post. A SystemC FSM is two SC_METHOD processes that share one sc_signal<state_t>:
- The state-register process is sensitive only to
clk.pos()(and torstif you want asynchronous reset). On every rising clock edge it executes one line:state.write(next_state.read()). That is the entire body. The buffered write commits in the next update phase, the value-changed event fires, and every process sensitive tostatewakes up in the following delta. The kernel's update-phase delay is exactly what makes this look like a registered transfer.
- The combinational process is sensitive to
stateand to every input the FSM reads. Its job is to compute two things:next_state(always, for every transition the FSM might make) and the FSM's outputs. For a Moore FSM, the outputs depend only onstate; for a Mealy FSM, the outputs depend onstateand the inputs.
The kernel's behavior per clock edge is then completely deterministic. Suppose the clock just had a rising edge at time t and the FSM is currently in state S_A. The clock edge fires the state-register process; it reads next_state (whose value was computed by the combinational process during the previous delta, after the last input change), writes that value into state, and returns. The kernel's update phase commits state's new value and fires its value-changed event. In the next delta, still at time t, the combinational process wakes, reads the new state, computes a fresh next_state and fresh outputs, writes both, and returns. Update phase commits those writes. The kernel finds the runnable set empty, advances time to the next clock edge, and the loop repeats. Two delta cycles per clock edge: the state-register hop, then the combinational hop.
Now the Moore-versus-Mealy distinction snaps into focus. In a Moore design, the FSM's outputs change only when state changes — that is, only in the delta after a clock edge. The output is registered. There is exactly one cycle of latency between an input change that causes a transition and the output that reflects that transition. In a Mealy design, the FSM's outputs change whenever state or an input changes. The output reacts combinationally to inputs. If an input arrives mid-cycle (between clock edges), the Mealy output changes immediately — no clock-edge wait, no one-cycle delay — but as a side effect, any glitch on the input produces a glitch on the output, and that glitch is a real value_changed_event in the SystemC kernel. Downstream processes sensitive to the output will see the glitch and react.
Per the standard's update-phase semantics, all of this is deterministic given a fixed order of clock edges and input arrivals. The implementation-defined runnable-set ordering inside a delta (Part 3, Advanced section) does not affect FSM correctness here because the FSM's two processes have non-overlapping sensitivity lists: the state-register process fires on clock; the combinational process fires on state plus inputs. They cannot both be runnable in the same delta. The state-register process runs in delta N (clock edge); the combinational process runs in delta N+1 (post state update). The reverse order is impossible because the combinational process's wake-up depends on the state-register process completing its update.
That two-process, one-shared-state-register pattern is the entire substance of FSM modeling in SystemC. Everything else in this post is variations on it: Moore vs Mealy is a choice about what the combinational process reads to compute its outputs; registered Mealy is the same pattern with one extra output register; async reset is a sensitivity-list tweak on the state-register process. If you internalize the two-process pattern, every FSM you ever write is just filling in the case statements.
stateDiagram-v2
direction LR
[*] --> RED
RED --> RED_YELLOW : tick == TR
RED_YELLOW --> GREEN : tick == TRY
GREEN --> YELLOW : tick == TG
YELLOW --> RED : tick == TY
note right of RED : output r=1, y=0, g=0 (Moore — registered)
note right of GREEN : output r=0, y=0, g=1
The state diagram above is the traffic-light controller we will build in the Beginner section. Read it as: four states, transitions on tick-counter thresholds, outputs purely a function of state (the classic Moore pattern). Notice that no transition is labeled with an output — that is the visual signature of a Moore machine. A Mealy diagram would have output annotations on the transition arrows, not the nodes.
Beginner: First Principles
The simplest interesting FSM is a traffic-light controller. Four states, four registered outputs, one tick counter to time the transitions. We will build it now, run it in our heads, predict the output line by line, then verify by running it on a real kernel.
// file: traffic_light.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// traffic_light.cpp -o traffic_light -lsystemc
// Run: LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./traffic_light
#include <systemc.h>
#include <iostream>
enum light_state_t {
ST_RED = 0,
ST_RED_YELLOW = 1,
ST_GREEN = 2,
ST_YELLOW = 3
};
SC_MODULE(TrafficLight) {
sc_in<bool> clk;
sc_in<bool> rst;
sc_out<bool> red;
sc_out<bool> yellow;
sc_out<bool> green;
// State register and next-state wire (both sc_signal — the kernel
// gives us the flip-flop semantics via the update-phase delay).
sc_signal<int> state;
sc_signal<int> next_state;
// Tick counter — counts clock cycles spent in the current state.
sc_signal<int> tick;
sc_signal<int> next_tick;
// How long each state lasts, in clock cycles.
static constexpr int T_RED = 3;
static constexpr int T_RED_YELLOW = 1;
static constexpr int T_GREEN = 3;
static constexpr int T_YELLOW = 1;
// --- State register process: clocked, sensitive only to clk.pos() ---
void state_register() {
if (rst.read()) {
state.write(ST_RED);
tick.write(0);
} else {
state.write(next_state.read());
tick.write(next_tick.read());
}
}
// --- Combinational process: next-state and outputs, both functions
// of the current state (Moore: outputs do NOT read inputs) ---
void next_state_and_outputs() {
int s = state.read();
int t = tick.read();
int ns = s;
int nt = t + 1;
switch (s) {
case ST_RED:
if (t + 1 >= T_RED) { ns = ST_RED_YELLOW; nt = 0; }
break;
case ST_RED_YELLOW:
if (t + 1 >= T_RED_YELLOW) { ns = ST_GREEN; nt = 0; }
break;
case ST_GREEN:
if (t + 1 >= T_GREEN) { ns = ST_YELLOW; nt = 0; }
break;
case ST_YELLOW:
if (t + 1 >= T_YELLOW) { ns = ST_RED; nt = 0; }
break;
}
next_state.write(ns);
next_tick.write(nt);
// Outputs — pure function of CURRENT state. This is Moore.
red.write (s == ST_RED || s == ST_RED_YELLOW);
yellow.write(s == ST_RED_YELLOW || s == ST_YELLOW);
green.write (s == ST_GREEN);
}
SC_CTOR(TrafficLight) {
SC_METHOD(state_register);
sensitive << clk.pos();
dont_initialize();
SC_METHOD(next_state_and_outputs);
sensitive << state << tick;
// No dont_initialize() — let it run at t=0 to establish outputs.
}
};
SC_MODULE(Monitor) {
sc_in<bool> clk;
sc_in<bool> red;
sc_in<bool> yellow;
sc_in<bool> green;
void watch() {
std::cout << "[" << sc_time_stamp() << "] "
<< "R=" << red.read()
<< " Y=" << yellow.read()
<< " G=" << green.read() << "\n";
}
SC_CTOR(Monitor) {
SC_METHOD(watch);
sensitive << clk.pos();
dont_initialize();
}
};
int sc_main(int /*argc*/, char* /*argv*/[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst;
sc_signal<bool> red, yellow, green;
TrafficLight tl("tl");
tl.clk(clk); tl.rst(rst);
tl.red(red); tl.yellow(yellow); tl.green(green);
Monitor mon("mon");
mon.clk(clk); mon.red(red); mon.yellow(yellow); mon.green(green);
// Hold reset high for one cycle, then release.
rst.write(true);
sc_start(10, SC_NS);
rst.write(false);
sc_start(90, SC_NS);
sc_stop();
return 0;
}
Compile and run, and you will see this output. Read it before the walkthrough — try to explain each line to yourself first.
Expected output:
[10 ns] R=1 Y=0 G=0
[20 ns] R=1 Y=0 G=0
[30 ns] R=1 Y=0 G=0
[40 ns] R=1 Y=1 G=0
[50 ns] R=0 Y=0 G=1
[60 ns] R=0 Y=0 G=1
[70 ns] R=0 Y=0 G=1
[80 ns] R=0 Y=1 G=0
[90 ns] R=1 Y=0 G=0
[100 ns] R=1 Y=0 G=0
Walk through it. The clock period is 10 ns, so positive edges happen at 10, 20, 30, … ns. Reset is held high from t=0 to t=10 ns. At the first positive edge (t=10 ns) the state-register process fires while rst is still high; it writes ST_RED into state and 0 into tick. The kernel's update phase commits both. The combinational process fires in the next delta — still at t=10 ns — sees state == ST_RED, drives red.write(true). The update phase commits the red output. The monitor, also clocked, fires at the same edge and prints R=1 Y=0 G=0. (Because of the implementation-defined runnable-set ordering, the monitor may print before or after the state-register process's effects propagate — but dont_initialize() on the monitor ensures it does not print at t=0, and by t=10 ns the registered outputs are stable.)
At t=20 ns and t=30 ns, rst is low. The state-register process reads next_state (which the combinational process computed in the previous cycle's delta) and writes it into state. tick increments. The FSM stays in ST_RED because tick has not yet reached T_RED - 1. The monitor prints R=1.
At t=40 ns, the combinational process computed next_state = ST_RED_YELLOW at the end of the t=30 ns cycle (when tick reached 2 and t + 1 >= T_RED evaluated true). The state-register process at t=40 ns reads that next_state and writes ST_RED_YELLOW into state. The combinational process then computes red=1, yellow=1, green=0. The monitor prints R=1 Y=1 G=0. The FSM is now in the brief transitional RED_YELLOW state.
The pattern continues: GREEN for 3 cycles, YELLOW for 1, back to RED. The full cycle is T_RED + T_RED_YELLOW + T_GREEN + T_YELLOW = 8 clock periods = 80 ns. You can verify by counting from t=10 ns (first RED cycle) to t=90 ns (next RED cycle): exactly 80 ns apart.
Two operational details are worth knowing before you continue.
First, the dont_initialize() call on the state-register process. Without it, the kernel would fire state_register() once during the initialization delta — before the first clock edge — and the body would execute. If rst is high at construction (which it is in our sc_main), the body would write state = ST_RED and tick = 0. That is harmless. But if rst were low at construction, the body would read next_state.read() — which is 0 (the default-constructed value of sc_signal<int>) — and write state = 0. That is ST_RED by coincidence, but in a general FSM with non-zero "first" state encoding, the spurious initialization write would corrupt your FSM before the first clock edge. Always call dont_initialize() on a clocked process. There is no exception.
Second, the absence of dont_initialize() on the combinational process is also deliberate. The combinational process runs once at initialization, reads state (which is 0 = ST_RED by default-construction of sc_signal<int>), and writes the corresponding outputs red=1, yellow=0, green=0. That establishes the FSM's outputs before the first clock edge — so any logic downstream that needs a defined output at t=0 has one. If you mistakenly call dont_initialize() on the combinational process, the FSM's outputs sit at the default-constructed false for bool and 0 for arithmetic types until the first input changes, which can produce "stuck low for one cycle after reset release" bugs that are subtle and frustrating to chase. We will see one in the Advanced section.
A common confusion at this point is to look at the two-process design and ask "why not write the state register and the outputs in one process?" You can — and it works for a simple Moore machine. But the moment the design needs Mealy outputs (output reads inputs), or asynchronous reset (different sensitivity list from clock), or registered Mealy (output passes through one more register), the single-process design fights you. The two-process pattern is the lingua franca of RTL FSM modeling — used in Verilog, SystemVerilog, VHDL, and SystemC alike — precisely because it scales. Start with two processes from day one and you never have to refactor.
Intermediate: How It Really Works
The Beginner example was deliberately tiny so the kernel trace fit in your head. Real FSMs do more interesting things — they react to inputs that arrive between clock edges, they emit outputs that need to be combinational for one design and registered for another, and they live inside larger control hierarchies where one FSM's output is the next FSM's input. The single most important real-world choice is Moore versus Mealy, and the only way to internalize that choice is to build the same FSM both ways and look at the trace.
The canonical pedagogical example is a serial bit detector for the pattern 1011. The detector reads a stream of single-bit inputs, one per clock cycle, and asserts a detected output when the most recent four bits form the sequence 1, 0, 1, 1. We will build it twice: once Moore, once Mealy.
Worked example 1: Moore-variant 1011 detector
stateDiagram-v2
direction LR
[*] --> S0
S0 --> S0 : b=0
S0 --> S1 : b=1
S1 --> S10 : b=0
S1 --> S1 : b=1
S10 --> S0 : b=0
S10 --> S101 : b=1
S101 --> S0_DETECTED : b=1
S101 --> S10 : b=0
S0_DETECTED --> S1 : b=1
S0_DETECTED --> S0 : b=0
note right of S0_DETECTED : output detected=1 (Moore — registered)
Five states. The detected output is asserted in the S0_DETECTED terminal state only. Note that the detector folds back from S0_DETECTED on the same transitions as S0 would — overlapping pattern matches are allowed (the input 10111011 should detect twice).
// file: detector_moore.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// detector_moore.cpp -o detector_moore -lsystemc
#include <systemc.h>
#include <iostream>
enum moore_state_t {
M_S0 = 0,
M_S1 = 1,
M_S10 = 2,
M_S101 = 3,
M_S0_DETECTED = 4
};
SC_MODULE(DetectorMoore) {
sc_in<bool> clk;
sc_in<bool> rst;
sc_in<bool> bit_in;
sc_out<bool> detected;
sc_signal<int> state;
sc_signal<int> next_state;
void state_register() {
if (rst.read()) state.write(M_S0);
else state.write(next_state.read());
}
void next_state_and_outputs() {
int s = state.read();
bool b = bit_in.read();
int ns = s;
switch (s) {
case M_S0: ns = b ? M_S1 : M_S0; break;
case M_S1: ns = b ? M_S1 : M_S10; break;
case M_S10: ns = b ? M_S101 : M_S0; break;
case M_S101: ns = b ? M_S0_DETECTED : M_S10; break;
case M_S0_DETECTED: ns = b ? M_S1 : M_S0; break;
}
next_state.write(ns);
// Moore output: function of CURRENT state only — does NOT read bit_in.
detected.write(s == M_S0_DETECTED);
}
SC_CTOR(DetectorMoore) {
SC_METHOD(state_register);
sensitive << clk.pos();
dont_initialize();
SC_METHOD(next_state_and_outputs);
sensitive << state << bit_in;
}
};
SC_MODULE(BitDriver) {
sc_in<bool> clk;
sc_out<bool> bit_out;
// Test pattern: 1 0 1 1 (detect should fire) 1 0 1 1 (overlapping)
// then 0 to flush.
const char* pattern = "10110110";
int idx = 0;
void drive() {
if (pattern[idx] == '\0') {
bit_out.write(false);
} else {
bit_out.write(pattern[idx] == '1');
idx++;
}
}
SC_CTOR(BitDriver) {
SC_METHOD(drive);
sensitive << clk.pos();
}
};
SC_MODULE(Mon) {
sc_in<bool> clk;
sc_in<bool> bit_in;
sc_in<bool> detected;
void watch() {
std::cout << "[" << sc_time_stamp() << "] bit=" << bit_in.read()
<< " detected=" << detected.read() << "\n";
}
SC_CTOR(Mon) {
SC_METHOD(watch);
sensitive << clk.pos();
dont_initialize();
}
};
int sc_main(int /*argc*/, char* /*argv*/[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst;
sc_signal<bool> bit_wire;
sc_signal<bool> detected_wire;
DetectorMoore det("det");
det.clk(clk); det.rst(rst); det.bit_in(bit_wire); det.detected(detected_wire);
BitDriver drv("drv");
drv.clk(clk); drv.bit_out(bit_wire);
Mon mon("mon");
mon.clk(clk); mon.bit_in(bit_wire); mon.detected(detected_wire);
rst.write(true);
sc_start(10, SC_NS);
rst.write(false);
sc_start(120, SC_NS);
sc_stop();
return 0;
}
Expected output:
[10 ns] bit=0 detected=0
[20 ns] bit=1 detected=0
[30 ns] bit=0 detected=0
[40 ns] bit=1 detected=0
[50 ns] bit=1 detected=0
[60 ns] bit=0 detected=1
[70 ns] bit=1 detected=0
[80 ns] bit=1 detected=0
[90 ns] bit=0 detected=1
[100 ns] bit=0 detected=0
[110 ns] bit=0 detected=0
[120 ns] bit=0 detected=0
[130 ns] bit=0 detected=0
Walk the trace. At t=10 ns reset releases and state = M_S0, bit = 0 (the driver writes 0 as the first character of "10110110" — wait, the first char is '1'. Look again: at t=10 ns the driver's drive() method runs on the clock edge but its index is 0, so it reads pattern[0] == '1' and writes bit_out = true. The monitor reads bit_in at the same clock edge — and because of sc_signal's update-phase delay, the monitor sees the previous value of bit_in, which is false from initialization. So the first monitor line shows bit=0 even though the driver wrote 1 in the same delta.) This is the registered-output behavior cascading through every clocked process in the design: every value is one cycle "behind" the write.
By t=20 ns the driver's bit_out=true from t=10 ns has propagated, and the monitor prints bit=1. The combinational process in the detector saw state=M_S0 and bit_in=1 (after the update phase from t=10 ns committed) and computed next_state=M_S1. At t=20 ns the state register writes M_S1.
Continue: bits arrive as 1, 0, 1, 1, 0, 1, 1, 0, 0, 0, ... (one per cycle, starting at the monitor timestamp). The detector follows the state diagram. At t=60 ns the monitor sees detected=1 because the state register transitioned to M_S0_DETECTED at the t=60 ns clock edge after the combinational logic computed it at the end of the t=50 ns cycle (state = M_S101, bit = 1). The Moore output detected = (state == M_S0_DETECTED) becomes true, the value-changed event fires, and the monitor on the next clock edge prints detected=1.
The crucial point: the input 1 that completed the pattern 1011 arrived at t=50 ns. The detector's detected output went high at t=60 ns. One full cycle of latency. That is the Moore tax. The output is clean (registered, no glitches), but it lags the causing input by one cycle.
Worked example 2: Mealy-variant 1011 detector
Same detection job, different output timing. The Mealy version reads the input bit when computing its output and asserts detected combinationally on the transition that completes the pattern — same delta as that input arrives.
stateDiagram-v2
direction LR
[*] --> S0
S0 --> S0 : b=0 / d=0
S0 --> S1 : b=1 / d=0
S1 --> S10 : b=0 / d=0
S1 --> S1 : b=1 / d=0
S10 --> S0 : b=0 / d=0
S10 --> S101 : b=1 / d=0
S101 --> S1 : b=1 / d=1
S101 --> S10 : b=0 / d=0
note left of S101 : d=1 fires on the 1-transition out of S101 (Mealy — combinational)
Four states (no separate "detected" state — the detection happens on the transition arc). The output detected is 1 on the arc from S101 to S1 taken on b=1, and 0 everywhere else.
// file: detector_mealy.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// detector_mealy.cpp -o detector_mealy -lsystemc
#include <systemc.h>
#include <iostream>
enum mealy_state_t {
ML_S0 = 0,
ML_S1 = 1,
ML_S10 = 2,
ML_S101 = 3
};
SC_MODULE(DetectorMealy) {
sc_in<bool> clk;
sc_in<bool> rst;
sc_in<bool> bit_in;
sc_out<bool> detected;
sc_signal<int> state;
sc_signal<int> next_state;
void state_register() {
if (rst.read()) state.write(ML_S0);
else state.write(next_state.read());
}
void next_state_and_outputs() {
int s = state.read();
bool b = bit_in.read();
int ns = s;
bool out = false;
switch (s) {
case ML_S0: ns = b ? ML_S1 : ML_S0; break;
case ML_S1: ns = b ? ML_S1 : ML_S10; break;
case ML_S10: ns = b ? ML_S101 : ML_S0; break;
case ML_S101:
if (b) { ns = ML_S1; out = true; } // <-- Mealy: output on transition
else { ns = ML_S10; out = false; }
break;
}
next_state.write(ns);
// Mealy output: function of state AND input.
detected.write(out);
}
SC_CTOR(DetectorMealy) {
SC_METHOD(state_register);
sensitive << clk.pos();
dont_initialize();
SC_METHOD(next_state_and_outputs);
sensitive << state << bit_in;
}
};
// (BitDriver and Mon modules identical to detector_moore.cpp — omitted here
// for brevity; in the standalone file paste them in unchanged.)
SC_MODULE(BitDriver) {
sc_in<bool> clk;
sc_out<bool> bit_out;
const char* pattern = "10110110";
int idx = 0;
void drive() {
if (pattern[idx] == '\0') bit_out.write(false);
else { bit_out.write(pattern[idx] == '1'); idx++; }
}
SC_CTOR(BitDriver) { SC_METHOD(drive); sensitive << clk.pos(); }
};
SC_MODULE(Mon) {
sc_in<bool> clk;
sc_in<bool> bit_in;
sc_in<bool> detected;
void watch() {
std::cout << "[" << sc_time_stamp() << "] bit=" << bit_in.read()
<< " detected=" << detected.read() << "\n";
}
SC_CTOR(Mon) {
SC_METHOD(watch);
sensitive << clk.pos();
dont_initialize();
}
};
int sc_main(int /*argc*/, char* /*argv*/[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst;
sc_signal<bool> bit_wire;
sc_signal<bool> detected_wire;
DetectorMealy det("det");
det.clk(clk); det.rst(rst); det.bit_in(bit_wire); det.detected(detected_wire);
BitDriver drv("drv");
drv.clk(clk); drv.bit_out(bit_wire);
Mon mon("mon");
mon.clk(clk); mon.bit_in(bit_wire); mon.detected(detected_wire);
rst.write(true);
sc_start(10, SC_NS);
rst.write(false);
sc_start(120, SC_NS);
sc_stop();
return 0;
}
Expected output:
[10 ns] bit=0 detected=0
[20 ns] bit=1 detected=0
[30 ns] bit=0 detected=0
[40 ns] bit=1 detected=0
[50 ns] bit=1 detected=1
[60 ns] bit=0 detected=0
[70 ns] bit=1 detected=0
[80 ns] bit=1 detected=1
[90 ns] bit=0 detected=0
[100 ns] bit=0 detected=0
[110 ns] bit=0 detected=0
[120 ns] bit=0 detected=0
[130 ns] bit=0 detected=0
The pattern-detection events fire at t=50 ns and t=80 ns — exactly the clock edges on which the input bit that completes the 1011 pattern arrives. Compare to the Moore variant which fired at t=60 ns and t=90 ns — one clock period later. The Mealy variant is one cycle faster.
Side-by-side trace comparison
Bit arriving: 1 0 1 1 0 1 1 0 0 0
Cycle (10ns): t10 t20 t30 t40 t50 t60 t70 t80 t90 t100
Moore detected: 0 0 0 0 0 1 0 0 1 0
▲ ▲
└ ONE CYCLE LATE
Mealy detected: 0 0 0 0 1 0 0 1 0 0
▲ ▲
└ SAME CYCLE AS INPUT
Both detectors detect the same two pattern matches in the input stream 10110110. The Moore detector reports them one cycle later than the Mealy detector. This is the core trade-off.
Decision table: Moore vs Mealy vs registered Mealy
Property Moore Mealy Registered Mealy
───────────────────────────────── ───────────────── ───────────── ────────────────
Output is function of... state only state + inputs state + inputs,
then passed through
one output register
Output timing 1 cycle after same cycle as 1 cycle after
causing input causing input causing input
Output glitches? No (registered) Yes (combina- No (re-registered)
tional, follows
every input edge)
State count for a given problem More Fewer Fewer (same as
Mealy)
Composes safely with downstream Yes Risky (downstream Yes
clocked logic sees glitches)
Synthesizable in standard flows Yes Yes (with timing Yes
closure care)
Use when... Output drives a Need fastest Need Mealy state
register or a possible output count + Moore
protocol that and downstream output timing
samples on clock samples on clock (= timing closure
edges only edges only on Mealy)
Avoid when... Need latency below Downstream is Output latency is
one cycle asynchronous or irrelevant
combinational
A registered Mealy is the third option that combines the state-count advantage of Mealy with the timing-closure advantage of Moore. It is the Mealy detector with one extra register on its output. Complete program:
// file: detector_registered_mealy.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// detector_registered_mealy.cpp -o detector_registered_mealy -lsystemc
#include <systemc.h>
#include <iostream>
enum reg_mealy_state_t { RM_S0=0, RM_S1=1, RM_S10=2, RM_S101=3 };
SC_MODULE(DetectorRegisteredMealy) {
sc_in<bool> clk;
sc_in<bool> rst;
sc_in<bool> bit_in;
sc_out<bool> detected;
sc_signal<int> state;
sc_signal<int> next_state;
sc_signal<bool> detected_comb; // The combinational Mealy output
sc_signal<bool> detected_reg; // The registered version
void state_register() {
if (rst.read()) {
state.write(RM_S0);
detected_reg.write(false);
} else {
state.write(next_state.read());
detected_reg.write(detected_comb.read()); // register the Mealy output
}
}
void next_state_and_outputs() {
int s = state.read();
bool b = bit_in.read();
int ns = s;
bool out = false;
switch (s) {
case RM_S0: ns = b ? RM_S1 : RM_S0; break;
case RM_S1: ns = b ? RM_S1 : RM_S10; break;
case RM_S10: ns = b ? RM_S101 : RM_S0; break;
case RM_S101:
if (b) { ns = RM_S1; out = true; }
else { ns = RM_S10; }
break;
}
next_state.write(ns);
detected_comb.write(out);
}
void output_driver() {
detected.write(detected_reg.read());
}
SC_CTOR(DetectorRegisteredMealy) {
SC_METHOD(state_register);
sensitive << clk.pos();
dont_initialize();
SC_METHOD(next_state_and_outputs);
sensitive << state << bit_in;
SC_METHOD(output_driver);
sensitive << detected_reg;
}
};
SC_MODULE(BitDriver) {
sc_in<bool> clk;
sc_out<bool> bit_out;
const char* pattern = "10110110";
int idx = 0;
void drive() {
if (pattern[idx] == '\0') bit_out.write(false);
else { bit_out.write(pattern[idx] == '1'); idx++; }
}
SC_CTOR(BitDriver) { SC_METHOD(drive); sensitive << clk.pos(); }
};
SC_MODULE(Mon) {
sc_in<bool> clk;
sc_in<bool> bit_in;
sc_in<bool> detected;
void watch() {
std::cout << "[" << sc_time_stamp() << "] bit=" << bit_in.read()
<< " detected=" << detected.read() << "\n";
}
SC_CTOR(Mon) {
SC_METHOD(watch);
sensitive << clk.pos();
dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst, bit_wire, detected_wire;
DetectorRegisteredMealy det("det");
det.clk(clk); det.rst(rst); det.bit_in(bit_wire); det.detected(detected_wire);
BitDriver drv("drv");
drv.clk(clk); drv.bit_out(bit_wire);
Mon mon("mon");
mon.clk(clk); mon.bit_in(bit_wire); mon.detected(detected_wire);
rst.write(true);
sc_start(10, SC_NS);
rst.write(false);
sc_start(120, SC_NS);
sc_stop();
return 0;
}
Expected output:
[10 ns] bit=0 detected=0
[20 ns] bit=1 detected=0
[30 ns] bit=0 detected=0
[40 ns] bit=1 detected=0
[50 ns] bit=1 detected=0
[60 ns] bit=0 detected=1
[70 ns] bit=1 detected=0
[80 ns] bit=1 detected=0
[90 ns] bit=0 detected=1
[100 ns] bit=0 detected=0
[110 ns] bit=0 detected=0
[120 ns] bit=0 detected=0
[130 ns] bit=0 detected=0
The registered-Mealy output appears at t=60 ns and t=90 ns — one cycle later than the pure Mealy variant but at the same cycle as the Moore variant. The internal state encoding is still four states (Mealy's compactness preserved), but the output is registered (Moore's timing closure recovered). Total combinational depth on the output path is now one signal hop instead of the full Mealy combinational chain.
The choice among the three is fundamentally a question of where in the design hierarchy you can absorb the one-cycle latency. A CPU controller FSM that drives the register-file write-enable wants Moore (the write-enable must be stable for the entire cycle when it is asserted, no glitches). A serial-protocol receiver FSM whose output is "received byte valid" might want Mealy (the consumer is downstream pipeline that samples on clock edges anyway, and one cycle saved on receive latency matters). A registered Mealy is the right choice when the FSM's natural state encoding is Mealy but the synthesis tool reports timing violations on the output combinational path.
Performance note: combinational sensitivity lists
The combinational process's sensitivity list is sensitive << state << bit_in — every input the process reads must appear in the sensitivity list. If you omit bit_in, the process will not wake on input changes; the FSM will only react on clock edges via the state-update chain, and Mealy semantics will fail (the output will not reflect input changes until the next clock edge, defeating the purpose of Mealy). Omit state and the next-state logic stops working after the first input change. SystemC has no equivalent of Verilog's always @(*) auto-sensitivity — you must list every read. Linting tools catch most omissions; experienced reviewers catch the rest.
This concludes the Intermediate section. You now have the kernel mental model of FSMs in SystemC, two complete worked examples that demonstrate the Moore-vs-Mealy timing difference cycle-by-cycle, and a decision framework for picking the right variant. The Advanced section drills into the LRM corners.
Advanced: Edge Cases & LRM Corners
The Beginner and Intermediate sections cover what 95% of FSM code does. This section is the other 5% — LRM corners that tutorials skip, bugs that take a senior engineer half a day to track down, and constructs you should know exist even if you do not reach for them weekly.
Corner 1: async-input glitching in Mealy FSMs
The Mealy detector in the Intermediate section worked because bit_in only changed on clock edges (the BitDriver was a clocked process). In real designs, FSM inputs often come from asynchronous sources — a button press through a debouncer that has not fully settled, a signal from another clock domain through a synchronizer that briefly sits at metastability, a glitch on an interrupt line during a power-supply transient. Any of those can cause bit_in to wiggle between clock edges. A Mealy output that reads bit_in will wiggle with it.
The following program demonstrates the bug. We change BitDriver to write its bit off the clock — at a fixed simulation time mid-cycle — and watch the Mealy detector's detected output glitch:
// file: mealy_glitch.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// mealy_glitch.cpp -o mealy_glitch -lsystemc
#include <systemc.h>
#include <iostream>
enum mealy_state_t { ML_S0=0, ML_S1=1, ML_S10=2, ML_S101=3 };
SC_MODULE(DetectorMealy) {
sc_in<bool> clk;
sc_in<bool> rst;
sc_in<bool> bit_in;
sc_out<bool> detected;
sc_signal<int> state, next_state;
void state_register() {
if (rst.read()) state.write(ML_S0);
else state.write(next_state.read());
}
void next_state_and_outputs() {
int s = state.read(); bool b = bit_in.read();
int ns = s; bool out = false;
switch (s) {
case ML_S0: ns = b ? ML_S1 : ML_S0; break;
case ML_S1: ns = b ? ML_S1 : ML_S10; break;
case ML_S10: ns = b ? ML_S101 : ML_S0; break;
case ML_S101:
if (b) { ns = ML_S1; out = true; }
else { ns = ML_S10; }
break;
}
next_state.write(ns);
detected.write(out);
}
SC_CTOR(DetectorMealy) {
SC_METHOD(state_register); sensitive << clk.pos(); dont_initialize();
SC_METHOD(next_state_and_outputs); sensitive << state << bit_in;
}
};
SC_MODULE(AsyncDriver) {
sc_out<bool> bit_out;
void run() {
// Walk through pattern 1011 then glitch around the detection moment.
bit_out.write(true); wait(10, SC_NS); // bit=1 — drive S0->S1
bit_out.write(false); wait(10, SC_NS); // bit=0 — drive S1->S10
bit_out.write(true); wait(10, SC_NS); // bit=1 — drive S10->S101
bit_out.write(true); wait(5, SC_NS); // bit=1 — drive S101->S1, detected!
bit_out.write(false); wait(1, SC_NS); // GLITCH for 1ns: drops to 0
bit_out.write(true); wait(4, SC_NS); // recovers to 1
bit_out.write(false);
wait(30, SC_NS);
sc_stop();
}
SC_CTOR(AsyncDriver) { SC_THREAD(run); }
};
SC_MODULE(Watcher) {
sc_in<bool> detected;
void log() {
std::cout << "[" << sc_time_stamp() << "] detected="
<< detected.read() << "\n";
}
SC_CTOR(Watcher) {
SC_METHOD(log); sensitive << detected; dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst, bit_wire, detected_wire;
DetectorMealy det("det");
det.clk(clk); det.rst(rst); det.bit_in(bit_wire); det.detected(detected_wire);
AsyncDriver drv("drv");
drv.bit_out(bit_wire);
Watcher w("w");
w.detected(detected_wire);
rst.write(true);
sc_start(10, SC_NS);
rst.write(false);
sc_start();
return 0;
}
Expected output:
[40 ns] detected=1
[45 ns] detected=0
[46 ns] detected=1
[50 ns] detected=0
Notice the four value_changed_events on detected in the span of just 10 ns. The intended detection happens at t=40 ns (Mealy fires combinationally on the input rising to 1 that completes the 1011 pattern). Then the input glitches: at t=45 ns the input drops to 0, the combinational process re-evaluates, the FSM's transition computation says "state S101 + input 0 → S10, output false", and detected drops to 0. At t=46 ns the input recovers to 1, the combinational process re-evaluates, the output goes back to 1. At t=50 ns the input drops to 0 for real, the combinational process re-evaluates, and the output settles to 0.
A downstream consumer sensitive to detected saw it fire twice instead of once. If that consumer is a counter, you count two pattern matches when only one happened. If it is a write-enable to a register, you write twice. If it is an interrupt to a CPU model, you take the interrupt twice. The bug is real and shows up in any real-world deployment where FSM inputs are not clock-synchronized.
The fix has two parts. First, synchronize asynchronous inputs to the FSM clock before they reach the FSM. The canonical fix is a two-flop synchronizer (an sc_signal<bool> clocked register, then another). Second, prefer Moore for any FSM whose inputs may be glitchy. Moore outputs are immune to input glitches by construction — the output reads only state, which only changes on clock edges. Use Mealy only when you can guarantee inputs are clock-synchronous, or when the consumer is itself clocked and samples only on clock edges (in which case the glitches are unobserved).
Corner 2: the dont_initialize() mis-placement bug
In the Beginner section we noted that dont_initialize() belongs on the clocked process, not on the combinational process. Here is what happens when you get it wrong.
// In TrafficLight's constructor, BUGGED VERSION:
SC_METHOD(state_register);
sensitive << clk.pos();
dont_initialize();
SC_METHOD(next_state_and_outputs);
sensitive << state << tick;
dont_initialize(); // <-- BUG: prevents output from initializing
With dont_initialize() on the combinational process, the simulation starts with state = 0, tick = 0, next_state = 0, next_tick = 0, red = false, yellow = false, green = false. The combinational process does not run at initialization, so the outputs sit at false — all three lights are off. At the first clock edge (t=10 ns), with rst high, the state-register process runs and writes state = ST_RED (which is 0) and tick = 0. The kernel's update phase commits both — but state did not change (it was already 0) and tick did not change either. The value-changed events on state and tick do not fire. The combinational process never wakes. The outputs stay at false, false, false.
By t=20 ns reset is low, the state register reads next_state = 0 (still default-constructed) and writes state = 0 — no change, no event, no combinational wake. The FSM is stuck with all outputs false forever, no matter how many clock edges arrive. The state machine has no internal contradiction; it simply never gets the kick of an initial combinational evaluation that establishes correct outputs.
This bug is particularly nasty because it is invisible in a single-process FSM (where the bugged dont_initialize() would suppress both the combinational and the clocked work, which would be obviously wrong on inspection), and it is invisible if the state register's reset path writes a different value than the default-constructed 0 (because then state does change at the first reset-released clock edge, the value-changed event fires, and the combinational process wakes — masking the bug for designs where ST_RESET != 0). The bug only surfaces in the specific case of "Moore FSM, two processes, default-constructed state happens to equal the reset state value, dont_initialize() accidentally on the wrong process". The fix is one-line: remove dont_initialize() from the combinational process.
Corner 3: "last writer wins" on next_state
Suppose you split the combinational process into two halves — one process computes next_state for transitions on input A, another computes next_state for transitions on input B. Both processes write next_state. This is the "multiple writers, one signal, same evaluate phase" case from Part 3.
// BUGGED two-driver next-state pattern
void next_state_A() {
if (a_input.read()) next_state.write(ST_FROM_A);
}
void next_state_B() {
if (b_input.read()) next_state.write(ST_FROM_B);
}
Per the standard, when two processes write next_state in the same evaluate phase, the kernel commits one value during the update phase — and which one is implementation-defined. The Accellera 2.3.4 PoC kernel uses insertion order (the second-registered process wins), but a different kernel may use LIFO order, and a randomized-for-stress-testing kernel will make this nondeterministic. The bug is that the FSM's next-state computation depends on the order independent processes ran inside the same delta, which is exactly the case Part 3 warned against.
The fix is always combine all next-state logic in one process, with one set of case arms or one cascading if/else chain that handles every combination of inputs. The pattern is canonical:
void next_state_and_outputs() {
int s = state.read();
bool a = a_input.read();
bool b = b_input.read();
int ns = s;
switch (s) {
case ST_X:
if (a) ns = ST_FROM_A;
else if (b) ns = ST_FROM_B;
else ns = ST_X;
break;
case ST_FROM_A:
ns = ST_X;
break;
case ST_FROM_B:
ns = ST_X;
break;
default:
ns = ST_X;
break;
}
next_state.write(ns); // single writer
}
One process, one writer, deterministic behavior on every kernel. Even when the FSM has 12 inputs and 30 states, the single-process pattern scales — the case statement grows, but it remains a single process. Resist any urge to "modularize" by splitting it.
Corner 4: registered-Mealy timing pipeline and where the cycle moves
The registered-Mealy variant introduces a one-cycle output delay to clean up timing. The subtle question is: which cycle does it delay? The naive answer is "the cycle that the combinational Mealy output appeared." But the correct answer involves a careful look at the FSM's state-edge alignment.
Consider the Mealy detector at t=40 ns where the input completes the pattern. The combinational process computes detected_comb = true at t=40 ns (in the delta after the t=30 ns clock edge updated state to ML_S101, with the input arriving at t=40 ns on the same clock edge as the new state). The output register samples detected_comb on the t=40 ns clock edge — but wait, which t=40 ns clock edge? The state-register process and the output-register update run from the same edge.
In SystemC the answer is determined by the update-phase semantics. The state-register process at t=40 ns reads detected_comb (the current value, which is the combinational output computed in the delta after t=30 ns) and writes detected_reg. The kernel's runnable-set ordering across the two writers (state and detected_reg) within the state-register process is sequential within the process — both writes happen, both buffer, both commit in the next update phase. The combinational process then re-runs because state changed, and computes a fresh detected_comb based on the new state.
The result: in a registered-Mealy detector, the output detected_reg at t=50 ns (the cycle after detection) reflects the combinational Mealy output that was valid during the t=30 → t=40 ns cycle. In other words, the output appears one cycle after the Moore variant — which is two cycles after the input. That extra cycle is the price of the output register; you trade timing closure for latency.
This is rarely what designers want. The more common and useful pattern is: structure the output register so it samples on the same clock edge that updates state, so the output appears at the same cycle as the Moore variant (one cycle after the input). This requires sampling detected_comb based on the next state, not the current state — which means the state-register process writes both state and detected_reg from next_state and a "next-detected" value computed by the combinational process. Implementation:
// file: detector_fast_registered_mealy.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// detector_fast_registered_mealy.cpp -o detector_fast_registered_mealy \
// -lsystemc
//
// In this "fast" registered-Mealy variant, the combinational process writes
// both next_state and next_detected. The state-register process samples
// BOTH on the same clock edge, so the registered output appears one cycle
// after the input (matching Moore timing) — instead of two cycles.
#include <systemc.h>
#include <iostream>
enum fm_state_t { FM_S0=0, FM_S1=1, FM_S10=2, FM_S101=3 };
SC_MODULE(DetectorFastRegMealy) {
sc_in<bool> clk;
sc_in<bool> rst;
sc_in<bool> bit_in;
sc_out<bool> detected;
sc_signal<int> state;
sc_signal<int> next_state;
sc_signal<bool> next_detected;
sc_signal<bool> detected_reg;
void state_register() {
if (rst.read()) {
state.write(FM_S0);
detected_reg.write(false);
} else {
state.write(next_state.read());
detected_reg.write(next_detected.read()); // sample next_detected, not detected_comb
}
}
void next_state_and_outputs() {
int s = state.read();
bool b = bit_in.read();
int ns = s;
bool out = false;
switch (s) {
case FM_S0: ns = b ? FM_S1 : FM_S0; break;
case FM_S1: ns = b ? FM_S1 : FM_S10; break;
case FM_S10: ns = b ? FM_S101 : FM_S0; break;
case FM_S101:
if (b) { ns = FM_S1; out = true; }
else { ns = FM_S10; }
break;
}
next_state.write(ns);
next_detected.write(out);
}
void output_driver() {
detected.write(detected_reg.read());
}
SC_CTOR(DetectorFastRegMealy) {
SC_METHOD(state_register);
sensitive << clk.pos();
dont_initialize();
SC_METHOD(next_state_and_outputs);
sensitive << state << bit_in;
SC_METHOD(output_driver);
sensitive << detected_reg;
}
};
SC_MODULE(BitDriver) {
sc_in<bool> clk;
sc_out<bool> bit_out;
const char* pattern = "10110110";
int idx = 0;
void drive() {
if (pattern[idx] == '\0') bit_out.write(false);
else { bit_out.write(pattern[idx] == '1'); idx++; }
}
SC_CTOR(BitDriver) { SC_METHOD(drive); sensitive << clk.pos(); }
};
SC_MODULE(Mon) {
sc_in<bool> clk;
sc_in<bool> bit_in;
sc_in<bool> detected;
void watch() {
std::cout << "[" << sc_time_stamp() << "] bit=" << bit_in.read()
<< " detected=" << detected.read() << "\n";
}
SC_CTOR(Mon) {
SC_METHOD(watch);
sensitive << clk.pos();
dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst, bit_wire, detected_wire;
DetectorFastRegMealy det("det");
det.clk(clk); det.rst(rst); det.bit_in(bit_wire); det.detected(detected_wire);
BitDriver drv("drv");
drv.clk(clk); drv.bit_out(bit_wire);
Mon mon("mon");
mon.clk(clk); mon.bit_in(bit_wire); mon.detected(detected_wire);
rst.write(true);
sc_start(10, SC_NS);
rst.write(false);
sc_start(120, SC_NS);
sc_stop();
return 0;
}
Expected output:
[10 ns] bit=0 detected=0
[20 ns] bit=1 detected=0
[30 ns] bit=0 detected=0
[40 ns] bit=1 detected=0
[50 ns] bit=1 detected=0
[60 ns] bit=0 detected=1
[70 ns] bit=1 detected=0
[80 ns] bit=1 detected=0
[90 ns] bit=0 detected=1
[100 ns] bit=0 detected=0
[110 ns] bit=0 detected=0
[120 ns] bit=0 detected=0
[130 ns] bit=0 detected=0
For this particular FSM the monitor sees identical detection cycles in both registered variants — both fire at t=60 ns and t=90 ns, one cycle after the Mealy variant. The reason is that for this state encoding, the value held in detected_reg at any given clock edge is the same regardless of whether you registered detected_comb (slow) or next_detected (fast); both correctly capture "did we just complete the pattern" on the edge after detection. The distinction matters — and the cycle alignment visibly differs — when the FSM's combinational output depends on next-cycle inputs (not the current input), or when you cascade two registered-Mealy FSMs and need the second one to receive a sample of the first's output that aligns with the second's state-register cycle. In those cases the "fast" pattern (register next_* values) avoids an extra cycle of pipeline latency that the "slow" pattern (register *_comb values) introduces.
The lesson: registered Mealy is not just "stick a register after the output"; it is "decide which clock edge samples the output, and structure the combinational process to write a value that is valid for that edge." Get this wrong by one cycle and your FSM has a one-cycle output offset bug that will not be caught until system integration.
Version differences
The kernel behavior described here — SC_METHOD semantics, sensitivity-list mechanics, sc_signal update-phase delays, dont_initialize() effect on initialization — is identical across SystemC 2.3.1, 2.3.3, and 2.3.4. No changes have been made to these mechanisms in these releases. Differences between these releases are in build system, C++17 conformance, TLM-2.0 utilities, and some sc_vector helpers. Anything in this post compiles and behaves the same on any 2.3.x release. Pre-2.3 releases (2.2.0 from 2007) had different sc_event binding semantics but the FSM patterns shown here did not depend on those; the recommendation is to upgrade rather than work around the differences.
Hands-on exercise
Build a TCP-style connection-establishment FSM. Four states: LISTEN, SYN_RCVD, ESTABLISHED, CLOSED. The FSM has three inputs (syn, ack, rst) and one output (conn_valid — high when in ESTABLISHED). The transitions:
- From
LISTEN: onsyn, go toSYN_RCVD. Onrst, go toCLOSED. Otherwise stay. - From
SYN_RCVD: onack, go toESTABLISHED. Onrst, go toCLOSED. Otherwise stay. - From
ESTABLISHED: onrst, go toCLOSED. Otherwise stay. - From
CLOSED: unconditionally go toLISTENon the next clock (so the FSM auto-recovers and is ready for a new connection).
Build this as a Moore FSM with registered outputs. Drive it with a clocked stimulus that sends a valid SYN → ACK handshake, then a rst, then a second SYN → ACK handshake. Watch conn_valid go high one cycle after the ACK arrives, stay high for the duration of the connection, go low one cycle after rst, and then come back high after the second handshake completes.
When you have that working, extend it. Add a fifth state, FIN_WAIT, that the FSM enters on a new input fin from either ESTABLISHED or SYN_RCVD. From FIN_WAIT, on an ack input, go to CLOSED. This models a graceful close. Predict the cycle count for "open + close" before you run, then verify.
When that works, rebuild the same logic as a Mealy variant. The output conn_valid should now go high in the same cycle as the ACK that completes the handshake, not one cycle later. Compare the trace.
No solution is provided. The lessons live in the building.
Hints
- Use the same
state_register+next_state_and_outputstwo-process pattern from the Beginner section. - Use an
enumfor the state encoding; do not try to pack the state into aboolor compute it from bits of the inputs. - The Moore output
conn_valid.write(s == ST_ESTABLISHED)is one line in the combinational process. - The transitions from each state should be a single
casearm in aswitch (s); if you find yourself writing two case arms for the same state, you have a bug. - Drive your inputs with a clocked
SC_METHODand a smallintstep counter (if (step==0) syn.write(true); else if (step==1) ...). AvoidSC_THREADfor the stimulus generator unless you specifically need to drive inputs at non-clock-edge times (which would expose you to the Mealy glitching issue). - For the Mealy variant: remember to include all inputs the combinational process reads (
syn,ack,rst,fin) in its sensitivity list. Omitting any input from the list is the most common source of "Mealy output does not update" bugs.
Common mistakes
- Forgetting
dont_initialize()on the clocked state-register process. Without it, the state register runs once at initialization before any clock edge — and readsnext_state.read()which is the default-constructed0. For most FSMs,0is the reset state, so the bug looks benign. But the moment your FSM's reset state is not0(e.g., you reordered the enum), initialization writes the wrong starting state and the FSM begins in a corrupt configuration. The bug is invisible in waveforms because the FSM still runs — it just starts in the wrong state. Fix: always calldont_initialize()on the state-register process.
- Wrong sensitivity list on the combinational process. SystemC has no
always @(*)auto-sensitivity. If you readbit_inin the combinational process but forgetsensitive << bit_in, the process will not wake on input changes. The FSM appears to "freeze" — it only reacts to clock edges via the state-update chain. The bug looks like a Mealy FSM that has degenerated into Moore semantics. Fix: every read inside the combinational process must correspond to one entry insensitive << ....
- Writing the state register from the combinational process. Inside the combinational process, write
next_state, neverstate. If you writestatedirectly, you bypass the clock and create an asynchronous latch network — the FSM has no clocked storage and the state changes whenever an input wiggles. The simulation hangs on the first input pattern that requires the FSM to remember a previous input. Fix: the onlystate.write(...)calls in the entire module live inside the clocked state-register process.
- Mealy output glitching from an asynchronous input. Demonstrated in Advanced Corner 1. A Mealy output reads the FSM's inputs combinationally. If an input glitches between clock edges, the output glitches with it. Downstream consumers that sample on every value-changed event see the glitch and react. Fix: either synchronize the input to the FSM clock (two-flop synchronizer) before it reaches the FSM, or use Moore semantics so the output reads only
state, which is glitch-free.
- Mixing reset polarities between the state register and the combinational process. The combinational process should not read
rst— it should compute next-state and outputs fromstateand inputs as if reset did not exist. Reset is handled only in the state-register process. If you accidentally readrstin the combinational process and gate outputs onrst == false, the FSM's outputs go low during reset (correct) but the next-state computation continues to use the pre-resetstate, leading to one-cycle stall bugs at reset release. Fix: keep reset logic entirely in the clocked state-register process; the combinational process is reset-agnostic.
Recap
After working through this post you can now:
- State the two-process FSM pattern (clocked state-register + combinational next-state-and-outputs) and explain why
sc_signal<state_t>is the SystemC flip-flop. - Predict the exact cycle on which a Moore FSM's output changes given an input transition (one cycle after the input).
- Predict the exact cycle on which a Mealy FSM's output changes given an input transition (same cycle as the input).
- Choose among Moore, Mealy, and registered-Mealy by reasoning about latency, glitch tolerance, and downstream consumer semantics.
- Build a complete two-process FSM from scratch — traffic-light controller, serial bit detector, TCP handshake — without consulting reference code.
- Diagnose and fix the four most common FSM bugs:
dont_initialize()mis-placement, wrong sensitivity list, combinational write to state register, and asynchronous-input glitching. - Read a state diagram and tell at a glance whether it is Moore (outputs labeled on nodes) or Mealy (outputs labeled on transition arrows).
- Recognize when "last writer wins" on
next_stateis about to cause an implementation-defined behavior bug, and refactor to a single-writer combinational process.
Further reading
Standards
- IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §5.2.16 (
SC_METHODsemantics), §5.2.18 (sensitivity lists anddont_initialize), §6.4 (sc_signalwrite semantics), §4.2.1.3 (update phase and delta cycles).
Vendor and consortium documents
- Accellera Systems Initiative, SystemC 2.3.x User Guide, sections on process kinds and channel mechanics.
Textbooks
- Bhasker, A SystemC Primer (2nd ed.), ch. 5 — two-process FSM idiom, Moore vs Mealy decomposition.
- Doulos, SystemC Golden Reference Guide, ch. 4 —
SC_METHODand sensitivity-list conventions. - Grötker, Liao, Martin, and Swan, System Design with SystemC — channel and process chapters relevant to FSM modeling.
Training notes
- Berkeley CS152 lecture notes, FSM design module — pedagogical framing of Moore vs Mealy timing and the registered-output trade-off.
- MIT 6.004 lecture notes, FSM section — RTL pedagogy for transitions, encoding, and output timing.
Real-world references
- CV32E40P
cv32e40p_controller.sv(public source, OpenHW Group) — production CPU controller FSM using the registered-output Moore pattern this post teaches. - Ibex
ibex_controller.sv(public source, lowRISC) — independent corroboration of the same FSM idiom in a different open-source RISC-V core.
Next in this section
→ Part 5: Memories & Register Files — modeling storage arrays as plain C++ arrays behind clocked write and combinational/registered read processes, the synchronous read/write timing pattern, multi-port register-file decomposition, and the worked example of a 32×32-bit RISC-V register file. Read it here: 12. SystemC Tutorial — Memories & Register Files.
Comments (0)
Leave a Comment