9. SystemC Tutorial - Decoders & Case Analysis
Why this matters
Rewritten 2026-06-05 with deeper first-principles material.
A decoder is the smallest piece of logic that earns the word "intelligence." It looks at a bundle of bits, decides which of several mutually exclusive situations the bits describe, and lights up exactly the outputs that situation demands. That pattern — read a code, select a case, drive a bundle of outputs — is the most common shape in all of digital design. The address decoder that decides which peripheral on a bus a transaction is for is a decoder. The opcode decoder at the front of every CPU that turns a 32-bit instruction word into a fistful of control signals is a decoder. The 7-segment driver that turns a 4-bit nibble into seven lamp-enable lines is a decoder. The one-hot select generator inside every multiplexer is a decoder. The priority encoder that picks the highest-numbered pending interrupt runs a decoder backwards. When you understand decoders, you understand case analysis in hardware — and case analysis is what separates a wire from a circuit.
This matters in SystemC specifically because the language gives you no guard rails. SystemVerilog has always_comb, which auto-builds the sensitivity list and warns you when a code path forgets to assign an output (an inferred latch). SystemC has neither. You list every input by hand in the sensitivity expression, and you are personally responsible for driving every output on every path through your case analysis. Forget one output in one case and SystemC will not warn you — it will silently hold that output at whatever value it had last, which is the single most common decoder bug in the language and the one that wastes the most engineer-hours in pre-silicon bring-up. A control signal that should have been low stays high for one extra instruction because the decoder forgot to clear it, the CPU model writes a register it should not have, and the failure surfaces three thousand cycles later as a corrupted result with no obvious cause.
By the end of this post you will be able to write a combinational decoder in SystemC from scratch, predict its outputs from its inputs without running the simulator, recognize and prevent the stale-output bug on sight, choose between one-hot and binary encodings for a given job, extract and sign-extend immediates from a scattered instruction encoding, and handle illegal inputs deliberately rather than by accident. We will build the ideas on a deliberately tiny generic decoder first, then apply every one of them to the worked example that DV engineers actually care about: the RV32I instruction decoder. That is the bar.
Prerequisites
- Part 1 — Modules, Ports & Signals. You need to be comfortable declaring an
SC_MODULE, bindingsc_signalchannels tosc_in/sc_outports, and registering a process withSC_METHODinsideSC_CTOR. A decoder is nothing but ports plus one method, so this is the load-bearing prerequisite. - Part 4 — Processes & Sensitivity (SC_METHOD vs SC_THREAD). You need to know why a combinational block is an
SC_METHOD(runs to completion, nowait()), how a static sensitivity list is built withsensitive << ..., and why the sensitivity list must name every input the process reads. - Part 8 — Combinational Modeling Patterns. The "set safe defaults, then compute" discipline and the rule that a combinational
SC_METHODmust assign all outputs on all paths come from this part. A decoder is a direct application of those rules to case analysis, so re-read it if the defaults-before-logic idea is hazy. - SystemC 2.3.x installed. All examples compile with a C++17 compiler (
g++9 or newer, orclang++10 or newer) against any 2.3.x SystemC build. The installer posts earlier in this series cover environment setup for Linux and macOS.
Three ideas from those parts get heavy use here: the combinational SC_METHOD (the process kind a decoder uses), static sensitivity (the sensitive << instr that wakes the decoder when its input changes), and sc_signal's update-phase delay (which is why a write you make in a decoder becomes visible to readers in the next delta, not instantly). If any is unfamiliar, revisit the relevant part before continuing.
Mental model (first principles)
Strip away the application and a decoder is one sentence of logic: given an input code, select exactly one case out of a fixed set, and for that case drive a complete bundle of outputs. Three properties are baked into that sentence, and every decoder bug is a violation of one of them.
The first property is mutual exclusivity: the cases must not overlap. For any input code, exactly one case applies. In hardware this is what lets a decoder be a tree of gates with no contention; in code this is what lets a switch work, because switch evaluates one arm. If two of your cases can match the same input, you do not have a decoder — you have a priority resolver, and you must make the priority explicit.
The second property is completeness of selection: every possible input code lands in some case, including the codes that are not supposed to occur. A 3-bit code has eight possible values; if you define behavior for six of them, the other two still arrive in simulation eventually (from a corrupted stream, a directed test, or a constrained-random generator), and your decoder must do something defined when they do. That "something" is the default case. A decoder without a default is a decoder that has undefined behavior on inputs you did not think about, which is exactly the inputs a verification engineer will throw at it.
The third property is completeness of output: within each case, every output the decoder owns must be assigned a value. This is the property SystemC will not enforce for you, and it is the heart of the whole post. Picture the decoder's outputs as a row of lamps. Each activation of the decoder must set every lamp — on or off — based on the selected case. If a case forgets a lamp, that lamp keeps glowing from whatever the previous input set it to. In real hardware that "memory" is a latch, an unintended storage element that synthesis tools flag as an error. In SystemC there is no flag: the output is an sc_signal, the case did not call write() on it, so it keeps its last committed value silently.
Here is the canonical structure that satisfies all three properties at once, expressed as the shape every decoder in this post will take:
void decode() {
// 1. SELECT field(s): slice the input into the code that picks the case.
unsigned code = /* bits of the input that choose the case */;
// 2. SAFE DEFAULTS: assign EVERY output a defined value up front.
// This single block guarantees output-completeness for all paths.
out_a.write(false);
out_b.write(0);
// ... one default per output ...
// 3. CASE ANALYSIS: one arm per defined code; each arm OVERRIDES only
// the outputs that differ from the defaults.
switch (code) {
case CODE_X: out_a.write(true); out_b.write(7); break;
case CODE_Y: out_b.write(3); break;
// ...
default: /* illegal code: defaults already NOP everything */ break;
}
}
The defaults block is the trick that makes output-completeness automatic. Because every output is written before the switch, every path through the switch — including the ones that forget to touch an output, and including default — leaves every output defined. Each case arm then only has to mention the outputs that differ from the safe NOP state. This is the same "defaults first, then specialize" pattern Part 8 taught for the ALU; a decoder is that pattern applied to a wider fan-out of control outputs.
In SystemC this whole thing is one SC_METHOD:
SC_METHOD(decode);
sensitive << in_code; // every input the method reads is listed here
// NO dont_initialize(): we WANT it to run once at t=0 to establish outputs
Note the absence of dont_initialize(). A clocked process calls it so it does not fire spuriously before the first clock edge. A pure-combinational decoder wants to fire at initialization, because that first run is what gives the outputs defined values at t=0 — before any input has changed. Omit the defaults, or add dont_initialize() by reflex, and the decoder's outputs sit at their default-constructed values until the first input change, which is its own subtle bug.
The kernel mechanics behind this are exactly the sc_signal update-phase semantics from Part 1 and Part 4. When in_code changes, the kernel commits the new value in an update phase and fires its value-changed event. The decoder, statically sensitive to in_code, becomes runnable in the next delta. It reads the new code, writes its outputs (those writes are buffered), and returns. The update phase commits the output writes and fires their value-changed events, waking anything downstream. So a combinational decoder introduces exactly one delta of delay between an input change and the corresponding output change — the same "delta tax" every sc_signal-mediated combinational block pays. It is not a clock cycle; it is a zero-time delta, invisible on a nanosecond timescale, and it is the SystemC modeling artifact that stands in for real gate propagation delay.
%%{init: {'theme':'base','themeVariables':{'primaryColor':'#dbeafe','primaryTextColor':'#1e293b','primaryBorderColor':'#3b82f6','lineColor':'#64748b','secondaryColor':'#f1f5f9'}}}%%
flowchart LR
IN["in_code"] --> SEL["case select\n(switch)"]
SEL --> OA["out_a"]
SEL --> OB["out_b"]
SEL --> OC["out_c"]
DEF["safe defaults\n(every output)"] -.-> SEL
Read the diagram as: the input feeds a case-select, the case-select drives a full bundle of outputs, and a defaults block (dashed, because it runs first and is overridden) guarantees no output is ever left undriven. That is the entire substance of decoder modeling in SystemC. Everything that follows — the RV32I worked example included — is this shape with a wider input, more outputs, and more case arms.
Beginner: First Principles
The simplest decoder that is still interesting is a 3-bit opcode decoder. Three input bits select one of eight operations for a tiny made-up machine; the decoder drives a small bundle of control outputs that the rest of the (imaginary) datapath would consume. We will define the operations, build the decoder, predict its output by hand, then run it on a real kernel to confirm.
Our toy machine has eight opcodes:
| Code (binary) | Mnemonic | Meaning | alu_en |
mem_en |
wr_en |
is_branch |
|---|---|---|---|---|---|---|
000 |
NOP | do nothing | 0 | 0 | 0 | 0 |
001 |
ADD | register add, write result | 1 | 0 | 1 | 0 |
010 |
SUB | register subtract, write result | 1 | 0 | 1 | 0 |
011 |
LOAD | read memory, write result | 0 | 1 | 1 | 0 |
100 |
STORE | write memory, no register write | 0 | 1 | 0 | 0 |
101 |
BRANCH | conditional branch, no write | 1 | 0 | 0 | 1 |
110 |
(reserved) | illegal | 0 | 0 | 0 | 0 |
111 |
(reserved) | illegal | 0 | 0 | 0 | 0 |
Notice the structure of the table before you read the code. Six codes have defined behavior; two are reserved (illegal). The four output columns are a bundle — every row sets all four. The two illegal rows are not blank; they are explicitly all-zero, which is the safe NOP state. A decoder that left those rows undefined would be a decoder with a latch on every output for two of its eight inputs. Here is the SystemC.
// file: opcode_decoder.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// opcode_decoder.cpp -o opcode_decoder -lsystemc
// Run: LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./opcode_decoder
#include <systemc.h>
#include <iostream>
// Opcode encodings for our toy machine (3-bit codes).
static constexpr unsigned OP_NOP = 0; // 000
static constexpr unsigned OP_ADD = 1; // 001
static constexpr unsigned OP_SUB = 2; // 010
static constexpr unsigned OP_LOAD = 3; // 011
static constexpr unsigned OP_STORE = 4; // 100
static constexpr unsigned OP_BRANCH = 5; // 101
// 6 (110) and 7 (111) are reserved/illegal.
SC_MODULE(OpcodeDecoder) {
sc_in<sc_uint<3>> opcode; // the 3-bit code that selects the case
sc_out<bool> alu_en; // enable the ALU
sc_out<bool> mem_en; // enable the memory port
sc_out<bool> wr_en; // write the register file
sc_out<bool> is_branch; // this is a conditional branch
sc_out<bool> illegal; // the opcode matched no defined case
void decode() {
unsigned code = opcode.read();
// --- SAFE DEFAULTS: drive EVERY output before the switch. ---
alu_en.write(false);
mem_en.write(false);
wr_en.write(false);
is_branch.write(false);
illegal.write(false);
// --- CASE ANALYSIS: each arm overrides only what differs. ---
switch (code) {
case OP_NOP:
// all defaults are correct; nothing to override
break;
case OP_ADD:
case OP_SUB:
alu_en.write(true);
wr_en.write(true);
break;
case OP_LOAD:
mem_en.write(true);
wr_en.write(true);
break;
case OP_STORE:
mem_en.write(true);
break;
case OP_BRANCH:
alu_en.write(true);
is_branch.write(true);
break;
default:
// Reserved codes 110 and 111 land here: NOP outputs + illegal flag.
illegal.write(true);
break;
}
}
SC_CTOR(OpcodeDecoder) {
SC_METHOD(decode);
sensitive << opcode; // the only input; the only sensitivity entry
// No dont_initialize(): run at t=0 so outputs are defined immediately.
}
};
SC_MODULE(Stimulus) {
sc_out<sc_uint<3>> opcode;
void drive() {
// Walk every code 0..7, one per nanosecond, so we exercise all 8 cases
// including the two illegal ones.
for (unsigned c = 0; c < 8; ++c) {
opcode.write(c);
wait(1, SC_NS);
}
sc_stop();
}
SC_CTOR(Stimulus) { SC_THREAD(drive); }
};
SC_MODULE(Monitor) {
sc_in<sc_uint<3>> opcode;
sc_in<bool> alu_en, mem_en, wr_en, is_branch, illegal;
void watch() {
std::cout << "[" << sc_time_stamp() << "] code=" << opcode.read()
<< " alu=" << alu_en.read()
<< " mem=" << mem_en.read()
<< " wr=" << wr_en.read()
<< " branch=" << is_branch.read()
<< " illegal="<< illegal.read() << "\n";
}
SC_CTOR(Monitor) {
SC_METHOD(watch);
// Sensitive to every decoder output so we print after each settles.
sensitive << alu_en << mem_en << wr_en << is_branch << illegal;
dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_signal<sc_uint<3>> s_opcode;
sc_signal<bool> s_alu, s_mem, s_wr, s_branch, s_illegal;
OpcodeDecoder dec("dec");
dec.opcode(s_opcode);
dec.alu_en(s_alu); dec.mem_en(s_mem); dec.wr_en(s_wr);
dec.is_branch(s_branch); dec.illegal(s_illegal);
Stimulus stim("stim");
stim.opcode(s_opcode);
Monitor mon("mon");
mon.opcode(s_opcode);
mon.alu_en(s_alu); mon.mem_en(s_mem); mon.wr_en(s_wr);
mon.is_branch(s_branch); mon.illegal(s_illegal);
sc_start();
return 0;
}
Run it in your head before reading the output. The stimulus drives codes 0 through 7, one per nanosecond. For each code the decoder selects a case and the monitor prints the output bundle. Codes 0–5 hit defined cases; codes 6 and 7 fall through to default and raise illegal.
Expected output:
[0 s] code=1 alu=1 mem=0 wr=1 branch=0 illegal=0
[1 ns] code=2 alu=1 mem=0 wr=1 branch=0 illegal=0
[2 ns] code=3 alu=0 mem=1 wr=1 branch=0 illegal=0
[3 ns] code=4 alu=0 mem=1 wr=0 branch=0 illegal=0
[4 ns] code=5 alu=1 mem=0 wr=0 branch=1 illegal=0
[5 ns] code=6 alu=0 mem=0 wr=0 branch=0 illegal=1
[6 ns] code=7 alu=0 mem=0 wr=0 branch=0 illegal=1
Two details in that trace deserve explanation. First, why does the monitor print code=1 on its first line instead of code=0? At t=0 the stimulus thread writes opcode = 0 (its first loop iteration). That write commits in the update phase, the decoder runs and produces the NOP bundle (all zeros). But the monitor is sensitive to the decoder's outputs, and at t=0 the outputs were already at their initialization values of all-zero — writing all-zero again is not a value change, so no value-changed event fires on the outputs, and the dont_initialize() monitor never wakes for code 0. By the time the monitor first wakes, the stimulus has advanced to code 1, whose alu_en/wr_en outputs do change, firing events. This is a teaching artifact of monitoring on outputs rather than the clock; it does not reflect a decoder error. The decoder did decode code 0 correctly — it simply produced the same all-zero bundle it started with.
Second, look at codes 6 and 7: every functional output is 0 and illegal is 1. That is the default arm doing its job. The safe-defaults block already set the four functional outputs to their NOP values; the default arm only had to raise illegal. If we had omitted the defaults block and written each output only inside the cases, codes 6 and 7 would have left alu_en, mem_en, wr_en, and is_branch holding their values from code 5 — alu_en=1, is_branch=1 — and our "illegal NOP" would have been silently executing a branch. That is the stale-output bug, and we will reproduce it deliberately in a moment so you recognize it forever.
The stale-output (inferred-latch) bug
Here is the same decoder with one mistake: the safe-defaults block is deleted, on the assumption that "every case sets what it needs."
// BUGGED decode(): no defaults block. Each case sets only "its" outputs.
void decode_buggy() {
unsigned code = opcode.read();
switch (code) {
case OP_NOP:
alu_en.write(false); mem_en.write(false);
wr_en.write(false); is_branch.write(false);
break;
case OP_ADD:
case OP_SUB:
alu_en.write(true); wr_en.write(true); // forgot mem_en, is_branch
break;
case OP_BRANCH:
alu_en.write(true); is_branch.write(true); // forgot wr_en, mem_en
break;
// ... and crucially, no default at all ...
}
}
Walk a two-input sequence through it. Drive OP_BRANCH (code 5): the case sets alu_en=1, is_branch=1, and leaves mem_en and wr_en at whatever they were. Now drive code 6 (reserved). There is no default, so the switch matches nothing and not a single output is written. Every output holds its value from the BRANCH that came before: alu_en stays 1, is_branch stays 1. The reserved opcode is now indistinguishable from a branch. In synthesis this code infers four latches; in SystemC simulation it produces stale, wrong outputs with no diagnostic whatsoever. The only fix is the discipline from the mental-model section: defaults before the switch, a default arm always. There is no exception, and no decoder in this series ever omits either.
always_comb would warn you about the inferred latches above, and unique case would warn about the missing default at runtime. SystemC gives you neither warning. The discipline that the SV tool enforces for you, you must enforce for yourself in SystemC. Treat "defaults first, default arm always" as a non-negotiable reflex.One-hot vs binary encodings
Our toy decoder took a binary-encoded opcode: three bits naming one of eight cases, decoded by a switch. There is a second common encoding worth understanding because real designs mix both: one-hot, where you use one wire per case and exactly one wire is high at a time.
A one-hot version of the same selection would have eight input wires sel[7:0], and "ADD" means sel == 0b00000010 (only bit 1 set). The decode logic per output becomes a simple OR of the relevant select bits — wr_en = sel[ADD] | sel[SUB] | sel[LOAD] — with no switch needed. The trade-offs are concrete:
| Property | Binary (switch on code) |
One-hot (one wire per case) |
|---|---|---|
| Input wires for N cases | ⌈log2 N⌉ | N |
| Per-output logic | a decode gate / case match | a flat OR of select bits |
| Illegal-code detection | code ≥ defined count, or default |
popcount(sel) ≠ 1 |
| Typical use | instruction opcodes, address fields | FSM state, mux selects, request grids |
| Fan-out / timing | denser wiring, deeper decode | wider wiring, shallow decode |
Neither is "better"; they answer different pressures. Binary minimizes wires and dominates instruction encodings (an opcode is a binary code by nature). One-hot minimizes logic depth and dominates FSM state registers and mux selects, where the flat-OR structure is fast and the illegal-state check (more than one bit set, or zero bits set) is trivial. A decoder that converts binary to one-hot — sel = 1 << code — is itself one of the most common building blocks in hardware; it is literally what a memory address decoder does to pick a row. Recognizing which encoding a given interface uses tells you immediately whether to reach for a switch (binary) or a set of OR expressions (one-hot).
With the generic decoder, the stale-output bug, and the encoding choice all established, we can now apply every one of these ideas to the worked example DV engineers care about.
Intermediate: How It Really Works
The toy decoder taught the shape. Now we apply that exact shape to a real, completely specified instruction set: RV32I, the 32-bit integer base of RISC-V. The decoder we build is the same SC_METHOD with safe defaults and a switch, only wider — the input is a 32-bit instruction word, the "code" that selects the case is a 7-bit opcode field, and the output bundle is the full set of control signals a single-cycle CPU needs, plus a sign-extended immediate. Everything you learned on the 3-bit decoder transfers directly; only the field count grows.
Why fixed-width encodings make decoding easy
RV32I uses fixed 32-bit instructions. The lowest 7 bits (the opcode field) identify the instruction's format; the remaining 25 bits carry register addresses, function codes, and immediate bits. This is a deliberate hardware-friendly choice. Contrast x86, where an instruction is 1 to 15 bytes long and the decoder must first determine the length before it can find the operands — a whole variable-length pre-decode stage. RISC-V eliminates that: the front end always knows it has exactly 32 bits to decode, so case analysis can begin on bit zero.
The deeper trick is that RV32I places the register-address fields at the same bit positions in every format:
rs1is always at bits[19:15]rs2is always at bits[24:20]rdis always at bits[11:7]
Because these positions never move, hardware can slice out the register addresses and start reading the register file before it knows the instruction's format. Decode and register-read overlap. This invariant is why the immediate bits in some formats look scrambled — the encoding scatters immediate bits around the fixed rs1/rs2/rd fields rather than disturbing them. Every rearrangement exists to keep those three fields tappable at constant positions.
The six instruction formats
RV32I defines six formats. Every instruction belongs to exactly one, and the opcode tells you which.
| Format | [31:25] |
[24:20] |
[19:15] |
[14:12] |
[11:7] |
[6:0] |
|---|---|---|---|---|---|---|
| R-type | funct7 | rs2 | rs1 | funct3 | rd | opcode |
| I-type | imm[11:0] (spans [31:20]) | rs1 | funct3 | rd | opcode | |
| S-type | imm[11:5] | rs2 | rs1 | funct3 | imm[4:0] | opcode |
| B-type | imm[12,10:5] | rs2 | rs1 | funct3 | imm[4:1,11] | opcode |
| U-type | imm[31:12] (spans [31:12]) | rd | opcode | |||
| J-type | imm[20,10:1,11,19:12] (spans [31:12]) | rd | opcode |
The immediate columns are where the formats differ in difficulty. I-type and U-type immediates are contiguous slices — easy. S-type splits its 12-bit immediate into two pieces ([11:5] and [4:0]) that straddle the rs2/rs1 fields. B-type and J-type scatter their immediate bits in a pattern that keeps the sign bit at bit 31 (so sign extension is uniform across all formats) and keeps rs1/rs2/rd fixed. The bit scrambling is not arbitrary cruelty; it is the cost of the fixed-field invariant.
Fields, opcodes, and the control bundle
Before the switch, the decoder slices the fixed fields out of the instruction word. These are pure bit operations identical to the binary-encoding slice in the toy example, just at known offsets:
| Field | Bits | Extraction |
|---|---|---|
| opcode | [6:0] |
raw & 0x7F |
| funct3 | [14:12] |
(raw >> 12) & 0x7 |
| funct7 | [31:25] |
(raw >> 25) & 0x7F |
| rs1 | [19:15] |
(raw >> 15) & 0x1F |
| rs2 | [24:20] |
(raw >> 20) & 0x1F |
| rd | [11:7] |
(raw >> 7) & 0x1F |
The opcode is the case selector — the RV32I equivalent of our toy 3-bit code, just 7 bits wide with ten defined values instead of six. The control bundle the decoder drives for each instruction group is:
| Group | alu_op |
alu_src |
mem_read |
mem_write |
reg_write |
branch |
jump |
wb_sel |
|---|---|---|---|---|---|---|---|---|
| R-type (ADD/SUB/…) | varies | REG | 0 | 0 | 1 | 0 | 0 | ALU |
| I-type ALU (ADDI/…) | varies | IMM | 0 | 0 | 1 | 0 | 0 | ALU |
| LOAD (LW/LH/LB/…) | ADD | IMM | 1 | 0 | 1 | 0 | 0 | MEM |
| STORE (SW/SH/SB) | ADD | IMM | 0 | 1 | 0 | 0 | 0 | ALU |
| BRANCH (BEQ/BNE/…) | SUB | REG | 0 | 0 | 0 | 1 | 0 | ALU |
| JAL | ADD | IMM | 0 | 0 | 1 | 0 | 1 | PC4 |
| JALR | ADD | IMM | 0 | 0 | 1 | 0 | 1 | PC4 |
| LUI | LUI | IMM | 0 | 0 | 1 | 0 | 0 | ALU |
| AUIPC | AUIPC | IMM | 0 | 0 | 1 | 0 | 0 | ALU |
| illegal | ADD | REG | 0 | 0 | 0 | 0 | 0 | ALU + illegal=1 |
alu_src: REG means "use rs2," IMM means "use the sign-extended immediate." wb_sel: ALU writes the ALU result, MEM writes loaded data, PC4 writes the return address PC+4. Read this table as the RV32I version of the toy machine's four-column table — same idea, wider bundle. The illegal row, exactly as before, is an explicit all-safe NOP with an illegal flag raised.
%%{init: {'theme':'base','themeVariables':{'primaryColor':'#dbeafe','primaryTextColor':'#1e293b','primaryBorderColor':'#3b82f6','lineColor':'#64748b','secondaryColor':'#f1f5f9'}}}%%
flowchart LR
INSTR["instr[31:0]"] --> FIELDS["slice fields\nopcode/funct3/funct7\nrs1/rs2/rd"]
FIELDS --> CTRL["switch(opcode)\ncontrol decode"]
FIELDS --> IMM["immediate gen\nper-format + sign-extend"]
CTRL --> BUNDLE["control bundle\nalu_op, alu_src, mem_*,\nreg_write, branch, jump, wb_sel"]
IMM --> IMMOUT["imm[31:0]"]
CTRL --> ILL["illegal"]
The RV32I decoder, complete
This is one self-contained program: a header-free single file with the decoder module, a small driver, a monitor, and sc_main. It is the toy decoder's structure scaled to RV32I.
// file: rv32i_decoder.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// rv32i_decoder.cpp -o rv32i_decoder -lsystemc
// Run: LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./rv32i_decoder
#include <systemc.h>
#include <cstdint>
#include <iostream>
// ---- ALU operation encoding (4 bits) ----
enum class AluOp : uint8_t {
ADD = 0, SUB = 1, AND = 2, OR = 3, XOR = 4,
SLL = 5, SRL = 6, SRA = 7, SLT = 8, SLTU = 9,
LUI = 10, AUIPC = 11
};
// ---- Write-back source select (2 bits) ----
enum class WbSel : uint8_t { ALU = 0, MEM = 1, PC4 = 2 };
// ---- Opcode constants (instruction bits [6:0]) ----
static constexpr unsigned OP_R = 0x33;
static constexpr unsigned OP_I_ALU = 0x13;
static constexpr unsigned OP_LOAD = 0x03;
static constexpr unsigned OP_STORE = 0x23;
static constexpr unsigned OP_BRANCH = 0x63;
static constexpr unsigned OP_JAL = 0x6F;
static constexpr unsigned OP_JALR = 0x67;
static constexpr unsigned OP_LUI = 0x37;
static constexpr unsigned OP_AUIPC = 0x17;
// ---- funct3 codes shared by R-type and I-type ALU ----
static constexpr unsigned F3_ADD_SUB = 0x0;
static constexpr unsigned F3_SLL = 0x1;
static constexpr unsigned F3_SLT = 0x2;
static constexpr unsigned F3_SLTU = 0x3;
static constexpr unsigned F3_XOR = 0x4;
static constexpr unsigned F3_SR = 0x5; // SRL vs SRA via funct7 bit 5
static constexpr unsigned F3_OR = 0x6;
static constexpr unsigned F3_AND = 0x7;
SC_MODULE(Rv32iDecoder) {
sc_in<sc_uint<32>> instr;
sc_out<sc_uint<5>> rs1_addr;
sc_out<sc_uint<5>> rs2_addr;
sc_out<sc_uint<5>> rd_addr;
sc_out<sc_uint<3>> funct3_out;
sc_out<sc_uint<32>> imm;
sc_out<sc_uint<4>> alu_op;
sc_out<bool> alu_src; // false = rs2, true = immediate
sc_out<bool> mem_read;
sc_out<bool> mem_write;
sc_out<bool> reg_write;
sc_out<bool> branch;
sc_out<bool> jump;
sc_out<sc_uint<2>> wb_sel;
sc_out<bool> illegal;
void decode() {
uint32_t raw = instr.read();
// --- Field slicing (fixed positions, every format) ---
unsigned opcode = raw & 0x7F;
unsigned f3 = (raw >> 12) & 0x7;
unsigned f7 = (raw >> 25) & 0x7F;
unsigned rs1 = (raw >> 15) & 0x1F;
unsigned rs2 = (raw >> 20) & 0x1F;
unsigned rd = (raw >> 7) & 0x1F;
// Register addresses and funct3 are pure pass-throughs: drive always.
rs1_addr.write(rs1);
rs2_addr.write(rs2);
rd_addr.write(rd);
funct3_out.write(f3);
// --- SAFE DEFAULTS for the whole control bundle ---
AluOp s_alu = AluOp::ADD;
bool s_src = false;
bool s_mr = false;
bool s_mw = false;
bool s_rw = false;
bool s_br = false;
bool s_jmp = false;
WbSel s_wb = WbSel::ALU;
uint32_t s_imm = 0;
bool s_ill = false;
// --- Immediate generators (one per format) ---
// I-type: sign-extend bits [31:20] (12 bits).
auto imm_i = [&]() -> uint32_t {
return (uint32_t)(int32_t)((int32_t)raw >> 20);
};
// S-type: imm[11:5] = raw[31:25], imm[4:0] = raw[11:7].
auto imm_s = [&]() -> uint32_t {
uint32_t v = (((raw >> 25) & 0x7F) << 5) | ((raw >> 7) & 0x1F);
uint32_t sign = (raw >> 31) & 0x1;
if (sign) v |= 0xFFFFF000u; // sign-extend bit 11
return v;
};
// B-type: imm[12]=raw[31], imm[11]=raw[7], imm[10:5]=raw[30:25],
// imm[4:1]=raw[11:8], imm[0]=0 (2-byte aligned).
auto imm_b = [&]() -> uint32_t {
uint32_t v = (((raw >> 31) & 0x1) << 12)
| (((raw >> 7) & 0x1) << 11)
| (((raw >> 25) & 0x3F) << 5)
| (((raw >> 8) & 0xF) << 1);
uint32_t sign = (raw >> 31) & 0x1;
if (sign) v |= 0xFFFFE000u; // sign-extend bit 12
return v;
};
// U-type: imm[31:12] = raw[31:12], low 12 bits zero, no sign-extend.
auto imm_u = [&]() -> uint32_t {
return raw & 0xFFFFF000u;
};
// J-type: imm[20]=raw[31], imm[10:1]=raw[30:21], imm[11]=raw[20],
// imm[19:12]=raw[19:12], imm[0]=0.
auto imm_j = [&]() -> uint32_t {
uint32_t v = (((raw >> 31) & 0x1) << 20)
| (((raw >> 21) & 0x3FF) << 1)
| (((raw >> 20) & 0x1) << 11)
| (((raw >> 12) & 0xFF) << 12);
uint32_t sign = (raw >> 31) & 0x1;
if (sign) v |= 0xFFE00000u; // sign-extend bit 20
return v;
};
// --- ALU-op sub-decode for R-type (funct3 + funct7 bit 5) ---
auto r_alu = [&]() -> AluOp {
switch (f3) {
case F3_ADD_SUB: return (f7 & 0x20) ? AluOp::SUB : AluOp::ADD;
case F3_SLL: return AluOp::SLL;
case F3_SLT: return AluOp::SLT;
case F3_SLTU: return AluOp::SLTU;
case F3_XOR: return AluOp::XOR;
case F3_SR: return (f7 & 0x20) ? AluOp::SRA : AluOp::SRL;
case F3_OR: return AluOp::OR;
default: return AluOp::AND; // F3_AND
}
};
// --- ALU-op sub-decode for I-type ALU (SRAI vs SRLI via bit 30) ---
auto i_alu = [&]() -> AluOp {
switch (f3) {
case F3_ADD_SUB: return AluOp::ADD; // ADDI
case F3_SLL: return AluOp::SLL; // SLLI
case F3_SLT: return AluOp::SLT; // SLTI
case F3_SLTU: return AluOp::SLTU; // SLTIU
case F3_XOR: return AluOp::XOR; // XORI
case F3_SR: return (raw & (1u << 30)) ? AluOp::SRA : AluOp::SRL;
case F3_OR: return AluOp::OR; // ORI
default: return AluOp::AND; // ANDI
}
};
// --- Main case analysis on the opcode ---
switch (opcode) {
case OP_R:
s_alu = r_alu(); s_src = false; s_rw = true; s_wb = WbSel::ALU;
break;
case OP_I_ALU:
s_alu = i_alu(); s_src = true; s_rw = true; s_wb = WbSel::ALU;
s_imm = imm_i();
break;
case OP_LOAD:
s_alu = AluOp::ADD; s_src = true; s_mr = true; s_rw = true;
s_wb = WbSel::MEM; s_imm = imm_i();
break;
case OP_STORE:
s_alu = AluOp::ADD; s_src = true; s_mw = true; s_wb = WbSel::ALU;
s_imm = imm_s();
break;
case OP_BRANCH:
s_alu = AluOp::SUB; s_src = false; s_br = true; s_wb = WbSel::ALU;
s_imm = imm_b();
break;
case OP_JAL:
s_alu = AluOp::ADD; s_src = true; s_jmp = true; s_rw = true;
s_wb = WbSel::PC4; s_imm = imm_j();
break;
case OP_JALR:
s_alu = AluOp::ADD; s_src = true; s_jmp = true; s_rw = true;
s_wb = WbSel::PC4; s_imm = imm_i();
break;
case OP_LUI:
s_alu = AluOp::LUI; s_src = true; s_rw = true; s_wb = WbSel::ALU;
s_imm = imm_u();
break;
case OP_AUIPC:
s_alu = AluOp::AUIPC; s_src = true; s_rw = true; s_wb = WbSel::ALU;
s_imm = imm_u();
break;
default:
// Illegal opcode: defaults already NOP everything; flag it.
s_ill = true;
break;
}
// --- Drive the full bundle (every output, every activation) ---
alu_op.write((uint8_t)s_alu);
alu_src.write(s_src);
mem_read.write(s_mr);
mem_write.write(s_mw);
reg_write.write(s_rw);
branch.write(s_br);
jump.write(s_jmp);
wb_sel.write((uint8_t)s_wb);
imm.write(s_imm);
illegal.write(s_ill);
}
SC_CTOR(Rv32iDecoder) {
SC_METHOD(decode);
sensitive << instr; // single input; single sensitivity entry
}
};
// Helper encoders so the testbench reads as assembly, not hex.
static uint32_t enc_r(unsigned f7, unsigned rs2, unsigned rs1,
unsigned f3, unsigned rd, unsigned op) {
return (f7 << 25) | (rs2 << 20) | (rs1 << 15) | (f3 << 12) | (rd << 7) | op;
}
static uint32_t enc_i(int32_t imm, unsigned rs1, unsigned f3,
unsigned rd, unsigned op) {
return (((uint32_t)imm & 0xFFF) << 20) | (rs1 << 15) | (f3 << 12)
| (rd << 7) | op;
}
static uint32_t enc_s(int32_t imm, unsigned rs2, unsigned rs1,
unsigned f3, unsigned op) {
uint32_t up = (((uint32_t)imm >> 5) & 0x7F) << 25;
uint32_t lo = ((uint32_t)imm & 0x1F) << 7;
return up | (rs2 << 20) | (rs1 << 15) | (f3 << 12) | lo | op;
}
static uint32_t enc_b(int32_t imm, unsigned rs2, unsigned rs1,
unsigned f3, unsigned op) {
uint32_t b12 = (((uint32_t)imm >> 12) & 0x1) << 31;
uint32_t b11 = (((uint32_t)imm >> 11) & 0x1) << 7;
uint32_t b105 = (((uint32_t)imm >> 5) & 0x3F) << 25;
uint32_t b41 = (((uint32_t)imm >> 1) & 0xF) << 8;
return b12 | b105 | (rs2 << 20) | (rs1 << 15) | (f3 << 12) | b41 | b11 | op;
}
static uint32_t enc_u(uint32_t imm, unsigned rd, unsigned op) {
return (imm & 0xFFFFF000u) | (rd << 7) | op;
}
static uint32_t enc_j(int32_t imm, unsigned rd, unsigned op) {
uint32_t b20 = (((uint32_t)imm >> 20) & 0x1) << 31;
uint32_t b101 = (((uint32_t)imm >> 1) & 0x3FF) << 21;
uint32_t b11 = (((uint32_t)imm >> 11) & 0x1) << 20;
uint32_t b1912 = (((uint32_t)imm >> 12) & 0xFF) << 12;
return b20 | b101 | b11 | b1912 | (rd << 7) | op;
}
SC_MODULE(Driver) {
sc_out<sc_uint<32>> instr;
void run() {
// A short instruction stream that hits several formats + an illegal word.
instr.write(enc_r(0x00, 2, 1, F3_ADD_SUB, 3, OP_R)); wait(1, SC_NS); // ADD
instr.write(enc_r(0x20, 2, 1, F3_ADD_SUB, 3, OP_R)); wait(1, SC_NS); // SUB
instr.write(enc_i(-4, 1, 0x0, 3, OP_LOAD)); wait(1, SC_NS); // LB
instr.write(enc_s(-8, 2, 1, 0x2, OP_STORE)); wait(1, SC_NS); // SW
instr.write(enc_b(16, 2, 1, 0x0, OP_BRANCH)); wait(1, SC_NS); // BEQ
instr.write(enc_j(1024, 3, OP_JAL)); wait(1, SC_NS); // JAL
instr.write(enc_u(0x12345000u, 3, OP_LUI)); wait(1, SC_NS); // LUI
instr.write(0xFFFFFFFFu); wait(1, SC_NS); // illegal
sc_stop();
}
SC_CTOR(Driver) { SC_THREAD(run); }
};
SC_MODULE(Mon) {
sc_in<sc_uint<32>> instr;
sc_in<sc_uint<4>> alu_op;
sc_in<bool> mem_read, mem_write, reg_write, branch, jump, illegal;
sc_in<sc_uint<32>> imm;
void watch() {
std::cout << "[" << sc_time_stamp() << "] "
<< "alu_op=" << alu_op.read()
<< " mr=" << mem_read.read()
<< " mw=" << mem_write.read()
<< " rw=" << reg_write.read()
<< " br=" << branch.read()
<< " jmp=" << jump.read()
<< " imm=" << (int32_t)(uint32_t)imm.read()
<< " ill=" << illegal.read() << "\n";
}
SC_CTOR(Mon) {
SC_METHOD(watch);
sensitive << alu_op << mem_read << mem_write << reg_write
<< branch << jump << imm << illegal;
dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_signal<sc_uint<32>> s_instr, s_imm;
sc_signal<sc_uint<5>> s_rs1, s_rs2, s_rd;
sc_signal<sc_uint<3>> s_f3;
sc_signal<sc_uint<4>> s_alu;
sc_signal<bool> s_src, s_mr, s_mw, s_rw, s_br, s_jmp, s_ill;
sc_signal<sc_uint<2>> s_wb;
Rv32iDecoder dec("dec");
dec.instr(s_instr);
dec.rs1_addr(s_rs1); dec.rs2_addr(s_rs2); dec.rd_addr(s_rd);
dec.funct3_out(s_f3); dec.imm(s_imm);
dec.alu_op(s_alu); dec.alu_src(s_src);
dec.mem_read(s_mr); dec.mem_write(s_mw); dec.reg_write(s_rw);
dec.branch(s_br); dec.jump(s_jmp); dec.wb_sel(s_wb); dec.illegal(s_ill);
Driver drv("drv");
drv.instr(s_instr);
Mon mon("mon");
mon.instr(s_instr);
mon.alu_op(s_alu); mon.mem_read(s_mr); mon.mem_write(s_mw);
mon.reg_write(s_rw); mon.branch(s_br); mon.jump(s_jmp);
mon.imm(s_imm); mon.illegal(s_ill);
sc_start();
return 0;
}
Expected output:
[0 s] alu_op=0 mr=0 mw=0 rw=1 br=0 jmp=0 imm=0 ill=0
[1 ns] alu_op=1 mr=0 mw=0 rw=1 br=0 jmp=0 imm=0 ill=0
[2 ns] alu_op=0 mr=1 mw=0 rw=1 br=0 jmp=0 imm=-4 ill=0
[3 ns] alu_op=0 mr=0 mw=1 rw=0 br=0 jmp=0 imm=-8 ill=0
[4 ns] alu_op=1 mr=0 mw=0 rw=0 br=1 jmp=0 imm=16 ill=0
[5 ns] alu_op=0 mr=0 mw=0 rw=1 br=0 jmp=1 imm=1024 ill=0
[6 ns] alu_op=10 mr=0 mw=0 rw=1 br=0 jmp=0 imm=305418240 ill=0
[7 ns] alu_op=0 mr=0 mw=0 rw=0 br=0 jmp=0 imm=0 ill=1
Walk the trace against the control table. At t=0, ADD: alu_op=0 (ADD), reg_write=1, everything else NOP, immediate 0 (R-type has none). At t=1 ns, SUB: identical bundle except alu_op=1 (SUB) — the funct7 bit 5 flipped the ALU op while every control signal stayed the same, exactly the silent ADD-vs-SUB distinction we will dwell on in the Advanced section. At t=2 ns, LB (a load with offset −4): mem_read=1, reg_write=1, imm=-4 correctly sign-extended. At t=3 ns, SW with offset −8: mem_write=1, reg_write=0 (stores do not write a register), imm=-8 reassembled from the split S-type fields and sign-extended. At t=4 ns, BEQ +16: branch=1, imm=16 from the scrambled B-type encoding. At t=5 ns, JAL +1024: jump=1, reg_write=1 (the link register), imm=1024 from the J-type scramble. At t=6 ns, LUI with upper immediate 0x12345: alu_op=10 (LUI), imm=0x12345000 which prints as the decimal 305418240. At t=7 ns, the all-ones illegal word: every control output NOP and illegal=1 — the default arm, exactly as in the toy decoder.
Every output is driven on every line. That is the safe-defaults discipline paying off across a 10-output bundle: the illegal instruction at t=7 ns does not inherit LUI's alu_op=10 from the previous cycle, because the defaults reset s_alu to ADD (0) before the switch ran and the default arm overrode nothing.
The immediate generator is just more case analysis
Notice that immediate extraction is itself a small decoder. The opcode selects a format, and the format selects which immediate-generation lambda runs. I-type and U-type are contiguous slices; S/B/J reassemble scattered bits and then sign-extend at the width of the format's immediate (12 bits for I/S, 13 for B, 20 for U, 21 for J). The sign-extension step is the one most often gotten wrong: you must replicate the format's sign bit, not bit 31 of some intermediate value, across the high bits. The code does this explicitly with the if (sign) v |= mask pattern, where the mask covers exactly the bits above the immediate's width. This is more transparent than a chained sc_int<N> cast and makes the "extend from this bit" decision visible to a reviewer.
This closes the Intermediate section. You have built a complete, correct RV32I decoder as a direct scaling of the toy decoder: slice fields, set safe defaults, switch on the opcode, a default for illegal, and a per-format immediate sub-decoder. The Advanced section drills the corners where decoders bite.
Advanced: Edge Cases & LRM Corners
The Beginner and Intermediate sections cover what a correct decoder does. This section is the 5% that takes a senior engineer half a day to track down: the silent confusions where two instructions differ in one bit, the precise SystemC mechanism behind the stale-output latch, the deliberate handling of illegal instructions, and the verification mindset that catches decoder bugs before they ship.
Corner 1: the silent one-bit confusions (SRAI vs SRLI, SUB vs ADD)
Decoder bugs do not crash. A wrong control signal produces a wrong result, and a wrong result looks exactly like a correct one until you compare it against a golden model. The most dangerous decoder bugs are the pairs of instructions that share an opcode and a funct3 and differ in a single bit elsewhere — because a test that exercises one of the pair and not the other will pass with a decoder that handles only one.
Two such pairs live in RV32I. ADD vs SUB both use opcode 0x33 and funct3 0x0; they differ only in funct7 bit 5 (instruction bit 30). SRLI vs SRAI both use opcode 0x13 and funct3 0x5; they differ only in instruction bit 30 (which lives inside the I-type immediate field, where it functions as imm[10]). The r_alu and i_alu lambdas in our decoder handle these with (f7 & 0x20) and (raw & (1u << 30)) respectively. Get the mask wrong — 0x40 instead of 0x20, or bit 29 instead of 30 — and the decoder silently maps SUB to ADD or SRAI to SRLI.
// The two single-bit distinctions, isolated:
// ADD vs SUB: opcode 0x33, funct3 0x0, differ in funct7 bit 5.
AluOp add_or_sub = (f7 & 0x20) ? AluOp::SUB : AluOp::ADD;
// SRLI vs SRAI: opcode 0x13, funct3 0x5, differ in instruction bit 30.
AluOp srl_or_sra = (raw & (1u << 30)) ? AluOp::SRA : AluOp::SRL;
The verification consequence is concrete: a decode test that includes ADD x3, x1, x2 but not SUB x3, x1, x2, or SRLI x1, x1, 4 but not SRAI x1, x1, 4, cannot detect a decoder that collapses the pair. You must encode both members of each pair, as adjacent test vectors, and assert that alu_op differs between them. This is not optional thoroughness; it is the only way the bit is exercised at all.
Corner 2: the stale-output latch, in SystemC kernel terms
We met the stale-output bug in the Beginner section. Here is precisely why it happens in terms of the sc_signal machinery from Part 1, because understanding the mechanism is what lets you recognize the bug from a waveform rather than from the source.
An sc_out<bool> is bound to an sc_signal<bool>. Per IEEE 1666-2011 §6.4, calling write() on that signal during the evaluate phase does not change the signal's value immediately; it requests an update. In the following update phase the kernel commits the requested value and, if it differs from the current value, fires the value-changed event. The crucial corollary: if write() is never called in an activation, no update is requested, and the signal retains its last committed value. There is no "reset to default between activations." The signal is a persistent storage cell from the kernel's point of view.
So when a decoder activation runs a switch arm that forgets to write mem_read, the kernel simply receives no update request for mem_read's signal, and mem_read holds the value the previous activation committed. Across two different inputs that select two different arms, one of which writes mem_read and one of which does not, mem_read carries state from the first input into the second. That carried state is exactly a transparent latch: output follows input when the arm writes it, holds when the arm does not. The defaults-before-switch pattern fixes it by guaranteeing that every activation issues a write request for every output, so no output can ever carry a value across activations.
// Why this is a latch in kernel terms:
// Activation N (input selects an arm that writes mem_read):
// mem_read.write(true) -> update requested -> commits true
// Activation N+1 (input selects an arm that does NOT write mem_read):
// (no write) -> no update requested -> mem_read STAYS true
// The output "remembers" — that memory is the inferred latch.
This is also why a pure-combinational decoder must not call dont_initialize(). Without the initial activation, the outputs never receive their first write request and sit at their default-constructed values (false/0) until the first input change — a different but related "stuck at initialization" defect.
Corner 3: illegal-instruction handling as a first-class output
In the toy decoder and the RV32I decoder, the default arm raises an illegal flag. This is not decoration; in a real CPU the illegal-instruction signal drives an exception, and getting it wrong has two failure modes that are worth naming.
The first failure is silently executing an illegal instruction. If you omit the illegal output and the default arm only NOPs the control bundle, an illegal opcode produces a harmless-looking all-zero bundle — reg_write=0, no memory access — and the CPU quietly skips it instead of trapping. A program that jumps into data, or a fuzzer that feeds random words, will "execute" garbage as NOPs and the test will not flag it. The illegal output is what turns a silent skip into an observable event.
The second failure is a too-narrow legal set. Our decoder flags any opcode outside its nine cases as illegal, but RV32I legality is finer-grained than opcode alone: an opcode-0x33 word with a funct3/funct7 combination that no R-type instruction defines is also illegal, even though the opcode is "legal." A fully conformant decoder checks funct3 and funct7 within each opcode case and raises illegal for undefined sub-encodings. Our teaching decoder maps undefined R-type sub-codes onto defined ALU ops (the default arms in r_alu/i_alu) rather than flagging them, which is fine for a single-cycle teaching core but would be a conformance gap in a production decoder. The principle: illegal is a function of the whole instruction word, not just the opcode field, and a rigorous decoder treats every undefined bit-pattern in every field as a path to illegal.
// Sketch of finer-grained legality inside an opcode case:
case OP_R: {
bool legal_r =
(f3 == F3_ADD_SUB && (f7 == 0x00 || f7 == 0x20)) ||
(f3 == F3_SR && (f7 == 0x00 || f7 == 0x20)) ||
(f3 != F3_ADD_SUB && f3 != F3_SR && f7 == 0x00);
if (!legal_r) { s_ill = true; break; } // undefined R sub-encoding
s_alu = r_alu(); s_src = false; s_rw = true; s_wb = WbSel::ALU;
break;
}
Corner 4: the decode-table-driven verification mindset
The right way to verify a decoder is the way the decoder was specified: as a table pairing input patterns with expected output bundles. This is the single most important verification idea in the post, and it generalizes far past RV32I — any decoder, in any protocol, is verified by a golden table.
The structure is a list of {instruction, expected-bundle} rows defined independently of the DUT. You drive each instruction, read back every output, and compare against the expected bundle. The independence is what makes it a real check: if you derived the expected values by running the DUT, you would only be testing that the DUT agrees with itself.
// A golden-table verification harness for the decoder above.
// (Same #includes, enums, opcode constants, encoders, and Rv32iDecoder
// module as rv32i_decoder.cpp; only sc_main is replaced.)
struct Expect {
uint32_t instr;
const char* name;
AluOp alu_op;
bool mem_read, mem_write, reg_write, branch, jump, illegal;
int32_t imm; // INT32_MIN means "don't check imm"
};
int sc_main(int, char*[]) {
sc_signal<sc_uint<32>> s_instr, s_imm;
sc_signal<sc_uint<5>> s_rs1, s_rs2, s_rd;
sc_signal<sc_uint<3>> s_f3;
sc_signal<sc_uint<4>> s_alu;
sc_signal<bool> s_src, s_mr, s_mw, s_rw, s_br, s_jmp, s_ill;
sc_signal<sc_uint<2>> s_wb;
Rv32iDecoder dec("dec");
dec.instr(s_instr);
dec.rs1_addr(s_rs1); dec.rs2_addr(s_rs2); dec.rd_addr(s_rd);
dec.funct3_out(s_f3); dec.imm(s_imm);
dec.alu_op(s_alu); dec.alu_src(s_src);
dec.mem_read(s_mr); dec.mem_write(s_mw); dec.reg_write(s_rw);
dec.branch(s_br); dec.jump(s_jmp); dec.wb_sel(s_wb); dec.illegal(s_ill);
const int32_t X = INT32_MIN;
std::vector<Expect> table = {
// The ADD/SUB pair — adjacent, must differ in alu_op.
{ enc_r(0x00,2,1,F3_ADD_SUB,3,OP_R), "ADD", AluOp::ADD, 0,0,1,0,0,0, X },
{ enc_r(0x20,2,1,F3_ADD_SUB,3,OP_R), "SUB", AluOp::SUB, 0,0,1,0,0,0, X },
// The SRLI/SRAI pair — adjacent, must differ in alu_op.
{ enc_i(4,1,F3_SR,1,OP_I_ALU), "SRLI",AluOp::SRL, 0,0,1,0,0,0, 4 },
{ (uint32_t)0x40405093u, "SRAI",AluOp::SRA, 0,0,1,0,0,0, X },
{ enc_i(-4,1,0x0,3,OP_LOAD), "LB", AluOp::ADD, 1,0,1,0,0,0, -4 },
{ enc_s(-8,2,1,0x2,OP_STORE), "SW", AluOp::ADD, 0,1,0,0,0,0, -8 },
{ enc_b(16,2,1,0x0,OP_BRANCH), "BEQ", AluOp::SUB, 0,0,0,1,0,0, 16 },
{ enc_j(1024,3,OP_JAL), "JAL", AluOp::ADD, 0,0,1,0,1,0, 1024 },
{ 0xFFFFFFFFu, "ILL", AluOp::ADD, 0,0,0,0,0,1, 0 },
};
int pass = 0, fail = 0;
for (const auto& e : table) {
s_instr.write(e.instr);
sc_start(1, SC_NS); // let the combinational decode settle
bool ok = true;
auto bad = [&](const char* w){ std::cout << " FAIL " << e.name
<< ": " << w << "\n"; ok = false; };
if ((uint8_t)s_alu.read() != (uint8_t)e.alu_op) bad("alu_op");
if (s_mr.read() != e.mem_read) bad("mem_read");
if (s_mw.read() != e.mem_write) bad("mem_write");
if (s_rw.read() != e.reg_write) bad("reg_write");
if (s_br.read() != e.branch) bad("branch");
if (s_jmp.read() != e.jump) bad("jump");
if (s_ill.read() != e.illegal) bad("illegal");
if (e.imm != X && (int32_t)(uint32_t)s_imm.read() != e.imm) bad("imm");
if (ok) { std::cout << " PASS " << e.name << "\n"; ++pass; }
else ++fail;
}
std::cout << "PASS=" << pass << " FAIL=" << fail << "\n";
std::cout << (fail ? "RESULT: FAIL\n" : "RESULT: PASS\n");
return fail ? 1 : 0;
}
Expected output:
PASS ADD
PASS SUB
PASS SRLI
PASS SRAI
PASS LB
PASS SW
PASS BEQ
PASS JAL
PASS ILL
PASS=9 FAIL=0
RESULT: PASS
The table deliberately places ADD next to SUB and SRLI next to SRAI, with differing expected alu_op values, so the single-bit distinctions are checked. It includes a store (to confirm reg_write=0), a branch and a jump (to confirm the right control bit), and an illegal word (to confirm the default/illegal path). This is the same golden-reference principle as reference-model-driven verification in UVM: the expected values are committed before, and independently of, the implementation. A real conformance suite would extend this table to all base RV32I instructions plus negative-immediate and boundary cases for each format's sign extension — but the structure is exactly this, and the structure is the lesson.
Hands-on exercise
Take the RV32I decoder from the Intermediate section and extend it with a feature real cores need: a is_compressed pre-check that distinguishes 16-bit compressed instructions from 32-bit standard ones. In RV32IC, an instruction is 16-bit (compressed) when its lowest two bits are not both 1, and 32-bit (standard) when bits [1:0] == 0b11.
- Add an
sc_out<bool> is_compressedport. Indecode(), set it true when(raw & 0x3) != 0x3. This is a one-line decoder in front of the opcode decoder — case analysis on a 2-bit field. - When
is_compressedis true, your existing opcodeswitchis decoding garbage (the upper bits belong to the next instruction). Guard the standard decode: if compressed, set the control bundle to NOP, leaveillegallow (it is a legal compressed instruction, you simply are not expanding it yet), and setis_compressedhigh. - Add a golden-table row for a compressed word — for example
0x0001isC.NOP— and assertis_compressed=1,illegal=0, and a NOP control bundle. - Stretch: add
illegaldetection for the reserved all-zero 16-bit pattern0x0000, which RV32IC defines as illegal. Nowis_compressedandillegalcan both be high — verify your defaults-and-override structure handles that combination without one clobbering the other.
The point of the exercise is that the compressed pre-check is another decoder stacked in front of the opcode decoder, and the discipline is identical: slice the selecting field, set defaults, case-analyze, drive every output. If your is_compressed logic ever leaves the standard control bundle holding stale values from a previous 32-bit instruction, you have reproduced the stale-output bug — and the fix is the same defaults-before-logic structure you have used all along.
Common mistakes
- Omitting the safe-defaults block. Writing each output only inside the case that needs it leaves every other case holding stale values — an inferred latch in synthesis, silent wrong outputs in SystemC. Always assign every output before the
switch. - No
defaultarm. Withoutdefault, an opcode that matches no case writes nothing and every output holds its previous value. Always include adefaultthat NOPs the bundle and raisesillegal. - Reflexively adding
dont_initialize()to a combinational decoder. That suppresses the initial activation, so outputs sit at default-constructed values until the first input change. A pure-combinationalSC_METHODshould not calldont_initialize(). - Incomplete sensitivity list. SystemC has no
always @(*). If you read a signal indecode(), it must appear insensitive << .... A missing input means the decoder does not re-run when that input changes — stale output, no warning. - Sign-extending from the wrong bit. Each format's immediate has its own width and its own sign bit (bit 11 for I/S, bit 12 for B, bit 20 for J). Replicate the format's sign bit, not bit 31 of an intermediate, across the high bits.
- Reading the I-type immediate slice for an S-type store. S-type splits its immediate around the rs2 field; reading bits
[31:20]gives the wrong value. Reassemble[11:5]and[4:0]separately, then sign-extend. - Testing only one of a single-bit pair. ADD without SUB, or SRLI without SRAI, leaves the distinguishing bit unexercised. Always test both members of each pair adjacently and assert they differ.
- Treating
illegalas opcode-only. Legality depends on the whole word; an undefined funct3/funct7 within a legal opcode is still illegal. A rigorous decoder checks sub-fields, not just the opcode.
Recap
After working through this post you can now:
- Define a decoder as pure combinational case analysis — read an input code, select exactly one of a fixed set of mutually exclusive cases, and drive a complete output bundle — and implement it as a single
SC_METHODsensitive to every input it reads, with nodont_initialize(). - State the three correctness properties of any decoder: mutual exclusivity of cases, completeness of selection (every code lands somewhere, including a
default), and completeness of output (every output driven on every path). - Apply the defaults-before-
switchpattern to guarantee output-completeness automatically, and explain why omitting it produces the stale-output / inferred-latch bug. - Explain that inferred-latch bug in kernel terms — an output whose
sc_signalreceived nowrite()request in an activation holds its last committed value — and recognize it from a waveform. - Choose between binary and one-hot encodings for a selection job by reasoning about wire count, per-output logic depth, and illegal-code detection.
- Build a complete RV32I instruction decoder from scratch: slice the fixed opcode/funct3/funct7/rs1/rs2/rd fields, set safe defaults,
switchon the 7-bit opcode, sub-decode the ALU op from funct3/funct7, and route everything outside the legal set to adefaultthat raisesillegal. - Extract and sign-extend immediates for all six RV32I formats, reassembling the scattered S/B/J bits and replicating the format's sign bit at the correct width.
- Catch the silent one-bit confusions (ADD vs SUB, SRLI vs SRAI) and illegal-instruction defects with a decode-table-driven harness that pairs each input pattern with an independently specified expected bundle.
Further reading
- IEEE Std 1666-2011, IEEE Standard for Standard SystemC Language Reference Manual, §5.2.16 (
SC_METHOD), §5.2.18 (sensitivity anddont_initialize), §6.4 (sc_signalrequest-update semantics). The normative source for why a combinational decoder behaves as it does. - The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA, v20191213, §2.2–2.3 — instruction formats and the immediate-encoding variants that explain B/J bit scrambling.
- D. Bhasker, A SystemC Primer, 2nd ed., chapters 4–5 — combinational
SC_METHODmodeling and the defaults-then-case idiom. - T. Grötker, S. Liao, G. Martin, S. Swan, System Design with SystemC, chapter 3 — combinational processes and signal mechanics.
- Berkeley CS152 lecture notes on single-cycle RV32I — the control-signal table this post's worked example follows.
- The Ibex
ibex_decoder.svand PicoRV32 sources — production RV32I decoders that use the same defaults-then-case structure and explicit illegal-instruction flagging.
Next in this section
→ Part 3: Registered & Sequential Logic — Program Counter & sequential modeling. The decoder you just built is stateless; the next part adds the storage element. We start from a generic 4-bit counter to establish clocked-process and reset semantics from first principles, then apply them to the RV32I program counter and instruction fetch — the module that feeds this decoder one instruction per cycle.
Comments (0)
Leave a Comment