12. SystemC Tutorial - Memories & Register Files
Why this matters
Rewritten 2026-06-05 with deeper first-principles material.
Storage is where every interesting design keeps its state. The register file inside a CPU holds the operands the next instruction will consume. The line buffers inside a cache hold the bytes a load will return. The scoreboard RAM inside a DRAM controller holds the open-row state of every bank. The descriptor ring inside a DMA engine, the credit counters inside a PCIe link layer, the FIFO that decouples a fast producer from a slow consumer — all of them are memories. When you model hardware in SystemC, you spend more lines describing how storage is read and written than almost anything else, and you get more subtle bugs there than almost anywhere else. A memory whose read returns the value from the wrong cycle will pass a hundred directed tests and then fail the one that writes and reads the same address back-to-back. A memory built from an array of sc_signal will simulate correctly and run ten times too slow, and nobody will know why until a profiler points at the channel update phase. A register file whose x0 is not properly hardwired to zero will boot Linux and then mysteriously corrupt a pointer the first time the compiler emits an instruction that writes the zero register as a discard.
This post fixes all of that. It is the fifth part of the RTL-patterns section, and it builds directly on the sequential-logic and FSM material that came before. The plan is deliberate: we teach memory modeling generically first, with a tiny 16-entry 8-bit memory that you can hold entirely in your head, and only once the read/write timing, the storage-representation choice, and the in-cycle write-read hazard are completely clear do we bring in the worked example — the 32×32-bit RISC-V register file with two read ports, one write port, and x0 hardwired to zero. By the end you will be able to model any storage array in SystemC from scratch, choose combinational versus registered reads on purpose rather than by accident, predict cycle-by-cycle what a write-then-read of the same address returns, build a multi-port memory, and explain to a reviewer exactly why the storage is a plain C++ array and not an array of channels. That is the bar.
Prerequisites
- Part 1 — Modules, Ports & Signals. You need to be comfortable declaring an
SC_MODULE, bindingsc_signalchannels tosc_in/sc_outports, and registering a process insideSC_CTOR. - Part 4 — Processes & Sensitivity (SC_METHOD vs SC_THREAD). You need to know the difference between an
SC_METHODand anSC_THREAD, why combinational logic uses methods, and how a static sensitivity list is built withsensitive << clk.pos()versussensitive << addr. - Part 10 — Registered & Sequential Logic (Program Counter). You need the registered-transfer pattern: a clocked process that copies a next-value into a state-holding element on each rising edge, gated by an active-high synchronous reset. A memory's write port is exactly this pattern, applied to one element of an array per cycle.
- Part 11 — Finite State Machines (Moore vs Mealy). The two-process split (clocked process for state, combinational process for outputs) and the one-delta
sc_signalupdate delay carry straight over. A registered-read memory uses the FSM's registered-output timing; a combinational-read memory uses its combinational-output timing. - SystemC 2.3.x installed. All examples compile with a C++17 compiler (
g++9 or newer, orclang++10 or newer) against any 2.3.x SystemC build. The clock is alwayssc_clockwith a 10 ns period and 50% duty; reset is always active-high and synchronous unless stated otherwise.
Two ideas from the earlier parts do the heavy lifting here. The first is that an ordinary C++ member variable inside a module is perfectly good internal state — it does not need to be an sc_signal just because it changes over simulated time. The second is that sc_signal's buffered write (covered in Part 1 and used throughout Parts 10 and 11) is what gives a registered read its one-cycle delay. Keep both in mind; the entire post turns on them.
Mental model (first principles)
Strip a memory down to its essence and there are exactly three pieces: an array of storage cells, a write port, and one or more read ports.
The array of storage cells is just that — N words of W bits each. In hardware it is a bank of flip-flops (for a register file) or an SRAM array (for a larger memory). In SystemC it is most naturally a plain C++ array held as a private member of the module: uint8_t mem[16]; or uint32_t regs[32];. This is the single most important modeling decision in the whole post, so it is worth saying plainly: the storage is a plain C++ array, not an array of sc_signal. The reasons come straight from what sc_signal is for. An sc_signal is a primitive channel built to carry a value between processes with the evaluate/update protocol — a buffered write, a deferred commit in the update phase, and a value_changed_event that wakes anything sensitive to it. That machinery is exactly what you want on a wire that crosses between modules or that you sample on a clock edge. It is exactly what you do not want on a thousand internal storage cells: every cell would carry its own event queue and request-update bookkeeping, every write would enter the update phase, and your simulation would crawl while delivering no benefit, because no external process reads those cells through a port anyway. They are private state. Private state belongs in plain C++ members. We will demonstrate the performance and semantic consequences concretely later; for now, internalize the rule: storage cells are plain C++; channels are for communication.
The write port is a clocked, registered transfer — the Part 10 pattern, narrowed to one array element. A clocked process, sensitive to clk.pos(), checks a write-enable; if it is asserted (and reset is not), it performs mem[write_addr] = write_data. Because this is a plain C++ assignment inside the clocked process, the cell updates immediately within that process's execution — which is exactly the behavior of a flip-flop capturing its D input on the rising edge. The synchronous reset, when asserted, clears the array instead.
The read port is where the only real design choice lives, and it is the choice between combinational and registered reads.
A combinational read is an SC_METHOD sensitive to the read-address port. Its body is one line: read_data.write(mem[read_addr.read()]). When the address changes, the method runs, indexes the array, and drives the output. There is no clock in its sensitivity list, so the read tracks the address with no clock delay — like a multiplexer selecting one of N words. The crucial and slightly counter-intuitive consequence: because mem is a plain C++ array, writing a cell generates no event, so the combinational read does not automatically re-fire when memory content changes under a fixed address. It re-fires only when the address changes. In single-cycle designs this is exactly right, because every cycle presents a fresh address.
A registered read latches the addressed word into an output sc_signal on the clock edge: inside a clocked process, read_data_reg.write(mem[read_addr.read()]). Now the output is delayed by one cycle — the value you present an address for this cycle appears at the output next cycle. This is the same one-delta-plus-clock-edge delay you saw on a Moore FSM's registered output in Part 11. It costs a cycle of latency but it cleans up timing: the output is stable for a whole cycle and free of mid-cycle glitches.
Here is the whole structure in one diagram. Read it as: address and write data in on the left, the storage array in the middle, read data out on the right, with the write port gated by clock and write-enable.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
flowchart LR
WADDR([write_addr]) --> MEM
WDATA([write_data]) --> MEM
WEN([write_en]) --> MEM
CLK([clk]) --> MEM
RADDR([read_addr]) --> MEM
MEM["storage array
N x W bits
(plain C++)"] --> RDATA([read_data])
Everything else in this post is a variation on that picture: combinational versus registered read is a choice of which process drives read_data; a multi-port memory adds more read methods over the same array; the register file adds a second read port and an x0-is-zero rule; write-through adds a bypass comparator between the write port and the read port. Internalize the three pieces and one timing choice, and every memory you ever model is just filling in the details.
Beginner: First Principles
The simplest interesting memory is a single-port RAM: one address, one data-in, one write-enable, one data-out. We will make ours 16 words of 8 bits each — small enough that the whole array fits in a printed trace, big enough that addressing is non-trivial. The write is synchronous (on the rising clock edge, gated by write_en); the read is combinational (the output tracks the address). Reset is active-high and synchronous, and clears the array to zero.
// file: memory16x8.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// memory16x8.cpp -o memory16x8 -lsystemc
// Run: LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./memory16x8
#include <systemc.h>
#include <cstdint>
#include <iostream>
// 16 words x 8 bits. Synchronous write, combinational read, sync reset.
SC_MODULE(Memory16x8) {
sc_in<bool> clk;
sc_in<bool> rst; // active-high synchronous reset
sc_in<sc_uint<4> > addr; // 4-bit address: 0..15
sc_in<sc_uint<8> > write_data;
sc_in<bool> write_en;
sc_out<sc_uint<8> > read_data; // combinational read output
// Storage: plain C++ array. NOT an array of sc_signal — this is
// private internal state, accessed only by this module's processes.
uint8_t mem[16] = {}; // value-initialized: defined before reset
// --- Synchronous write process: clocked, gated by write_en + reset ---
void write_proc() {
if (rst.read()) {
for (int i = 0; i < 16; i++) mem[i] = 0;
} else if (write_en.read()) {
mem[addr.read()] = write_data.read();
}
}
// --- Combinational read process: output tracks the address ---
void read_proc() {
read_data.write(mem[addr.read()]);
}
SC_CTOR(Memory16x8) {
SC_METHOD(write_proc);
sensitive << clk.pos();
dont_initialize(); // clocked process: never run at t=0
SC_METHOD(read_proc);
sensitive << addr; // combinational: re-fires on address change
// No dont_initialize(): let it establish read_data at t=0.
}
};
SC_MODULE(Driver) {
sc_in<bool> clk;
sc_out<bool> rst;
sc_out<sc_uint<4> > addr;
sc_out<sc_uint<8> > write_data;
sc_out<bool> write_en;
sc_in<sc_uint<8> > read_data;
int step = 0;
void drive() {
// One step per rising clock edge. We write three cells, then read
// them back by changing only the address (write_en low).
switch (step) {
case 0: rst.write(false);
addr.write(3); write_data.write(0xAA); write_en.write(true); break;
case 1: addr.write(7); write_data.write(0xBB); write_en.write(true); break;
case 2: addr.write(12); write_data.write(0xCC); write_en.write(true); break;
case 3: write_en.write(false); addr.write(3); break; // read back x3
case 4: addr.write(7); break; // read back x7
case 5: addr.write(12); break; // read back x12
default: break;
}
step++;
}
void monitor() {
std::cout << "[" << sc_time_stamp() << "] step=" << step
<< " addr=" << addr.read()
<< " wen=" << write_en.read()
<< " read_data=0x" << std::hex << read_data.read().to_uint()
<< std::dec << "\n";
}
SC_CTOR(Driver) {
SC_METHOD(drive);
sensitive << clk.pos();
dont_initialize();
SC_METHOD(monitor);
sensitive << clk.pos();
dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS); // 10 ns period, 50% duty
sc_signal<bool> rst;
sc_signal<sc_uint<4> > addr;
sc_signal<sc_uint<8> > write_data;
sc_signal<bool> write_en;
sc_signal<sc_uint<8> > read_data;
Memory16x8 mem("mem");
mem.clk(clk); mem.rst(rst); mem.addr(addr);
mem.write_data(write_data); mem.write_en(write_en); mem.read_data(read_data);
Driver drv("drv");
drv.clk(clk); drv.rst(rst); drv.addr(addr);
drv.write_data(write_data); drv.write_en(write_en); drv.read_data(read_data);
// Hold reset high for one cycle, then let the driver take over.
rst.write(true);
write_en.write(false);
addr.write(0);
sc_start(10, SC_NS); // reset cycle (edge at 10 ns clears the array)
rst.write(false);
sc_start(70, SC_NS); // 7 more edges
sc_stop();
return 0;
}
Before reading the walkthrough, predict the output for yourself. Three writes (0xAA to address 3, 0xBB to 7, 0xCC to 12), then three reads of those same addresses with write_en low.
Expected output:
[0 s] step=0 addr=0 wen=0 read_data=0x0
[10 ns] step=1 addr=3 wen=1 read_data=0x0
[20 ns] step=2 addr=7 wen=1 read_data=0x0
[30 ns] step=3 addr=12 wen=1 read_data=0x0
[40 ns] step=4 addr=3 wen=0 read_data=0xaa
[50 ns] step=5 addr=7 wen=0 read_data=0xbb
[60 ns] step=6 addr=12 wen=0 read_data=0xcc
[70 ns] step=7 addr=12 wen=0 read_data=0xcc
Walk it carefully, because the one-cycle commit delay of every sc_signal write is the whole story. The clock period is 10 ns, so rising edges occur at 0, 10, 20, 30, … ns (an sc_clock that starts low has its first posedge at t=0). At t=0 the driver runs step=0 and writes its first setup onto the signals (rst→false, addr→3, write_data→0xAA, write_en→1) — but those writes are buffered and commit only after this edge. The monitor at t=0 therefore still sees the pre-edge committed values (addr=0, read_data=0x0, from the sc_main initialization), and prints the [0 s] line with step=0. Reset is high at t=0, so the write process zeroes the array.
At t=10 ns the t=0 writes have committed: rst=0, addr=3, write_data=0xAA, write_en=1. The write process performs mem[3] = 0xAA. But the combinational read fired when addr became 3 and read mem[3] as it was before this edge's write committed — which is 0. So the monitor at t=10 ns prints addr=3 read_data=0x0. This is the read-old behavior: the freshly written value is not yet visible to a read that already evaluated this cycle. The driver meanwhile sets up the second write (addr=7, 0xBB).
At t=20 ns: mem[7] = 0xBB commits, the read of address 7 still shows the old 0x0. At t=30 ns: mem[12] = 0xCC, read of address 12 shows 0x0. So far every read shows 0x0 because each read address was only just driven and the cell it points at had not yet been written when the read evaluated. Then the driver drops write_en and walks the address back over 3, 7, 12. At t=40 ns the read of address 3 returns the 0xAA written earlier; at t=50 ns address 7 returns 0xBB; at t=60 ns address 12 returns 0xCC. The final edge at t=70 ns holds address 12 and still reads 0xCC. The memory remembers.
Two operational details deserve attention.
First, the dont_initialize() on the write process. Without it the clocked write process would run once during the initialization delta, before any clock edge — and sample default-constructed inputs, possibly corrupting state. The rule from Part 10 stands: always call dont_initialize() on the clocked write process.
Second, the absence of dont_initialize() on the read process is deliberate. The combinational read runs once at initialization, indexes mem[0] (here 0, because we value-initialized the array with = {}), and establishes a defined read_data of 0x0 at t=0 — which is exactly the [0 s] line you see. A combinational output should be defined from the start, exactly as you saw with the FSM's combinational outputs in Part 11. (We will return in the Advanced section to why the array is value-initialized and what happens if it is not.)
A natural question at this point: why split write and read into two processes at all? Because they have different sensitivities. The write is clocked (sensitive to clk.pos()); the read is combinational (sensitive to addr). You cannot express both sensitivities in one SC_METHOD. The two-process split — clocked writer, combinational reader — is the canonical memory idiom, and it is the direct analog of the two-process FSM you already know.
Intermediate: How It Really Works
The Beginner example used a combinational read and a write/read sequence that never wrote and read the same address in the same cycle. Real designs do both of the things we sidestepped: they sometimes want a registered read (one cycle of latency, glitch-free output), and they routinely write one address while reading the same address in the same cycle. This section builds both, side by side, so the timing is unambiguous.
Combinational read vs registered read
Take the same 16×8 storage and give it a registered read: instead of a combinational SC_METHOD that drives the output the instant the address changes, latch the addressed word into an output sc_signal on the clock edge. The storage array is identical; only the read process moves from combinational to clocked.
// file: memory16x8_regread.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// memory16x8_regread.cpp -o memory16x8_regread -lsystemc
#include <systemc.h>
#include <cstdint>
#include <iostream>
SC_MODULE(Memory16x8Reg) {
sc_in<bool> clk;
sc_in<bool> rst;
sc_in<sc_uint<4> > addr;
sc_in<sc_uint<8> > write_data;
sc_in<bool> write_en;
sc_out<sc_uint<8> > read_data; // REGISTERED read output
uint8_t mem[16] = {}; // value-initialized: defined before reset
// Write and registered-read share the clock edge. Both are clocked.
void clocked_proc() {
if (rst.read()) {
for (int i = 0; i < 16; i++) mem[i] = 0;
read_data.write(0);
} else {
// Read latches the value addressed THIS cycle; appears next cycle.
read_data.write(mem[addr.read()]);
// Write commits this cycle (visible to reads from next cycle on).
if (write_en.read()) mem[addr.read()] = write_data.read();
}
}
SC_CTOR(Memory16x8Reg) {
SC_METHOD(clocked_proc);
sensitive << clk.pos();
dont_initialize();
}
};
SC_MODULE(Driver) {
sc_in<bool> clk;
sc_out<bool> rst;
sc_out<sc_uint<4> > addr;
sc_out<sc_uint<8> > write_data;
sc_out<bool> write_en;
sc_in<sc_uint<8> > read_data;
int step = 0;
void drive() {
switch (step) {
case 0: rst.write(false);
addr.write(5); write_data.write(0x42); write_en.write(true); break;
case 1: write_en.write(false); addr.write(5); break; // present read addr
case 2: addr.write(5); break; // hold; reg read emerges
case 3: addr.write(9); break; // change addr
default: break;
}
step++;
}
void monitor() {
std::cout << "[" << sc_time_stamp() << "] step=" << step
<< " addr=" << addr.read()
<< " read_data=0x" << std::hex << read_data.read().to_uint()
<< std::dec << "\n";
}
SC_CTOR(Driver) {
SC_METHOD(drive); sensitive << clk.pos(); dont_initialize();
SC_METHOD(monitor); sensitive << clk.pos(); dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst, write_en;
sc_signal<sc_uint<4> > addr;
sc_signal<sc_uint<8> > write_data;
sc_signal<sc_uint<8> > read_data;
Memory16x8Reg mem("mem");
mem.clk(clk); mem.rst(rst); mem.addr(addr);
mem.write_data(write_data); mem.write_en(write_en); mem.read_data(read_data);
Driver drv("drv");
drv.clk(clk); drv.rst(rst); drv.addr(addr);
drv.write_data(write_data); drv.write_en(write_en); drv.read_data(read_data);
rst.write(true); write_en.write(false); addr.write(0);
sc_start(10, SC_NS);
rst.write(false);
sc_start(50, SC_NS);
sc_stop();
return 0;
}
Expected output:
[0 s] step=0 addr=0 read_data=0x0
[10 ns] step=1 addr=5 read_data=0x0
[20 ns] step=2 addr=5 read_data=0x0
[30 ns] step=3 addr=5 read_data=0x42
[40 ns] step=4 addr=9 read_data=0x42
[50 ns] step=5 addr=9 read_data=0x0
This is the registered-read timing, and the one-cycle lag is the whole point. At t=0 the driver runs step 0 (deassert reset, set up the write of 0x42 to address 5); reset is still high at this edge so the clocked process zeroes the array and drives read_data=0. At t=10 ns the write setup has committed; the clocked process latches the registered read first (read_data = mem[5], which is still 0 because nothing has written cell 5 yet) and then performs mem[5] = 0x42. So the monitor at t=10 ns and again at t=20 ns prints read_data=0x0 — the addressed value as it was before this cycle's write, delayed one more cycle by the output register. At t=30 ns the registered output finally shows 0x42: it latched mem[5] at the t=20 ns edge, by which point mem[5] already held 0x42, and that latched value emerges one cycle later. It stays 0x42 while address 5 is held and through the t=40 ns edge after the address moves to 9 (the register still holds last cycle's sample). At t=50 ns the read of mem[9]=0 finally emerges. The output always lags the addressed value by one cycle relative to the Beginner combinational read. That is the registered-read tax — and it is exactly the Moore-output timing from Part 11.
Contrast this with the Beginner combinational read, where read_data followed the address within the same cycle (no lag). The trade is identical to Moore-vs-Mealy: combinational reads are fast but track inputs immediately (and would glitch if the address glitched); registered reads cost a cycle but produce a clean, stable, glitch-free output. Choose combinational reads when the consumer is combinational and needs the data this cycle (a CPU's register file feeding the ALU in a single-cycle design); choose registered reads when the memory is large, the read path is timing-critical, and the consumer samples on a clock edge anyway (most SRAM macros are registered-read for exactly this reason).
Write-then-read of the same address in one cycle
The hazard that bites everyone: what does a read of address A return in the same cycle a write lands on address A? The answer depends on the read style, and for combinational reads it depends on whether you build a bypass. There are two defensible behaviors:
- Read-old (no bypass): the read returns the value stored before this cycle's write. The new value becomes visible only from the next cycle.
- Write-through / bypass (read-new): the read returns the value being written this cycle, by comparing the read address against the write address and muxing the write data straight onto the read output.
Let us build both on the combinational-read 16×8 memory and watch them diverge. The driver presents the same address on the write and read paths in one cycle and we observe what the combinational read returns within that cycle.
// file: memory16x8_bypass.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// memory16x8_bypass.cpp -o memory16x8_bypass -lsystemc
#include <systemc.h>
#include <cstdint>
#include <iostream>
// Read-OLD memory: combinational read does NOT see this cycle's write.
SC_MODULE(MemReadOld) {
sc_in<bool> clk, rst, write_en;
sc_in<sc_uint<4> > waddr, raddr;
sc_in<sc_uint<8> > write_data;
sc_out<sc_uint<8> > read_data;
uint8_t mem[16] = {}; // value-initialized: defined before reset
void write_proc() {
if (rst.read()) { for (int i = 0; i < 16; i++) mem[i] = 0; }
else if (write_en.read()) mem[waddr.read()] = write_data.read();
}
void read_proc() {
// Pure read: no comparison to the write address.
read_data.write(mem[raddr.read()]);
}
SC_CTOR(MemReadOld) {
SC_METHOD(write_proc); sensitive << clk.pos(); dont_initialize();
SC_METHOD(read_proc); sensitive << raddr;
}
};
// Write-THROUGH memory: combinational read bypasses this cycle's write
// when raddr == waddr and write_en is asserted.
SC_MODULE(MemWriteThrough) {
sc_in<bool> clk, rst, write_en;
sc_in<sc_uint<4> > waddr, raddr;
sc_in<sc_uint<8> > write_data;
sc_out<sc_uint<8> > read_data;
uint8_t mem[16] = {}; // value-initialized: defined before reset
void write_proc() {
if (rst.read()) { for (int i = 0; i < 16; i++) mem[i] = 0; }
else if (write_en.read()) mem[waddr.read()] = write_data.read();
}
void read_proc() {
// Bypass: if writing the very cell we are reading this cycle,
// forward write_data straight to the read output.
if (write_en.read() && raddr.read() == waddr.read())
read_data.write(write_data.read());
else
read_data.write(mem[raddr.read()]);
}
SC_CTOR(MemWriteThrough) {
SC_METHOD(write_proc); sensitive << clk.pos(); dont_initialize();
// Read must wake on raddr, waddr, write_en, AND write_data because the
// bypass path reads all of them.
SC_METHOD(read_proc);
sensitive << raddr << waddr << write_en << write_data;
}
};
SC_MODULE(Driver) {
sc_in<bool> clk;
sc_out<bool> rst, write_en;
sc_out<sc_uint<4> > waddr, raddr;
sc_out<sc_uint<8> > write_data;
sc_in<sc_uint<8> > old_rd, wt_rd;
int step = 0;
void drive() {
switch (step) {
case 0: rst.write(false);
waddr.write(4); raddr.write(4);
write_data.write(0x77); write_en.write(true); break; // W&R same addr
case 1: write_en.write(false); raddr.write(0); break; // move read addr away
case 2: raddr.write(4); break; // come back to 4 -> combinational read re-fires
default: break;
}
step++;
}
void monitor() {
std::cout << "[" << sc_time_stamp() << "] step=" << step
<< " raddr=" << raddr.read() << " wen=" << write_en.read()
<< " read_old=0x" << std::hex << old_rd.read().to_uint()
<< " write_through=0x" << wt_rd.read().to_uint()
<< std::dec << "\n";
}
SC_CTOR(Driver) {
SC_METHOD(drive); sensitive << clk.pos(); dont_initialize();
SC_METHOD(monitor); sensitive << clk.pos(); dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst, write_en;
sc_signal<sc_uint<4> > waddr, raddr;
sc_signal<sc_uint<8> > write_data;
sc_signal<sc_uint<8> > old_rd, wt_rd;
MemReadOld m_old("m_old");
m_old.clk(clk); m_old.rst(rst); m_old.write_en(write_en);
m_old.waddr(waddr); m_old.raddr(raddr);
m_old.write_data(write_data); m_old.read_data(old_rd);
MemWriteThrough m_wt("m_wt");
m_wt.clk(clk); m_wt.rst(rst); m_wt.write_en(write_en);
m_wt.waddr(waddr); m_wt.raddr(raddr);
m_wt.write_data(write_data); m_wt.read_data(wt_rd);
Driver drv("drv");
drv.clk(clk); drv.rst(rst); drv.write_en(write_en);
drv.waddr(waddr); drv.raddr(raddr); drv.write_data(write_data);
drv.old_rd(old_rd); drv.wt_rd(wt_rd);
rst.write(true); write_en.write(false);
waddr.write(0); raddr.write(0); write_data.write(0);
sc_start(10, SC_NS);
rst.write(false);
sc_start(40, SC_NS);
sc_stop();
return 0;
}
Expected output:
[0 s] step=0 raddr=0 wen=0 read_old=0x0 write_through=0x0
[10 ns] step=1 raddr=4 wen=1 read_old=0x0 write_through=0x77
[20 ns] step=2 raddr=0 wen=0 read_old=0x0 write_through=0x0
[30 ns] step=3 raddr=4 wen=0 read_old=0x77 write_through=0x77
[40 ns] step=4 raddr=4 wen=0 read_old=0x77 write_through=0x77
Here is the divergence in one line. At t=10 ns we are simultaneously writing 0x77 to cell 4 and reading cell 4, with write_en asserted. The read-old memory returns 0x0 — the value cell 4 held before this cycle's write committed (the combinational read already evaluated mem[4] for this cycle when raddr became 4, and the clocked mem[4]=0x77 assignment happening this same edge does not retroactively change that already-computed read). The write-through memory returns 0x77 — its read process sees write_en && raddr==waddr and forwards write_data directly. That is the divergence: same stimulus, opposite in-cycle answer.
Then the driver moves the read address to 0 (t=20 ns: both read 0x0, the x0/cell-0 value) and back to 4 (t=30 ns). Now write_en is low and the address has changed back to 4, so the read-old memory's combinational read re-fires and finally sees the committed mem[4]=0x77. From t=30 ns on both memories agree: cell 4 holds 0x77. Note why we had to move the read address away and back — if we had simply held raddr=4, the read-old combinational read would never re-fire (its address never changed) and would show the stale 0x0 forever. That is the stale-combinational-read behavior of Corner 1 in the Advanced section, previewed here.
The two behaviors model genuinely different hardware. A plain flip-flop register file is read-old: the read multiplexer sees the current stored value, and the write does not propagate until the edge. A write-through register file adds a bypass comparator and a 2:1 mux on each read port so that an in-flight write is forwarded — this is exactly the forwarding path a pipelined CPU uses to avoid a read-after-write stall. The lesson: decide which behavior you want and build it explicitly. There is no default that is correct for every design.
raddr, waddr, write_en, and write_data in its sensitivity list. Every signal the process reads must appear, or the bypass output will not re-evaluate when those inputs change. This is the same "list every read" discipline you learned for combinational FSM processes in Part 11.Multi-port memories
A multi-port memory is the single-port pattern with more read methods over the same array. Each read port is an independent combinational SC_METHOD (or an independent registered process) with its own address input and data output, all indexing the one shared uint8_t mem[16] (or uint32_t regs[32]). The write port stays single — one clocked process. Because the storage is plain C++, sharing it across several read methods costs nothing: there is no channel contention, no resolution, no extra event traffic. Two read ports is exactly what the RISC-V register file needs, because almost every instruction reads two source operands at once. We build that next.
Advanced: Edge Cases & LRM Corners
Everything so far was generic memory modeling. Now we apply it to the worked example this section has been building toward: the RV32I register file. It is a 32-entry, 32-bit-wide memory with two combinational read ports (for the two source operands) and one synchronous write port (for the destination), plus one architectural twist — register x0 is hardwired to zero. Reads of x0 always return 0; writes to x0 are silently discarded. This is not a software convention; it is a hardware invariant the register file must enforce.
The register file is just the multi-port memory from the Intermediate section with N=32, W=32, two read ports, and the x0 rule layered on top. The address width is exactly 5 bits — sc_uint<5> indexes 0..31 and structurally cannot go out of bounds, which is the indexing-safety point we will return to.
The 32×32 register file
// file: regfile32x32.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// regfile32x32.cpp -o regfile32x32 -lsystemc
#include <systemc.h>
#include <cstdint>
#include <iostream>
// RV32I register file: 32 x 32-bit, 2 read ports, 1 write port.
// Combinational reads, synchronous write, active-high sync reset,
// x0 hardwired to zero.
SC_MODULE(RegFile) {
sc_in<bool> clk;
sc_in<bool> rst;
// Read port 1 (rs1) and read port 2 (rs2): combinational.
sc_in<sc_uint<5> > rs1_addr;
sc_out<sc_uint<32> > rs1_data;
sc_in<sc_uint<5> > rs2_addr;
sc_out<sc_uint<32> > rs2_data;
// Write port (rd): synchronous.
sc_in<sc_uint<5> > rd_addr;
sc_in<sc_uint<32> > rd_data;
sc_in<bool> wr_en;
// Storage: plain C++ array. 32 words of 32 bits. regs[0] exists but
// is never written (x0 invariant enforced in the write guard) and
// always reads as 0 (enforced in the read methods).
uint32_t regs[32] = {}; // value-initialized: defined before reset
// --- Synchronous write port ---
void write_proc() {
if (rst.read()) {
for (int i = 0; i < 32; i++) regs[i] = 0;
} else if (wr_en.read() && rd_addr.read() != 0) { // x0 write discarded
regs[rd_addr.read()] = rd_data.read();
}
}
// --- Combinational read port 1 ---
void read1_proc() {
sc_uint<5> a = rs1_addr.read();
rs1_data.write(a == 0 ? 0u : regs[a]); // x0 reads zero
}
// --- Combinational read port 2 ---
void read2_proc() {
sc_uint<5> a = rs2_addr.read();
rs2_data.write(a == 0 ? 0u : regs[a]); // x0 reads zero
}
SC_CTOR(RegFile) {
SC_METHOD(write_proc);
sensitive << clk.pos();
dont_initialize();
SC_METHOD(read1_proc);
sensitive << rs1_addr;
SC_METHOD(read2_proc);
sensitive << rs2_addr;
}
};
SC_MODULE(Driver) {
sc_in<bool> clk;
sc_out<bool> rst;
sc_out<sc_uint<5> > rs1_addr, rs2_addr, rd_addr;
sc_out<sc_uint<32> > rd_data;
sc_out<bool> wr_en;
sc_in<sc_uint<32> > rs1_data, rs2_data;
int step = 0;
void drive() {
switch (step) {
case 0: rst.write(false);
rd_addr.write(1); rd_data.write(0x11111111); wr_en.write(true); break;
case 1: rd_addr.write(2); rd_data.write(0x22222222); wr_en.write(true); break;
case 2: rd_addr.write(0); rd_data.write(0xDEADBEEF); wr_en.write(true); break; // x0 write
case 3: wr_en.write(false); rs1_addr.write(1); rs2_addr.write(2); break; // read x1,x2
case 4: rs1_addr.write(0); rs2_addr.write(2); break; // read x0,x2
default: break;
}
step++;
}
void monitor() {
std::cout << "[" << sc_time_stamp() << "] step=" << step
<< " rs1(@" << rs1_addr.read() << ")=0x" << std::hex
<< rs1_data.read().to_uint()
<< " rs2(@" << std::dec << rs2_addr.read() << ")=0x" << std::hex
<< rs2_data.read().to_uint() << std::dec << "\n";
}
SC_CTOR(Driver) {
SC_METHOD(drive); sensitive << clk.pos(); dont_initialize();
SC_METHOD(monitor); sensitive << clk.pos(); dont_initialize();
}
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<bool> rst, wr_en;
sc_signal<sc_uint<5> > rs1_addr, rs2_addr, rd_addr;
sc_signal<sc_uint<32> > rd_data, rs1_data, rs2_data;
RegFile rf("rf");
rf.clk(clk); rf.rst(rst);
rf.rs1_addr(rs1_addr); rf.rs1_data(rs1_data);
rf.rs2_addr(rs2_addr); rf.rs2_data(rs2_data);
rf.rd_addr(rd_addr); rf.rd_data(rd_data); rf.wr_en(wr_en);
Driver drv("drv");
drv.clk(clk); drv.rst(rst);
drv.rs1_addr(rs1_addr); drv.rs2_addr(rs2_addr); drv.rd_addr(rd_addr);
drv.rd_data(rd_data); drv.wr_en(wr_en);
drv.rs1_data(rs1_data); drv.rs2_data(rs2_data);
rst.write(true); wr_en.write(false);
rs1_addr.write(0); rs2_addr.write(0); rd_addr.write(0); rd_data.write(0);
sc_start(10, SC_NS);
rst.write(false);
sc_start(50, SC_NS);
sc_stop();
return 0;
}
Expected output:
[0 s] step=0 rs1(@0)=0x0 rs2(@0)=0x0
[10 ns] step=1 rs1(@0)=0x0 rs2(@0)=0x0
[20 ns] step=2 rs1(@0)=0x0 rs2(@0)=0x0
[30 ns] step=3 rs1(@0)=0x0 rs2(@0)=0x0
[40 ns] step=4 rs1(@1)=0x11111111 rs2(@2)=0x22222222
[50 ns] step=5 rs1(@0)=0x0 rs2(@2)=0x22222222
Walk the trace. Reset is high through t=0 (the driver's step 0 deasserts it for the next edge), so all 32 registers are zeroed at the t=0 edge. Then over three cycles the driver writes 0x11111111 to x1 (setup at the t=0 edge, committed and written at t=10 ns), 0x22222222 to x2 (written at t=20 ns), and attempts to write 0xDEADBEEF to x0 (at t=30 ns). The x0 write is silently discarded by the rd_addr != 0 guard in write_proc. Through these cycles the read addresses are still 0, so both read ports return 0x0 — that is the x0-reads-zero rule in action (and it is why the first four monitor lines show zeros). At t=40 ns the driver has dropped wr_en and presents rs1_addr=1, rs2_addr=2; the two combinational read ports return 0x11111111 and 0x22222222 simultaneously. At t=50 ns rs1_addr goes to 0 and the rs1 port returns 0x0 (x0 hardwired), while rs2 still reads x2 as 0x22222222. The x0 write of 0xDEADBEEF left no trace anywhere — exactly the invariant we needed.
Note three things this worked example demonstrates that the generic memory did not. First, two read ports operate independently and simultaneously over the one shared array — read1 and read2 are separate SC_METHODs with separate sensitivities (rs1_addr and rs2_addr), and they never interfere because the storage is plain C++ with no channel contention. Second, x0 is enforced in two places: the write guard (rd_addr != 0) prevents x0 from ever being written, and the read methods (a == 0 ? 0u : regs[a]) force a zero result even if regs[0] were somehow non-zero. The read-side guard is defense-in-depth; with the write guard in place regs[0] stays 0 after reset anyway, but enforcing it on both sides means a future bug on one path cannot violate the invariant. Third, the address type is sc_uint<5>, which can only represent 0..31 — there is no way to form an out-of-bounds index into regs[32] from a 5-bit address. That is the cheapest possible bounds check: make the index type structurally unable to overflow the array.
Corner 1: combinational read does not re-fire on a write
This is the subtle behavior we flagged in the Mental Model and must now confront head-on. Because regs is a plain C++ array, writing a register fires no SystemC event. The combinational read methods are sensitive only to their address ports. So if you write x5 on a clock edge and the read address has been sitting at 5 the whole time (it did not change), the read method does not automatically re-run, and rs1_data does not update to reflect the new value of x5 until the next time rs1_addr changes.
In a single-cycle CPU this never causes a bug, because every cycle decodes a new instruction, which drives new (or at least re-driven) rs1_addr/rs2_addr values, and that address change re-fires the read. But in a testbench it is a classic trap: write a register, leave the read address unchanged, and assert that the read shows the new value — the assertion fails, not because the model is wrong, but because nothing told the combinational read to re-evaluate. The fix in a testbench is to nudge the address (or step the simulation by SC_ZERO_TIME after re-presenting the address). The fix in a real design is: there is nothing to fix — the model is faithful to combinational-read hardware, where the new cell value propagates to the read output through gate delay only when the read decoder selects that cell.
Corner 2: why not sc_signal<sc_uint<32>> regs[32]?
It is tempting — especially when you want to see every register in a waveform — to declare the storage as an array of channels: sc_signal<sc_uint<32>> regs[32];. It compiles and it simulates correctly. But it is the wrong default for three reasons grounded in the channel semantics from Part 1 and the LRM's §6.4 update protocol.
First, performance. Every regs[i].write(...) is a buffered request-update: the kernel records the pending write, processes it in the update phase, and fires a value_changed_event. Thirty-two channels each carry that machinery. In a SoC where the register file is instantiated dozens of times, the update-phase traffic from internal storage that no external process observes is pure overhead.
Second, spurious wakeups. If you made the read methods sensitive to the regs[] signals (to get the read to re-fire on write — see Corner 1), every write would wake the reads, producing extra delta cycles and changing the very combinational-read timing we worked to establish. If you did not make them sensitive, you gained nothing over the plain array except the overhead.
Third, semantic clarity. sc_signal says "this value crosses between processes via a channel." Internal storage does not cross a channel boundary — it is read and written by processes inside the same module. Using a channel for it muddies the design's intent. The standard positions sc_signal as an inter-process communication primitive (§6.4.7); the right tool for module-private state is a plain member.
The one legitimate reason to reach for channels here is tracing: sc_trace cannot trace a plain C++ array element directly. If you must watch individual registers in GTKWave, the cleaner approach is to keep the plain array for the model and add a small debug-only process that copies the registers you care about onto traced signals, rather than paying the channel cost on the storage itself.
Corner 3: where sc_vector does and does not belong
sc_vector<T> (LRM §7.5) is SystemC's facility for a vector of named, hierarchical elements — typically modules, ports, or channels that need to be elaborated, named, and bound. If you were building a memory whose every cell had to be an independently bound sub-channel or sub-module (for example, a bank of memory modules you instantiate in a loop), sc_vector is exactly right: it gives each element a hierarchical name and participates correctly in elaboration. But for raw storage cells — bits that are read and written by arithmetic indexing inside one module — sc_vector<sc_signal<...>> buys you the same channel overhead as a raw array of sc_signal, plus elaboration cost, for no modeling benefit. The decision rule: use sc_vector when the elements need binding and hierarchy; use a plain C++ array when the elements are internal state indexed by value. A register file's 32 words are internal state. They are a plain array.
Corner 4: initialization and reset of memory contents
Two distinct concerns: what the array holds before reset, and what reset does to it.
Before reset, a plain C++ array member left fully untouched is not guaranteed zero — a member array of trivial type that you never initialize holds indeterminate values, and a combinational read that fires at t=0 (before the first reset edge) would read that garbage onto the output. This is exactly why every example in this post declares the storage as uint32_t regs[32] = {}; (or uint8_t mem[16] = {};) — the = {} value-initializes every cell to zero at construction, so the t=0 combinational read produces a clean 0x0 (which is the [0 s] line in each trace) rather than garbage. That is one defense; the other, which we also use, is to assert reset for at least one full cycle before any meaningful read so the first clock edge zeroes the array regardless. Belt and suspenders: value-initialize and reset. Real silicon register files frequently are not reset (zeroing thousands of flops costs area and power), and software is expected to initialize registers before reading them — but for a deterministic, reproducible simulation, zeroing on construction and on reset is the safe default and what we use.
Reset itself is the synchronous, active-high pattern from Part 10: when rst is high at the rising edge, the write process loops over the array and clears it, instead of performing the normal write. Because it is synchronous, the clear happens on the clock edge, not the instant rst rises. If you needed asynchronous reset (clear the moment rst asserts, between edges), you would move to an SC_CTHREAD with async_reset_signal_is, but for a register file synchronous reset is the common and simpler choice, and it matches the rest of this section's convention.
Version differences
The behavior described here — SC_METHOD sensitivity, sc_signal update-phase delay, plain-array internal state, sc_vector binding semantics, and dont_initialize() — is identical across SystemC 2.3.1, 2.3.3, and 2.3.4. Nothing in these memory and register-file patterns depends on a feature that changed across those releases; the examples compile and behave the same on any 2.3.x build with a C++17 compiler.
Hands-on exercise
Build a small byte-addressable scratchpad memory and then extend it, predicting every result before you run.
Start with a 32×8 single-port memory: one 5-bit address, one 8-bit data-in, one write-enable, one 8-bit data-out, active-high synchronous reset, and a combinational read. Drive it with a clocked stimulus that writes a recognizable pattern (mem[i] = i * 3 for a handful of addresses), then reads them back with write_en low. Confirm the reads return what you wrote, and confirm a read of an un-written cell returns 0 (because reset zeroed it). Before you run, write down the expected output line for each cycle; then check.
Next, add a registered-read variant of the same memory (latch the addressed byte into an output sc_signal on the clock edge). Drive both variants from the same stimulus and print their outputs side by side. Predict, before running, how many cycles the registered read lags the combinational read. You should see exactly one cycle of lag, matching the Moore-output timing from Part 11.
Then add a write-through bypass to the combinational-read variant: when the read address equals the write address and write_en is asserted, forward the write data straight to the read output. Construct a stimulus cycle that writes and reads the same address simultaneously, and verify the read-old variant returns the old value that cycle while the write-through variant returns the new value. This is the in-cycle hazard from the Intermediate section; reproducing it yourself is the point.
Finally, rebuild the structure as a 16×16 two-read-port, one-write-port memory — a miniature register file, but without the x0 rule. Then add the x0 rule (writes to address 0 discarded, reads of address 0 return 0) and confirm with a directed test that a write of 0xFFFF to address 0 leaves address 0 reading as 0. Predict the cycle on which each read result appears; verify against the simulator.
Hints
- Reuse the two-process pattern: a clocked
SC_METHODfor the write port (sensitive << clk.pos(), withdont_initialize()), and a combinationalSC_METHODper read port (sensitive << read_addr). - Storage is a plain C++ array (
uint8_t mem[32]/uint16_t regs[16]), never an array ofsc_signal. - For the write-through read process, remember to list every signal it reads in the sensitivity list:
read_addr,write_addr,write_en, andwrite_data. - For the registered-read variant, decide on purpose whether the read latches before or after the write line in the clocked process (read-old vs read-new) and document your choice.
- Use
sc_uint<5>for the 32-entry address andsc_uint<4>for the 16-entry address so the index type structurally cannot exceed the array bounds. - Drive your stimulus with a clocked
SC_METHODand anint stepcounter (the pattern used in every example above), not anSC_THREAD, unless you specifically want to drive inputs off the clock edge.
No solution is provided. The understanding lives in building it and predicting the traces before you run them.
Common mistakes
- Using an array of
sc_signalfor storage. It compiles, simulates correctly, and runs far slower than it should, because every cell carries the channel update-phase machinery for state no external process observes through a port. Use a plain C++ array for internal storage; reservesc_signalfor values that genuinely cross a process or module boundary. If you need to trace individual cells, copy them onto traced signals in a debug-only process rather than paying the channel cost on the storage itself.
- Expecting a combinational read to re-fire when memory content changes. Writing a plain C++ array element generates no event, so a combinational read
SC_METHOD(sensitive only to its address) does not re-run when you write the cell it is currently addressing. In a single-cycle design this is correct (a new address arrives every cycle). In a testbench it traps you: write a cell, leave the read address fixed, and the read output appears stale. Fix: change the address (or advance simulation) to force the read to re-evaluate — the model is faithful, the testbench just needs to present a new address.
- Confusing read-old with write-through on a same-cycle write and read. A plain synchronous-write / combinational-read memory returns the old value when you write and read the same address in one cycle; getting the new value requires an explicit bypass comparator and mux. Assuming write-through behavior from a read-old memory (or vice versa) produces off-by-one-cycle data bugs that survive directed testing and surface only on the back-to-back write/read pattern. Decide which behavior you want and build it explicitly.
- Forgetting to enforce x0 on both the write and read paths. The x0-hardwired-zero invariant needs the write guard (
rd_addr != 0, so x0 is never written) and ideally the read guard (return 0 when the read address is 0). Relying on only one path is fragile: a decoder bug that assertswr_enwithrd_addr=0is caught by the write guard; a stray value left inregs[0]is masked by the read guard. Enforce both, as defense-in-depth.
- Omitting
dont_initialize()on the clocked write process, or adding it to the combinational read. The clocked write process must not run at t=0 (it would sample default-constructed inputs); always calldont_initialize()on it. The combinational read process should not calldont_initialize(), so it establishes a defined output value at t=0. Swapping these is the same mis-placement bug you saw with FSMs in Part 11, and it leaves either a corrupted initial state or an undefined read output for the first cycle.
- Using an oversized, unchecked address type. Index a
regs[32]array with asc_uint<8>(or anint) that is not range-checked, and a stray value of 40 indexes out of bounds and corrupts adjacent memory. Make the address type exactly as wide as the array needs (sc_uint<5>for 32 entries) so the index structurally cannot exceed the bounds — the cheapest possible safety check.
Recap
After working through this post you can now:
- Model any storage array in SystemC as a plain C++ array member plus a clocked write process and one or more read processes, and explain from
sc_signal's §6.4 update semantics why the storage is a plain array and not an array of channels. - Build a synchronous-write, combinational-read single-port memory from scratch, with active-high synchronous reset, and predict its trace cycle-by-cycle.
- Choose between a combinational read (tracks the address this cycle) and a registered read (one cycle of latency, glitch-free), recognizing the trade as the same one you learned for Moore vs Mealy outputs.
- Predict what a write-then-read of the same address returns, and build both read-old (no bypass) and write-through (bypass comparator and mux) behaviors on purpose.
- Compose a multi-port memory as several independent read processes over one shared plain array, with a single clocked write port.
- Build the RV32I 32×32 register file as the worked example: two combinational read ports, one synchronous write port, x0 hardwired to zero on both the write and read paths,
sc_uint<5>addresses, and a verified simulation trace. - Reason about memory initialization and synchronous reset, and know when
sc_vectoris the right tool (binding and hierarchy) versus a plain array (internal state indexed by value). - Diagnose the common memory bugs on sight:
sc_signalstorage, stale combinational reads, read-old/write-through confusion, single-sided x0 enforcement,dont_initialize()mis-placement, and unchecked address widths.
Further reading
Standards
- IEEE Std 1666-2011, IEEE Standard for Standard SystemC Language Reference Manual, §5.2.16 (
SC_METHODsemantics), §5.2.18 (sensitivity lists anddont_initialize), §6.4 and §6.4.7 (sc_signalwrite semantics and its role as an inter-process channel), §7.5 (sc_vector).
Vendor and consortium documents
- Accellera Systems Initiative, SystemC 2.3.x User Guide — process kinds, channel mechanics, and
sc_vectorusage. - Doulos, SystemC Golden Reference Guide, ch. 4–5 —
SC_METHODsensitivity conventions andsc_vector.
Textbooks
- Bhasker, A SystemC Primer (2nd ed.), ch. 6 — memory and register-file modeling, plain-array internal state.
- Grötker, Liao, Martin, and Swan, System Design with SystemC, ch. 4–5 — channels versus internal state, port binding.
Training notes
- Berkeley CS152 lecture notes — RV32I register-file structure (2R1W ports, x0 hardwired), combinational versus registered reads.
- MIT 6.004 lecture notes — register-file and SRAM read/write timing pedagogy.
Real-world references
- PicoRV32 (public source, YosysHQ) — a compact RV32I core whose register file shows the x0-write-discard idiom this post teaches.
- Ibex
ibex_register_file_ff.sv(public source, lowRISC) — a flip-flop register file with explicit x0 handling, corroborating the read-zero / write-discard pattern.
Next in this section
→ Part 6: Memory-Interface Modeling — moving from raw storage to a memory interface: address-decode, byte enables, read/write strobes, and the request/response handshake that a CPU's load/store path uses to talk to memory. We teach the generic memory-mapped peripheral interface first, then build the RV32I data memory as the worked example.
Comments (0)
Leave a Comment