13. SystemC Tutorial - Memory-Interface Modeling
Why this matters
Rewritten 2026-06-05 with deeper first-principles material.
Almost everything a processor does that you can observe from the outside is a memory access. The CPU you are validating fetches an instruction — a read. It loads an operand — a read. It stores a result — a write. It reads a status register on a peripheral, polls a FIFO depth, kicks off a DMA descriptor, takes an interrupt and reads the cause register — every one of those is a transaction across a memory interface. When a virtual platform boots an operating system, the very first thing the boot ROM does after reset is a string of loads and stores against memory-mapped configuration registers, and if your model gets the interface contract of even one of those registers wrong — wrong read latency, wrong byte lane, wrong default value for an unmapped address — the boot wanders off into the weeds and the engineer staring at the trace spends a day discovering that the bug was a single byte enable that shifted the wrong way.
The memory array is the easy part. A C++ array holds the bytes; you can write a working storage model in five lines. The hard part — the part this post is about — is the interface: the precise agreement between the side that requests a transaction and the side that responds. Who drives the address, and on what edge? When is read data valid — the same cycle the address is presented, or the cycle after? Which bytes of the data bus actually get written when the requester only wants to update one byte? What happens when the address falls in a hole in the memory map? What does an unaligned access do? Get the array right and the contract wrong, and your model produces plausible-looking waveforms that are subtly, expensively incorrect.
This post teaches the interface contract from first principles using a deliberately tiny memory-mapped peripheral — a 16-byte register block on a generic bus — so that every concept (request/ready signals, address decode, byte enables, read-timing) lands without any RISC-V baggage. Only after the contract is solid do we bring in the worked example: the RV32I data memory, with its eight load/store widths (LW, SW, LB, LBU, LH, LHU, SB, SH), its byte-enable generation, its alignment rules, and its sign-versus-zero extension. By the end you will be able to specify a memory interface's timing contract precisely, decode an address against a memory map, generate correct byte enables for any sub-word access at any offset, choose between a combinational-read and a registered-read interface and state the latency each commits you to, and debug the three classic memory-interface bugs — the wrong byte enable for a store-halfword at offset 2, the accidental sign extension of an unsigned load, and the unaligned access that silently spans two words — on sight. That is the bar.
Prerequisites
- Part 1 — Modules, Ports & Signals. You must be comfortable declaring an
SC_MODULE, bindingsc_signalchannels tosc_in/sc_outports, and registering processes insideSC_CTOR. A memory interface is, structurally, a module with a handful of ports and two or three processes — nothing you have not seen, but assembled with unusual care about timing. - Part 4 — Processes & Sensitivity (SC_METHOD vs SC_THREAD). You need to know why a combinational read port is an
SC_METHODsensitive to its inputs, why a synchronous write port is anSC_METHODsensitive toclk.pos()(or anSC_CTHREAD), and whatdont_initialize()does. The read-timing contract is entirely a question of which process drives the read-data output and what it is sensitive to. - Part 12 — Memories & Register Files. That post built the storage array — the synchronous-write, combinational-read register file, and the
uint8_t mem[]array model. This post builds the interface around such an array. If the idea of a C++ array as the storage element andsc_signalports as the access path is hazy, re-read it first; everything here assumes it. - SystemC 2.3.x installed. All examples compile with a C++17 compiler (
g++9 or newer, orclang++10 or newer) against any 2.3.x SystemC build. The clock is alwayssc_clock("clk", 10, SC_NS)— a 10 ns period, 50% duty — and reset is active-high synchronous unless stated otherwise.
Three ideas from earlier parts do the heavy lifting here. First, sc_signal<T>: the buffered channel whose update-phase delay (Part 3) is what creates a registered read. Second, the C++ storage array: a plain uint8_t mem[] that mutates immediately on a C++ assignment, with no delta delay — distinct from a signal write. Third, the two-process split (Part 12): combinational read on one process, synchronous write on another. If any of those is shaky, pause and re-read before continuing.
Mental model (first principles)
Strip a memory interface down to its essence and it is a contract between a requester and a responder, expressed as a bundle of signals plus a timing agreement. The signals are easy to enumerate; the timing agreement is where all the subtlety lives. Forget RISC-V for now — picture any memory-mapped thing on a bus: a block of RAM, a UART's registers, a timer, a GPIO port. The requester (a CPU, a DMA engine, a test bench) wants to read or write some location inside that thing. The contract says: here is how you ask, and here is when you will get an answer.
A minimal contract needs these signals:
addr— which location. Driven by the requester.- a direction — read or write. This can be one
wr_enplus onerd_en(or a singlewritebit, or a transaction-type code). Driven by the requester. wdata— the data to write, on a write. Driven by the requester.byte_en— which lanes of the data bus participate. One bit per byte. Driven by the requester.rdata— the data read back, on a read. Driven by the responder.
That last line — driven by the responder — is the crux. Every signal except rdata is driven by the requester; rdata flows the other way. The interface is the seam where the two directions meet, and "who drives what when" is the first thing you must pin down before writing a line of code.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
flowchart LR
REQ["Requester
(CPU / DMA / TB)"]
RESP["Responder
(memory / peripheral)"]
REQ -->|addr| RESP
REQ -->|wr_en / rd_en| RESP
REQ -->|byte_en| RESP
REQ -->|wdata| RESP
RESP -->|rdata| REQ
Now the timing agreement. There are two questions, and the answers to both are design choices the contract must state explicitly.
Question one: when is read data valid? Two common answers. In a combinational-read interface, the responder drives rdata as a pure function of the current addr and control inputs — the moment the address settles, the data is there, in the same delta cycle, no clock edge required. This is how an asynchronous SRAM read port behaves and how a register file's read port usually behaves. In a registered-read interface, the responder captures the address on a clock edge and drives rdata one cycle later. This is how a synchronous SRAM behaves, and it is what you get from on-chip SRAM macros and most real memory controllers. The difference is exactly one clock of latency, and it is load-bearing: a requester pipeline built for one read-latency will be off by a cycle on the other. The contract must say which it is, and the requester must obey.
Question two: is the responder always ready, or can it stall? The simplest contract is fixed latency: the responder is always ready, every transaction completes in a known number of cycles. A more general contract adds a ready/valid handshake: the requester raises valid to say "I have a request," the responder raises ready to say "I can take it," and the transaction happens only on the cycle both are high. Handshakes let a slow responder apply back-pressure — exactly what you need when a memory has wait states or a FIFO is full. We will model a single-cycle handshake at RTL level in this post; the full transaction-level abstraction of this idea (sockets, blocking/non-blocking transport) is where Section 3 picks it up.
In SystemC, the storage is a C++ array and the interface is the ports plus the processes that honor the contract:
- The read process drives
rdata. If the contract is combinational-read, this is anSC_METHODsensitive toaddr, the read-enable, and any width/byte-enable inputs it reads — it recomputesrdatawhenever its inputs change, in the same delta. If the contract is registered-read, the read process is instead clocked (sensitive << clk.pos()), and it captures the addressed data intordataon the edge, so the value appears one cycle later via thesc_signalupdate-phase delay.
- The write process mutates the storage array. It is clocked —
sensitive << clk.pos(), withdont_initialize()— so writes land on rising edges. It readswr_enandbyte_enand writes only the enabled lanes:if (byte_en & 1) mem[base+0] = wdata bits 7..0;and so on. Becausemem[]is a plain C++ array, the write takes effect immediately when the process runs — there is no delta delay on the array itself; the delta delay lives only on thesc_signalports.
- Address decode is logic — usually inside whichever process reads the address — that compares the high bits of
addragainst a memory map to decide which responder (or which sub-region) owns the transaction, then indexes the chosen array with the low bits.
That is the whole mental model. A memory interface is a contract (signals + timing) honored by a small set of processes around a C++ storage array. Combinational-vs-registered read is a choice about which process drives rdata and what it is sensitive to. Byte enables are a mask the write process applies lane by lane. Address decode is a comparison on the high bits. Alignment and sign/zero extension — which dominate the RV32I worked example later — are shaping operations the responder applies to fit the contract's width and signedness rules. Everything else is filling in the details.
Beginner: First Principles
The simplest interesting memory interface is a small register block on a generic bus. We will build a 16-byte block — call it a peripheral, because in a real SoC it might be a UART's or a timer's register file — exposed through the contract signals from the mental model: addr, wr_en, rd_en, byte_en, wdata, rdata. We will make the read combinational (data valid the same cycle the address is presented) and the write synchronous (lands on the rising clock edge). No RISC-V, no funct3, no sign extension — just the bare contract, so you see the skeleton before any flesh.
The block holds 16 bytes, addressed 0x00–0x0F. The data bus is 32 bits wide, so a transaction touches an aligned word (4 bytes) and byte_en[3:0] selects which of those four bytes participate. The requester drives a word-aligned addr (bits [1:0] zero) and a byte_en mask; the responder reads or writes exactly the masked lanes.
// file: peripheral_block.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// peripheral_block.cpp -o peripheral_block -lsystemc
// Run: LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./peripheral_block
#include <systemc.h>
#include <cstdint>
#include <cstring>
#include <iostream>
#include <iomanip>
// A 16-byte register block on a 32-bit bus.
// - Combinational read: rdata valid the same delta the address settles.
// - Synchronous write: enabled bytes land on the rising clock edge.
SC_MODULE(PeripheralBlock) {
sc_in<bool> clk;
sc_in<sc_uint<32>> addr; // word-aligned byte address, 0x00..0x0C
sc_in<bool> wr_en; // request a write this cycle
sc_in<bool> rd_en; // request a read this cycle
sc_in<sc_uint<4>> byte_en; // one bit per byte lane
sc_in<sc_uint<32>> wdata; // write data
sc_out<sc_uint<32>> rdata; // read data (responder-driven)
static const int BLOCK_BYTES = 16;
uint8_t mem[BLOCK_BYTES];
// --- Combinational read: pure function of addr + rd_en ---
void read_proc() {
if (!rd_en.read()) { rdata.write(0); return; }
uint32_t a = (uint32_t)addr.read() & (BLOCK_BYTES - 1) & ~0x3u; // word base
uint32_t w = (uint32_t)mem[a]
| ((uint32_t)mem[a + 1] << 8)
| ((uint32_t)mem[a + 2] << 16)
| ((uint32_t)mem[a + 3] << 24);
rdata.write(w);
}
// --- Synchronous write: enabled lanes land on the rising edge ---
void write_proc() {
if (!wr_en.read()) return;
uint32_t a = (uint32_t)addr.read() & (BLOCK_BYTES - 1) & ~0x3u;
uint32_t d = (uint32_t)wdata.read();
uint32_t be = (uint32_t)byte_en.read();
if (be & 0x1) mem[a + 0] = (uint8_t)(d & 0xFF);
if (be & 0x2) mem[a + 1] = (uint8_t)((d >> 8) & 0xFF);
if (be & 0x4) mem[a + 2] = (uint8_t)((d >> 16) & 0xFF);
if (be & 0x8) mem[a + 3] = (uint8_t)((d >> 24) & 0xFF);
}
SC_CTOR(PeripheralBlock) {
memset(mem, 0, sizeof(mem));
SC_METHOD(read_proc);
sensitive << addr << rd_en << byte_en; // combinational: list every read input
SC_METHOD(write_proc);
sensitive << clk.pos(); // synchronous: clocked
dont_initialize(); // do not run before the first edge
}
};
SC_MODULE(Driver) {
sc_in<bool> clk;
sc_out<sc_uint<32>> addr;
sc_out<bool> wr_en;
sc_out<bool> rd_en;
sc_out<sc_uint<4>> byte_en;
sc_out<sc_uint<32>> wdata;
sc_in<sc_uint<32>> rdata;
void stim() {
// Cycle 0: idle.
addr.write(0); wr_en.write(false); rd_en.write(false);
byte_en.write(0); wdata.write(0);
wait();
// Cycle 1: write full word 0xDEADBEEF to offset 0x00 (be = 1111).
addr.write(0x00); wdata.write(0xDEADBEEF); byte_en.write(0xF);
wr_en.write(true); rd_en.write(false);
wait();
// Cycle 2: read it back (combinational — rdata valid this cycle).
addr.write(0x00); wr_en.write(false); rd_en.write(true);
wait(SC_ZERO_TIME); wait(SC_ZERO_TIME); // let the combinational read settle
std::cout << "[" << sc_time_stamp() << "] read 0x00 -> 0x"
<< std::hex << std::setw(8) << std::setfill('0')
<< (uint32_t)rdata.read() << std::dec << "\n";
wait();
// Cycle 3: write only byte lane 2 (be = 0100) with 0xFF.
addr.write(0x00); wdata.write(0x00FF0000); byte_en.write(0x4);
wr_en.write(true); rd_en.write(false);
wait();
// Cycle 4: read back — only lane 2 changed (0xDEADBEEF -> 0xDEFFBEEF).
addr.write(0x00); wr_en.write(false); rd_en.write(true);
wait(SC_ZERO_TIME); wait(SC_ZERO_TIME);
std::cout << "[" << sc_time_stamp() << "] read 0x00 -> 0x"
<< std::hex << std::setw(8) << std::setfill('0')
<< (uint32_t)rdata.read() << std::dec << "\n";
wait();
rd_en.write(false);
wait();
sc_stop();
}
SC_CTOR(Driver) {
SC_THREAD(stim);
sensitive << clk.pos();
}
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<sc_uint<32>> addr, wdata, rdata;
sc_signal<bool> wr_en, rd_en;
sc_signal<sc_uint<4>> byte_en;
PeripheralBlock dut("dut");
dut.clk(clk); dut.addr(addr); dut.wr_en(wr_en); dut.rd_en(rd_en);
dut.byte_en(byte_en); dut.wdata(wdata); dut.rdata(rdata);
Driver drv("drv");
drv.clk(clk); drv.addr(addr); drv.wr_en(wr_en); drv.rd_en(rd_en);
drv.byte_en(byte_en); drv.wdata(wdata); drv.rdata(rdata);
sc_start();
return 0;
}
Compile and run. Try to predict the two printed lines before you read the expected output.
Expected output:
[10 ns] read 0x00 -> 0xdeadbeef
[30 ns] read 0x00 -> 0xdeffbeef
Walk through it. The clock period is 10 ns, so rising edges are at 10, 20, 30, … ns. The driver is an SC_THREAD sensitive to clk.pos(), so each plain wait() parks it until the next rising edge. At the cycle-1 edge (t=10 ns) the driver asserts a full-word write of 0xDEADBEEF with byte_en = 1111. The write process fires on that same edge, sees wr_en high and all four lane bits set, and mutates mem[0..3] to 0xEF, 0xBE, 0xAD, 0xDE (little-endian: least-significant byte at the lowest address). Because mem[] is a C++ array, that write is immediate — there is no delta delay on the storage.
Still at t=10 ns, the driver raises rd_en and keeps addr = 0x00. The two wait(SC_ZERO_TIME) calls let the combinational read_proc evaluate and let its buffered rdata write commit through the update phase: read_proc is sensitive to rd_en, which just changed, so it wakes, reads mem[0..3], assembles 0xDEADBEEF, and writes rdata; the second zero-time wait lets that write commit so the thread reads the fresh value. The driver prints 0xdeadbeef at t=10 ns — the same cycle the read was requested. That is the combinational read-timing contract: data available the same cycle, no extra clock of latency. (One wait(SC_ZERO_TIME) is not enough — it samples rdata before the buffered write commits, returning the stale value. Two zero-time waits step past the update phase. This is the sc_signal request-update delay from Part 3 in action.)
At the cycle-3 edge (t=20 ns) the driver does a partial write: byte_en = 0100 (only lane 2), wdata = 0x00FF0000. The write process writes only mem[2] = 0xFF; lanes 0, 1, and 3 are untouched. This is the whole point of byte enables — a sub-word write that leaves its neighbors alone. The previous word 0xDEADBEEF becomes 0xDEFFBEEF: byte 2 (which was 0xAD) is now 0xFF, everything else unchanged. The read at t=30 ns confirms it.
Two operational details deserve a beat.
First, dont_initialize() on the write process. Without it, the write process would run once during the initialization phase — before the first clock edge — and if wr_en happened to be high at construction it would corrupt memory before the simulation properly starts. A clocked process that mutates state must always call dont_initialize(). There is no exception.
Second, the absence of dont_initialize() on the read process is deliberate. Letting the combinational read run once at initialization establishes a defined rdata (here, 0, because rd_en starts low) before any clock edge, so any downstream consumer has a defined value at t=0 rather than an indeterminate one. This mirrors exactly the combinational-process convention from the FSM and register-file posts.
sensitive << addr << rd_en << byte_en. SystemC has no always @(*) auto-sensitivity. If you read an input but forget to list it, the read will not re-evaluate when that input changes, and rdata will go stale. Every read inside a combinational SC_METHOD must correspond to one entry in the sensitive << chain.A natural question at this point: "why is the read combinational but the write clocked — why the asymmetry?" Because that is the contract this particular block commits to, and it is a common one: it models an asynchronous-read, synchronous-write SRAM, the same shape as the register file from Part 12. The read has no clock because the consumer wants the data now; the write is clocked because storage must change deterministically on a defined edge, not whenever an input wiggles. The next section shows the alternative — a registered read — and exactly how the timing contract changes when you pick it.
Intermediate: How It Really Works
The Beginner block committed to one read-timing contract (combinational) and lived at one fixed address. Real interfaces force two more decisions: which read-timing contract, and how the address selects among several responders. This section builds both, with traces that show the cycle-by-cycle difference.
The read-timing contract: combinational vs registered
The single most consequential interface decision is when read data is valid. Below is the same 16-byte block, but exposed two ways behind a switch. In the combinational variant the read port is an SC_METHOD sensitive to its inputs — rdata is valid the same delta the address settles. In the registered variant the read port is clocked — it captures the addressed word on the rising edge into rdata, which (because rdata is an sc_signal) becomes visible one delta after the edge, i.e. one full cycle after the address was presented.
// file: read_timing.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// read_timing.cpp -o read_timing -lsystemc
//
// Demonstrates the two read-timing contracts on the same storage:
// CombMem : rdata valid the same cycle the address is presented.
// RegisteredMem: rdata valid one cycle later (synchronous read).
#include <systemc.h>
#include <cstdint>
#include <cstring>
#include <iostream>
#include <iomanip>
// Combinational-read memory: rdata is a pure function of addr + rd_en.
SC_MODULE(CombMem) {
sc_in<bool> clk;
sc_in<sc_uint<32>> addr;
sc_in<bool> rd_en;
sc_out<sc_uint<32>> rdata;
static const int N = 16;
uint8_t mem[N];
void read_proc() {
if (!rd_en.read()) { rdata.write(0); return; }
uint32_t a = (uint32_t)addr.read() & (N - 1) & ~0x3u;
rdata.write((uint32_t)mem[a] | ((uint32_t)mem[a+1] << 8)
| ((uint32_t)mem[a+2] << 16) | ((uint32_t)mem[a+3] << 24));
}
SC_CTOR(CombMem) {
memset(mem, 0, sizeof(mem));
// Preload word 0x11223344 at offset 0.
mem[0] = 0x44; mem[1] = 0x33; mem[2] = 0x22; mem[3] = 0x11;
SC_METHOD(read_proc);
sensitive << addr << rd_en; // combinational
}
};
// Registered-read memory: address captured on the edge, rdata one cycle later.
SC_MODULE(RegisteredMem) {
sc_in<bool> clk;
sc_in<sc_uint<32>> addr;
sc_in<bool> rd_en;
sc_out<sc_uint<32>> rdata;
static const int N = 16;
uint8_t mem[N];
void read_proc() { // clocked
if (!rd_en.read()) { rdata.write(0); return; }
uint32_t a = (uint32_t)addr.read() & (N - 1) & ~0x3u;
rdata.write((uint32_t)mem[a] | ((uint32_t)mem[a+1] << 8)
| ((uint32_t)mem[a+2] << 16) | ((uint32_t)mem[a+3] << 24));
}
SC_CTOR(RegisteredMem) {
memset(mem, 0, sizeof(mem));
mem[0] = 0x44; mem[1] = 0x33; mem[2] = 0x22; mem[3] = 0x11;
SC_METHOD(read_proc);
sensitive << clk.pos(); // registered: one-cycle read latency
dont_initialize();
}
};
SC_MODULE(Probe) {
sc_in<bool> clk;
sc_out<sc_uint<32>> addr;
sc_out<bool> rd_en;
sc_in<sc_uint<32>> comb_rdata;
sc_in<sc_uint<32>> reg_rdata;
void run() {
// Present the request just after a posedge, in the cycle [10,20) ns, so the
// registered memory has NOT yet seen a capturing edge with the request high.
// Sample on negedges (mid-cycle), when all posedge activity has settled.
addr.write(0); rd_en.write(false);
wait(clk.posedge_event()); // t=10 ns
addr.write(0x00); rd_en.write(true); // request valid during [10,20) ns
wait(clk.negedge_event()); // t=15 ns: same cycle as the request
std::cout << "[" << sc_time_stamp() << "] cycle of the request; "
<< "comb=0x" << std::hex << std::setw(8) << std::setfill('0')
<< (uint32_t)comb_rdata.read() << " reg=0x"
<< std::setw(8) << std::setfill('0')
<< (uint32_t)reg_rdata.read() << std::dec << "\n";
wait(clk.negedge_event()); // t=25 ns: one cycle later
std::cout << "[" << sc_time_stamp() << "] one cycle later; "
<< "comb=0x" << std::hex << std::setw(8) << std::setfill('0')
<< (uint32_t)comb_rdata.read() << " reg=0x"
<< std::setw(8) << std::setfill('0')
<< (uint32_t)reg_rdata.read() << std::dec << "\n";
rd_en.write(false);
wait(clk.posedge_event());
sc_stop();
}
SC_CTOR(Probe) { SC_THREAD(run); }
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<sc_uint<32>> addr, comb_rdata, reg_rdata;
sc_signal<bool> rd_en;
CombMem cm("cm");
cm.clk(clk); cm.addr(addr); cm.rd_en(rd_en); cm.rdata(comb_rdata);
RegisteredMem rm("rm");
rm.clk(clk); rm.addr(addr); rm.rd_en(rd_en); rm.rdata(reg_rdata);
Probe p("p");
p.clk(clk); p.addr(addr); p.rd_en(rd_en);
p.comb_rdata(comb_rdata); p.reg_rdata(reg_rdata);
sc_start();
return 0;
}
Expected output:
[5 ns] cycle of the request; comb=0x11223344 reg=0x00000000
[15 ns] one cycle later; comb=0x11223344 reg=0x11223344
Read the trace carefully — it is the whole point of the section. The probe asserts rd_en and addr = 0x00 just after a rising edge, so the request is valid throughout that clock cycle, and we sample on the following negedge (mid-cycle, after all edge activity has settled). In the cycle of the request, the combinational memory's read process — sensitive to rd_en and addr, both of which just changed — has already fired in the same delta and driven comb_rdata = 0x11223344. The registered memory's read process is sensitive to clk.pos(); it does not fire from the input change, and on the edge that began this cycle the request was not yet asserted, so it captured nothing — reg_rdata is still its initial 0. One cycle later, the registered memory's clocked read process fires on the next rising edge (with the request now high), captures the addressed word, and one delta after drives reg_rdata = 0x11223344. The absolute timestamps (5 ns, 15 ns) depend on the clock's start phase; what matters is the relationship — the registered read lands exactly one clock period after the combinational read.
That one-cycle gap is the difference between the two contracts, and it is not cosmetic. If you build a CPU pipeline that presents a load address in cycle N and expects the data in cycle N — that is, you assumed combinational read — and you wire it to a registered-read SRAM, every load reads stale data and the CPU computes wrong answers, with no error message anywhere. If you assumed registered read and wired it to a combinational memory, you waste a pipeline stage waiting for data that was already there. The contract is a promise the requester and responder must agree on in writing.
When do you pick which? Combinational read models register files, small look-up ROMs, and asynchronous SRAM — anywhere the consumer needs the data within the same cycle and the array is small enough to read in one combinational path. Registered read models synchronous SRAM macros, on-chip memory blocks, and anything large enough that a same-cycle read would blow the timing budget — you trade one cycle of latency for a clean, registered output and a relaxed timing path. Real CPUs almost always use registered-read data memory and absorb the cycle in the pipeline; the RV32I single-cycle CPU in this series uses a combinational read precisely because it has no pipeline to absorb a cycle in. That is a deliberate simplification we will revisit when the CPU gains a pipeline in a later section.
Address decoding and the memory map
A real bus has more than one responder hanging off it. The requester drives a single address; address decode is the logic that examines the high bits of that address and selects which responder owns the transaction. The set of "which address range belongs to which device" rules is the memory map.
The example below puts two responders behind one address: a 16-byte RAM region at base 0x0000, and a single read-only 32-bit STATUS register at 0x1000. A read to anything in 0x0000–0x000F hits the RAM; a read to 0x1000 returns the status word; a read to any unmapped address returns a defined sentinel 0xBADADD12 (so that bus errors are visible rather than silently returning zero). This "decode high bits, index with low bits, default for holes" structure is exactly what every bus fabric does, scaled down.
// file: address_decode.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// address_decode.cpp -o address_decode -lsystemc
//
// One bus, two responders selected by address:
// 0x0000..0x000F : 16-byte RAM
// 0x1000 : read-only STATUS register
// anything else : unmapped -> sentinel 0xBADADD12
#include <systemc.h>
#include <cstdint>
#include <cstring>
#include <iostream>
#include <iomanip>
SC_MODULE(BusResponder) {
sc_in<bool> clk;
sc_in<sc_uint<32>> addr;
sc_in<bool> wr_en;
sc_in<bool> rd_en;
sc_in<sc_uint<4>> byte_en;
sc_in<sc_uint<32>> wdata;
sc_out<sc_uint<32>> rdata;
static const uint32_t RAM_BASE = 0x0000;
static const uint32_t RAM_BYTES = 16;
static const uint32_t STATUS_ADDR = 0x1000;
static const uint32_t UNMAPPED = 0xBADADD12;
uint8_t ram[RAM_BYTES];
uint32_t status; // read-only status register
bool in_ram(uint32_t a) const {
return a >= RAM_BASE && a < (RAM_BASE + RAM_BYTES);
}
void read_proc() {
if (!rd_en.read()) { rdata.write(0); return; }
uint32_t a = (uint32_t)addr.read();
if (in_ram(a)) {
uint32_t base = (a - RAM_BASE) & ~0x3u;
rdata.write((uint32_t)ram[base] | ((uint32_t)ram[base+1] << 8)
| ((uint32_t)ram[base+2] << 16) | ((uint32_t)ram[base+3] << 24));
} else if (a == STATUS_ADDR) {
rdata.write(status);
} else {
rdata.write(UNMAPPED); // defined value for a hole in the map
}
}
void write_proc() {
if (!wr_en.read()) return;
uint32_t a = (uint32_t)addr.read();
if (in_ram(a)) {
uint32_t base = (a - RAM_BASE) & ~0x3u;
uint32_t d = (uint32_t)wdata.read();
uint32_t be = (uint32_t)byte_en.read();
if (be & 0x1) ram[base+0] = (uint8_t)(d & 0xFF);
if (be & 0x2) ram[base+1] = (uint8_t)((d >> 8) & 0xFF);
if (be & 0x4) ram[base+2] = (uint8_t)((d >> 16) & 0xFF);
if (be & 0x8) ram[base+3] = (uint8_t)((d >> 24) & 0xFF);
}
// Writes to STATUS_ADDR or unmapped space are silently dropped (read-only).
}
SC_CTOR(BusResponder) {
memset(ram, 0, sizeof(ram));
status = 0x0000C0DE; // some fixed status value
SC_METHOD(read_proc);
sensitive << addr << rd_en;
SC_METHOD(write_proc);
sensitive << clk.pos();
dont_initialize();
}
};
SC_MODULE(Probe) {
sc_in<bool> clk;
sc_out<sc_uint<32>> addr;
sc_out<bool> wr_en;
sc_out<bool> rd_en;
sc_out<sc_uint<4>> byte_en;
sc_out<sc_uint<32>> wdata;
sc_in<sc_uint<32>> rdata;
void read_at(uint32_t a, const char* label) {
addr.write(a); rd_en.write(true); wr_en.write(false);
wait(SC_ZERO_TIME); wait(SC_ZERO_TIME); // settle combinational read + commit
std::cout << "[" << sc_time_stamp() << "] read " << label
<< " (0x" << std::hex << a << ") -> 0x"
<< std::setw(8) << std::setfill('0') << (uint32_t)rdata.read()
<< std::dec << "\n";
rd_en.write(false);
}
void run() {
addr.write(0); wr_en.write(false); rd_en.write(false); byte_en.write(0);
wait();
// Write a word into RAM, then read all three regions.
addr.write(0x0000); wdata.write(0xCAFEF00D); byte_en.write(0xF);
wr_en.write(true); rd_en.write(false);
wait();
wr_en.write(false);
read_at(0x0000, "RAM");
wait();
read_at(0x1000, "STATUS");
wait();
read_at(0x2000, "UNMAPPED");
wait();
sc_stop();
}
SC_CTOR(Probe) { SC_THREAD(run); sensitive << clk.pos(); }
};
int sc_main(int, char*[]) {
sc_clock clk("clk", 10, SC_NS);
sc_signal<sc_uint<32>> addr, wdata, rdata;
sc_signal<bool> wr_en, rd_en;
sc_signal<sc_uint<4>> byte_en;
BusResponder dut("dut");
dut.clk(clk); dut.addr(addr); dut.wr_en(wr_en); dut.rd_en(rd_en);
dut.byte_en(byte_en); dut.wdata(wdata); dut.rdata(rdata);
Probe p("p");
p.clk(clk); p.addr(addr); p.wr_en(wr_en); p.rd_en(rd_en);
p.byte_en(byte_en); p.wdata(wdata); p.rdata(rdata);
sc_start();
return 0;
}
Expected output:
[20 ns] read RAM (0x0) -> 0xcafef00d
[30 ns] read STATUS (0x1000) -> 0x0000c0de
[40 ns] read UNMAPPED (0x2000) -> 0xbadadd12
The decode is the three-way if/else if/else in read_proc: high-bit comparison selects the region, the low bits index within it, and the final else gives a defined value for every address that belongs to no device. That last clause matters more than it looks. A decode that returns 0 for unmapped space hides bugs — a typo in a base address reads zero, looks like an uninitialized-but-plausible value, and the real fault surfaces three modules downstream. Returning a loud sentinel (or, in a stricter model, asserting a bus-error response) turns a silent wrong-address bug into an obvious one. Real fabrics issue a DECERR/SLVERR response for exactly this reason.
Foreshadowing the handshake
Everything above used a fixed-latency contract: the responder is always ready, so there is no "can you take this now?" negotiation. The general case adds a ready/valid handshake — the requester raises valid, the responder raises ready, and the transfer happens only on a cycle where both are high — which lets a slow responder apply back-pressure. At RTL you would add valid/ready ports and gate the write/read on valid && ready. The transaction-level form of this same idea — where a single function call carries an entire transaction and the responder can model arbitrary latency without per-cycle signaling — is the subject of Section 3.
Advanced: Edge Cases & LRM Corners
Now the worked example. The RV32I data memory is one concrete instance of the contract from the Mental Model: a combinational-read, synchronous-write byte-addressable memory whose width-and-signedness code happens to be the instruction's funct3 field. Everything specific to RISC-V — eight access widths, byte-enable generation, alignment, sign versus zero extension — is shaping applied on top of the generic interface you already understand.
The RV32I width/sign contract
RISC-V encodes access width and signedness entirely in the 3-bit funct3 field. The decoder (covered in an earlier part) extracts it; the data memory consumes it. The same funct3 value means different things for loads versus stores, with the read/write direction supplied by separate mem_read/mem_write control bits.
funct3 |
Load | Store | Width | Extension |
|---|---|---|---|---|
000 |
LB |
SB |
byte (8-bit) | sign-extend (load) |
001 |
LH |
SH |
half (16-bit) | sign-extend (load) |
010 |
LW |
SW |
word (32-bit) | none |
100 |
LBU |
— | byte (8-bit) | zero-extend |
101 |
LHU |
— | half (16-bit) | zero-extend |
Two structural facts fall straight out of the generic interface. First, the byte-enable the store needs is a function of the access width and the byte offset within the aligned word — exactly the byte_en mask from the Beginner block, now generated from funct3 and addr[1:0] instead of being driven directly. Second, sign/zero extension is the responder shaping read data to the contract: the array holds raw bytes; the load shapes them into a 32-bit value whose upper bits follow the signedness rule.
Byte-enable generation
The store byte-enable mask is be = width_mask << addr[1:0], where the width mask is 0b0001 for a byte, 0b0011 for a halfword, and 0b1111 for a word. This is the single most error-prone line in the whole module, so look at it closely:
| Access | addr[1:0] |
Width mask | be = mask << offset |
|---|---|---|---|
SB |
00 |
0001 |
0001 |
SB |
01 |
0001 |
0010 |
SB |
10 |
0001 |
0100 |
SB |
11 |
0001 |
1000 |
SH |
00 |
0011 |
0011 |
SH |
10 |
0011 |
1100 |
SW |
00 |
1111 |
1111 |
The trap is SH at offset 2: the byte enable is 1100, not 0011. An engineer who thinks "halfword means the low two bytes" hard-codes 0011 and silently corrupts bytes 0 and 1 while leaving bytes 2 and 3 — the ones the program actually wanted to write — untouched. The shift-by-offset formula gets it right for every legal offset.
The RV32I data memory, complete
Below is a self-contained, runnable RV32I data memory using the combinational-read / synchronous-write contract. The read process shapes load data per funct3; the write process generates byte enables per funct3 and addr[1:0] and writes only the enabled lanes. An alignment check warns (it does not abort) on a misaligned halfword or word — the base ISA requires natural alignment, and a real core would raise an exception, which we foreshadow with a warning since exception plumbing belongs to the control unit.
// file: rv32i_dmem.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
// rv32i_dmem.cpp -o rv32i_dmem -lsystemc
// Run: LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./rv32i_dmem
//
// RV32I data memory: combinational load, synchronous store.
// Loads : LB(000) LH(001) LW(010) LBU(100) LHU(101) -- sign/zero extension
// Stores: SB(000) SH(001) SW(010) -- byte-enable generation
// Byte-addressable, little-endian. SystemC 2.3.x.
#include <systemc.h>
#include <cstdint>
#include <cstring>
#include <iostream>
#include <iomanip>
enum funct3_t : uint32_t {
F3_LB = 0x0, // also SB
F3_LH = 0x1, // also SH
F3_LW = 0x2, // also SW
F3_LBU = 0x4,
F3_LHU = 0x5,
};
SC_MODULE(DataMemory) {
sc_in<bool> clk;
sc_in<sc_uint<32>> addr; // byte address
sc_in<sc_uint<32>> wdata; // store data
sc_out<sc_uint<32>> rdata; // load data (responder-driven)
sc_in<bool> mem_read; // perform a load this cycle
sc_in<bool> mem_write; // perform a store this cycle
sc_in<sc_uint<3>> funct3; // width + sign selector
static const int MEM_BYTES = 4096;
uint8_t mem[MEM_BYTES];
// Warn (do not abort) on a misaligned access.
void check_align(uint32_t a, uint32_t width, const char* op) const {
if ((a & (width - 1)) != 0)
std::cerr << "[dmem] ALIGN WARNING: " << op << " at 0x"
<< std::hex << a << " not " << std::dec << width
<< "-byte aligned\n";
}
// --- Combinational load: shape raw bytes per funct3 ---
void read_proc() {
if (!mem_read.read()) { rdata.write(0); return; }
uint32_t a = (uint32_t)addr.read();
uint32_t f = (uint32_t)funct3.read();
uint32_t result = 0;
switch (f) {
case F3_LB: { // sign-extend byte
uint8_t raw = mem[a];
result = (uint32_t)(int32_t)(int8_t)raw; // int8_t -> int32_t sign-extends
break;
}
case F3_LH: { // sign-extend halfword
check_align(a, 2, "LH");
uint16_t raw = (uint16_t)mem[a] | ((uint16_t)mem[a+1] << 8);
result = (uint32_t)(int32_t)(int16_t)raw;
break;
}
case F3_LW: { // full word, no extension
check_align(a, 4, "LW");
result = (uint32_t)mem[a] | ((uint32_t)mem[a+1] << 8)
| ((uint32_t)mem[a+2] << 16) | ((uint32_t)mem[a+3] << 24);
break;
}
case F3_LBU: { // zero-extend byte
result = (uint32_t)mem[a]; // uint8_t -> uint32_t zero-extends
break;
}
case F3_LHU: { // zero-extend halfword
check_align(a, 2, "LHU");
result = (uint32_t)mem[a] | ((uint32_t)mem[a+1] << 8);
break;
}
default:
result = 0;
break;
}
rdata.write(result);
}
// --- Synchronous store: generate byte enables, write enabled lanes ---
void write_proc() {
if (!mem_write.read()) return;
uint32_t a = (uint32_t)addr.read();
uint32_t data = (uint32_t)wdata.read();
uint32_t f = (uint32_t)funct3.read();
uint32_t offset = a & 0x3; // addr[1:0]
uint32_t width_mask; // 0b0001 / 0b0011 / 0b1111
uint32_t width;
switch (f) {
case F3_LB: width_mask = 0x1; width = 1; break; // SB
case F3_LH: width_mask = 0x3; width = 2; check_align(a,2,"SH"); break; // SH
case F3_LW: width_mask = 0xF; width = 4; check_align(a,4,"SW"); break; // SW
default: width_mask = 0x0; width = 0; break;
}
uint32_t be = (width_mask << offset) & 0xF; // byte-enable mask
uint32_t word_base = a & ~0x3u; // aligned word base
// Lane the data so byte k of the word carries data byte (k - offset).
for (uint32_t lane = 0; lane < 4; ++lane) {
if (be & (1u << lane)) {
uint32_t data_byte_index = lane - offset; // which source byte feeds this lane
mem[word_base + lane] = (uint8_t)((data >> (8 * data_byte_index)) & 0xFF);
}
}
(void)width; // width is the alignment-check width; kept for clarity
}
SC_CTOR(DataMemory) {
memset(mem, 0, sizeof(mem));
SC_METHOD(read_proc);
sensitive << addr << mem_read << funct3;
SC_METHOD(write_proc);
sensitive << clk.pos();
dont_initialize();
}
};
static int passes = 0, fails = 0;
static void check(const char* name, uint32_t got, uint32_t exp) {
if (got == exp) { std::cout << " PASS " << name << " = 0x" << std::hex
<< std::setw(8) << std::setfill('0') << got
<< std::dec << "\n"; passes++; }
else { std::cout << " FAIL " << name << " got=0x" << std::hex
<< std::setw(8) << std::setfill('0') << got << " exp=0x"
<< std::setw(8) << std::setfill('0') << exp << std::dec
<< "\n"; fails++; }
}
SC_MODULE(TB) {
sc_clock clk{"clk", 10, SC_NS};
sc_signal<sc_uint<32>> addr, wdata, rdata;
sc_signal<bool> mem_read, mem_write;
sc_signal<sc_uint<3>> funct3;
DataMemory dut{"dut"};
void store(uint32_t a, uint32_t d, uint32_t f) {
addr.write(a); wdata.write(d); funct3.write(f);
mem_write.write(true); mem_read.write(false);
wait(clk.posedge_event()); wait(SC_ZERO_TIME);
mem_write.write(false);
}
uint32_t load(uint32_t a, uint32_t f) {
addr.write(a); funct3.write(f);
mem_read.write(true); mem_write.write(false);
wait(SC_ZERO_TIME); wait(SC_ZERO_TIME); // settle read + commit buffered rdata
uint32_t r = (uint32_t)rdata.read();
mem_read.write(false);
return r;
}
void run() {
mem_read.write(false); mem_write.write(false);
addr.write(0); wdata.write(0); funct3.write(0);
wait(clk.posedge_event());
std::cout << "\n=== SW / LW (word) ===\n";
store(0x00, 0xDEADBEEF, F3_LW);
check("LW @0x00", load(0x00, F3_LW), 0xDEADBEEF);
std::cout << "\n=== SB at all four byte lanes ===\n";
store(0x10, 0x00000000, F3_LW); // clear the word
store(0x10, 0xAA, F3_LB); // lane 0
store(0x11, 0xBB, F3_LB); // lane 1
store(0x12, 0xCC, F3_LB); // lane 2
store(0x13, 0xDD, F3_LB); // lane 3
check("LW @0x10 after 4x SB", load(0x10, F3_LW), 0xDDCCBBAA);
std::cout << "\n=== SH at offset 0 and offset 2 (byte-enable corner) ===\n";
store(0x20, 0x00000000, F3_LW);
store(0x20, 0xABCD, F3_LH); // low halfword, be = 0011
store(0x22, 0x1234, F3_LH); // high halfword, be = 1100 (NOT 0011)
check("LW @0x20 after SH lo+hi", load(0x20, F3_LW), 0x1234ABCD);
std::cout << "\n=== Sign vs zero extension (LB/LBU, LH/LHU) ===\n";
store(0x30, 0x80, F3_LB); // store byte 0x80
check("LB 0x80 (sign)", load(0x30, F3_LB), 0xFFFFFF80);
check("LBU 0x80 (zero)", load(0x30, F3_LBU), 0x00000080);
store(0x34, 0x8000, F3_LH); // store halfword 0x8000
check("LH 0x8000 (sign)", load(0x34, F3_LH), 0xFFFF8000);
check("LHU 0x8000 (zero)", load(0x34, F3_LHU), 0x00008000);
store(0x38, 0x7F, F3_LB); // positive byte must NOT sign-extend
check("LB 0x7F (positive)", load(0x38, F3_LB), 0x0000007F);
std::cout << "\n=== Partial overwrite (SB inside a word) ===\n";
store(0x40, 0x11223344, F3_LW);
store(0x42, 0xFF, F3_LB); // overwrite only lane 2
check("LW @0x40 after SB lane2", load(0x40, F3_LW), 0x11FF3344);
std::cout << "\n=== Misaligned LH (expect ALIGN WARNING, no crash) ===\n";
store(0x50, 0xAABBCCDD, F3_LW);
(void)load(0x51, F3_LH); // misaligned; prints a warning
std::cout << " [continued past misaligned access]\n";
passes++;
std::cout << "\n========================================\n";
std::cout << " PASS: " << passes << " FAIL: " << fails << "\n";
std::cout << "========================================\n";
sc_stop();
}
SC_CTOR(TB) {
dut.clk(clk); dut.addr(addr); dut.wdata(wdata); dut.rdata(rdata);
dut.mem_read(mem_read); dut.mem_write(mem_write); dut.funct3(funct3);
SC_THREAD(run);
}
};
int sc_main(int, char*[]) {
TB tb{"tb"};
sc_start();
return 0;
}
Expected output:
=== SW / LW (word) ===
PASS LW @0x00 = 0xdeadbeef
=== SB at all four byte lanes ===
PASS LW @0x10 after 4x SB = 0xddccbbaa
=== SH at offset 0 and offset 2 (byte-enable corner) ===
PASS LW @0x20 after SH lo+hi = 0x1234abcd
=== Sign vs zero extension (LB/LBU, LH/LHU) ===
PASS LB 0x80 (sign) = 0xffffff80
PASS LBU 0x80 (zero) = 0x00000080
PASS LH 0x8000 (sign) = 0xffff8000
PASS LHU 0x8000 (zero) = 0x00008000
PASS LB 0x7F (positive) = 0x0000007f
=== Partial overwrite (SB inside a word) ===
PASS LW @0x40 after SB lane2 = 0x11ff3344
=== Misaligned LH (expect ALIGN WARNING, no crash) ===
[dmem] ALIGN WARNING: LH at 0x51 not 2-byte aligned
[continued past misaligned access]
========================================
PASS: 10 FAIL: 0
========================================
(The ALIGN WARNING line is written to stderr; the rest is stdout. Depending on how your terminal interleaves the two streams it may appear at a slightly different position, but it always prints exactly once, for the one misaligned LH.)
Reading the load/store trace
Three results in that trace are worth dwelling on because each one is a bug magnet.
The four-SB word. Storing 0xAA, 0xBB, 0xCC, 0xDD to byte addresses 0x10, 0x11, 0x12, 0x13 and then reading the word back yields 0xDDCCBBAA, not 0xAABBCCDD. This is little-endian laid bare: the byte at the lowest address (0xAA at 0x10) is the least-significant byte of the word, so it sits in bits [7:0]. The byte-enable generation placed each store in its correct lane: SB at offset 0 set be = 0001, at offset 1 be = 0010, at offset 2 be = 0100, at offset 3 be = 1000, and the data-lane shift fed each store byte to the matching lane.
The SH-at-offset-2 corner. Storing 0xABCD at offset 0 (be = 0011) and 0x1234 at offset 2 (be = 1100) assembles 0x1234ABCD. Had the byte enable for the offset-2 store been hard-coded to 0011, the high store would have clobbered 0xABCD and left the high half of the word zero — a classic, hard-to-spot corruption. The width_mask << offset formula produced 1100 and routed 0x1234 into bytes 2 and 3.
Sign versus zero extension. The byte 0x80 loaded as LB becomes 0xFFFFFF80 (sign-extended: bit 7 is set, so the upper 24 bits fill with ones), but loaded as LBU becomes 0x00000080 (zero-extended). The bytes in memory are identical; only the load's signedness rule differs. The C++ cast chain (int32_t)(int8_t)raw does the sign extension because the standard guarantees that widening a signed type sign-extends; (uint32_t)raw zero-extends because widening an unsigned type fills with zeros. Get the intermediate cast wrong — write (int32_t)(uint8_t)raw for LB — and you zero-extend a signed load: correct for 0x00–0x7F, wrong for 0x80–0xFF, and invisible in any test whose data happens to be positive.
LRM corner: the array write is not a signal write
One subtlety from IEEE 1666-2011 §6.4 underpins the whole module. The store does mem[a] = ..., a plain C++ array assignment, which takes effect immediately — there is no request-update buffering, no delta delay. The sc_signal ports (addr, rdata, …) do buffer and commit in the update phase. This split is exactly what you want: the storage mutates synchronously on the clock edge inside the write process, and the combinational read sees the new bytes the next time it evaluates, while the ports still obey signal timing. If you mistakenly modeled each memory byte as an sc_signal, every store would incur an extra delta of latency and the model would be both slower and harder to reason about. Keep storage in a C++ array; keep only the interface in signals.
LRM corner: combinational read needs every input in its sensitivity list
The read process is sensitive << addr << mem_read << funct3 — all three of its inputs. Drop funct3 and a load that changes only the width code (say LB then LBU at the same address) will not re-evaluate, so rdata keeps the previous width's value and the sign/zero distinction silently fails. SystemC has no always @(*); the list is manual and must be complete. This is the same discipline the FSM post hammered on, applied to a memory's read port.
Hands-on exercise
Build a memory-mapped GPIO peripheral and exercise it through the generic contract, then layer the RV32I shaping rules onto a second memory.
Part A — the peripheral. Model a GPIO block with three 32-bit registers behind a bus interface (addr, wr_en, rd_en, byte_en, wdata, rdata), combinational read, synchronous write:
DIRat offset0x00— read/write direction register (1 = output, 0 = input).OUTat offset0x04— read/write output data register.INat offset0x08— read-only input data register; writes to it are silently dropped.
Decode the offset to pick the register, honor byte_en on writes to DIR and OUT, and return a sentinel (0xDEADC0DE) for any offset outside 0x00–0x08. Drive it with a probe that writes DIR = 0x000000FF, writes OUT = 0x000000A5, reads both back, attempts a write to IN (verify it does not change), and reads an unmapped offset (verify the sentinel). Predict every printed value before you run.
Part B — registered-read variant. Re-expose the same three registers behind a registered-read port (clocked read process, dont_initialize()). Drive the identical stimulus and watch every read appear one cycle later than in Part A. Confirm that a requester that reads on the same cycle it presents the address now sees stale data — this is the contract mismatch from the Intermediate section, reproduced by your own hand.
Part C — RV32I shaping on top. Take the DataMemory module from the Advanced section and add a tenth check to its testbench: store 0xFFFFFFFF as a word at 0x60, then load byte 0 as LB (expect 0xFFFFFFFF) and as LBU (expect 0x000000FF), and load the low halfword as LH (expect 0xFFFFFFFF) and as LHU (expect 0x0000FFFF). Then store a halfword 0xBEEF at offset 2 of a cleared word at 0x64 and confirm the word reads back 0xBEEF0000, proving your byte enable was 1100. No solution is provided.
Hints
- Reuse the
read_proc/write_proctwo-process split from the Beginner block verbatim; only the decode and the register set change. - For the registered-read variant, the only change is moving the read process's sensitivity from
addr/rd_entoclk.pos()and addingdont_initialize(). The body is identical. That one edit is the entire combinational-to-registered conversion. - Generate the GPIO byte enables exactly as the RV32I store does:
be = width_mask << (addr & 0x3). Even though the registers are word-aligned, writing the mask out makes the sub-word write path explicit. - For Part C, remember that
LBof0xFFis0xFFFFFFFF(sign bit set) butLBUof the same byte is0x000000FF. If they come out equal, your intermediate cast is wrong. - Drive stimulus from an
SC_THREADthatwait()s onclk.pos(); avoid driving inputs at non-clock-edge times unless you specifically want to study glitching.
Common mistakes
- Hard-coding the halfword byte enable to
0011. A store-halfword at byte offset 2 needsbe = 1100, not0011. The correct formula isbe = 0b0011 << addr[1:0], which yields1100at offset 2. Hard-coding0011writes the wrong two bytes — it corrupts bytes 0 and 1 and leaves the intended bytes 2 and 3 untouched. The bug is invisible whenever halfwords happen to be word-aligned (offset 0), so it survives casual testing and only bites on offset-2 stores. Fix: always shift the width mask by the byte offset.
- Sign-extending an unsigned load.
LBU/LHUzero-extend;LB/LHsign-extend. The C++ idiom for sign extension is(int32_t)(int8_t)raw(signed intermediate); for zero extension it is(uint32_t)raw(unsigned widening). Writing(int32_t)(int8_t)forLBU, or forgetting theint8_tintermediate forLB, produces results that are correct for values0x00–0x7Fand wrong for0x80–0xFF. Because most test data is small and positive, the bug hides until a value with the top bit set shows up. Fix: match the cast to the instruction's signedness, and always test a value with the high bit set (0x80,0xFF,0x8000).
- Assuming combinational read when the memory is registered (or vice versa). A combinational-read interface returns data the same cycle the address is presented; a registered-read interface returns it one cycle later. Wire a requester built for one contract to a responder honoring the other and every read is off by a cycle — stale data on a combinational consumer of a registered memory, or a wasted cycle the other way. There is no error message. Fix: write the read-latency into the interface contract and make both sides obey it; when in doubt, add an assertion that the consumer samples data on the cycle the contract specifies.
- Letting an unaligned access silently span two words. A word load at byte address
0x1001is not naturally aligned; in real hardware it either faults (base RV32I) or is split into two bus beats. A model that indexesmem[a..a+3]from an unalignedareads across a word boundary and returns a value no real aligned-only core would ever produce, masking a program bug. Fix: check alignment ((a & (width-1)) == 0) and either warn, fault, or explicitly model the split — but never let it pass silently as if it were a normal access.
- Modeling every memory byte as an
sc_signal. The storage array should be a plain C++ array (uint8_t mem[]); only the ports are signals. Wrapping each byte in ansc_signaladds a delta of latency to every store, slows simulation by orders of magnitude on a large memory, and buys nothing — the C++ array already gives correct synchronous-write/combinational-read behavior when the write happens inside a clocked process. Fix: storage in a C++ array, interface in signals.
- Returning
0for unmapped addresses. A decode whose default clause returns0hides addressing bugs: a wrong base address reads a plausible-looking zero, and the real fault surfaces far downstream. Fix: return a loud sentinel or assert a bus-error response so a hole in the memory map is immediately visible.
Recap
After working through this post you can now:
- State a memory interface as a contract — a bundle of signals (
addr, direction,byte_en,wdata,rdata) plus a timing agreement — and say who drives each signal. - Explain the difference between a combinational-read and a registered-read interface, predict the exact cycle read data is valid in each, and name the bug that results from mismatching the two.
- Decode an address against a memory map: compare high bits to select a responder, index with low bits, and return a defined value for unmapped holes.
- Generate a correct byte-enable mask for any sub-word access at any offset using
be = width_mask << addr[1:0], and explain whySHat offset 2 needs1100. - Implement the RV32I data memory's eight load/store widths, including little-endian byte assembly and the sign-versus-zero extension rule, as a worked instance of the generic contract.
- Diagnose the classic memory-interface bugs on sight: the wrong halfword byte enable, the sign-extended unsigned load, the silent unaligned access, the over-modeled
sc_signalstorage, and the silent unmapped read. - Recognize where a ready/valid handshake would slot in to add back-pressure, and why the transaction-level form of that idea is deferred to Section 3.
Further reading
Standards
- IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §4.2.1.3 (update phase and delta cycles), §5.2.16 (
SC_METHODsemantics), §5.2.18 (sensitivity lists anddont_initialize), §6.4 (sc_signalrequest-update write semantics). - The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA, §2.6 — load and store instructions, natural-alignment requirement, little-endian default for the base ISA.
Vendor and consortium documents
- ARM, AMBA AHB-Lite Protocol Specification (IHI0033) —
HSIZE/HADDRwidth-and-alignment contract and byte-strobe semantics, the industry analogue of thefunct3/byte_eninterface modeled here. - ARM, AMBA AXI Protocol Specification (AXI4) —
WSTRBwrite-strobe (byte-enable) lane model. - Accellera Systems Initiative, SystemC 2.3.x User Guide — channel and process mechanics relevant to memory modeling.
Textbooks
- Bhasker, A SystemC Primer (2nd ed.), ch. 6 — modeling memories, ports, and channels.
- Grötker, Liao, Martin, and Swan, System Design with SystemC, ch. 4–5 — the request-update model and channel design.
Real-world references
- Ibex
ibex_load_store_unit.sv(public source, lowRISC) — production RV32I load/store unit: byte-enable generation and load data extension in synthesizable RTL. - PicoRV32 (public source) — compact RV32I core whose memory interface shows the same width/alignment handling at the leaf level.
Next in this section
→ Part 7: Structural Composition & Hierarchy — composing the register file, decoder, PC, instruction memory, and the data memory built here into a working single-cycle RV32I datapath plus control unit, with correct port binding, sensitivity, and reset across a module hierarchy. The memory interface from this post becomes one block in that top-level composition.
Comments (0)
Leave a Comment