13. SystemC Tutorial - Memory-Interface Modeling

Why this matters

Rewritten 2026-06-05 with deeper first-principles material.

Almost everything a processor does that you can observe from the outside is a memory access. The CPU you are validating fetches an instruction — a read. It loads an operand — a read. It stores a result — a write. It reads a status register on a peripheral, polls a FIFO depth, kicks off a DMA descriptor, takes an interrupt and reads the cause register — every one of those is a transaction across a memory interface. When a virtual platform boots an operating system, the very first thing the boot ROM does after reset is a string of loads and stores against memory-mapped configuration registers, and if your model gets the interface contract of even one of those registers wrong — wrong read latency, wrong byte lane, wrong default value for an unmapped address — the boot wanders off into the weeds and the engineer staring at the trace spends a day discovering that the bug was a single byte enable that shifted the wrong way.

The memory array is the easy part. A C++ array holds the bytes; you can write a working storage model in five lines. The hard part — the part this post is about — is the interface: the precise agreement between the side that requests a transaction and the side that responds. Who drives the address, and on what edge? When is read data valid — the same cycle the address is presented, or the cycle after? Which bytes of the data bus actually get written when the requester only wants to update one byte? What happens when the address falls in a hole in the memory map? What does an unaligned access do? Get the array right and the contract wrong, and your model produces plausible-looking waveforms that are subtly, expensively incorrect.

This post teaches the interface contract from first principles using a deliberately tiny memory-mapped peripheral — a 16-byte register block on a generic bus — so that every concept (request/ready signals, address decode, byte enables, read-timing) lands without any RISC-V baggage. Only after the contract is solid do we bring in the worked example: the RV32I data memory, with its eight load/store widths (LW, SW, LB, LBU, LH, LHU, SB, SH), its byte-enable generation, its alignment rules, and its sign-versus-zero extension. By the end you will be able to specify a memory interface's timing contract precisely, decode an address against a memory map, generate correct byte enables for any sub-word access at any offset, choose between a combinational-read and a registered-read interface and state the latency each commits you to, and debug the three classic memory-interface bugs — the wrong byte enable for a store-halfword at offset 2, the accidental sign extension of an unsigned load, and the unaligned access that silently spans two words — on sight. That is the bar.

Prerequisites

  • Part 1 — Modules, Ports & Signals. You must be comfortable declaring an SC_MODULE, binding sc_signal channels to sc_in/sc_out ports, and registering processes inside SC_CTOR. A memory interface is, structurally, a module with a handful of ports and two or three processes — nothing you have not seen, but assembled with unusual care about timing.
  • Part 4 — Processes & Sensitivity (SC_METHOD vs SC_THREAD). You need to know why a combinational read port is an SC_METHOD sensitive to its inputs, why a synchronous write port is an SC_METHOD sensitive to clk.pos() (or an SC_CTHREAD), and what dont_initialize() does. The read-timing contract is entirely a question of which process drives the read-data output and what it is sensitive to.
  • Part 12 — Memories & Register Files. That post built the storage array — the synchronous-write, combinational-read register file, and the uint8_t mem[] array model. This post builds the interface around such an array. If the idea of a C++ array as the storage element and sc_signal ports as the access path is hazy, re-read it first; everything here assumes it.
  • SystemC 2.3.x installed. All examples compile with a C++17 compiler (g++ 9 or newer, or clang++ 10 or newer) against any 2.3.x SystemC build. The clock is always sc_clock("clk", 10, SC_NS) — a 10 ns period, 50% duty — and reset is active-high synchronous unless stated otherwise.

Three ideas from earlier parts do the heavy lifting here. First, sc_signal<T>: the buffered channel whose update-phase delay (Part 3) is what creates a registered read. Second, the C++ storage array: a plain uint8_t mem[] that mutates immediately on a C++ assignment, with no delta delay — distinct from a signal write. Third, the two-process split (Part 12): combinational read on one process, synchronous write on another. If any of those is shaky, pause and re-read before continuing.

Mental model (first principles)

Strip a memory interface down to its essence and it is a contract between a requester and a responder, expressed as a bundle of signals plus a timing agreement. The signals are easy to enumerate; the timing agreement is where all the subtlety lives. Forget RISC-V for now — picture any memory-mapped thing on a bus: a block of RAM, a UART's registers, a timer, a GPIO port. The requester (a CPU, a DMA engine, a test bench) wants to read or write some location inside that thing. The contract says: here is how you ask, and here is when you will get an answer.

A minimal contract needs these signals:

  • addr — which location. Driven by the requester.
  • a direction — read or write. This can be one wr_en plus one rd_en (or a single write bit, or a transaction-type code). Driven by the requester.
  • wdata — the data to write, on a write. Driven by the requester.
  • byte_en — which lanes of the data bus participate. One bit per byte. Driven by the requester.
  • rdata — the data read back, on a read. Driven by the responder.

That last line — driven by the responder — is the crux. Every signal except rdata is driven by the requester; rdata flows the other way. The interface is the seam where the two directions meet, and "who drives what when" is the first thing you must pin down before writing a line of code.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
flowchart LR
    REQ["Requester
(CPU / DMA / TB)"] RESP["Responder
(memory / peripheral)"] REQ -->|addr| RESP REQ -->|wr_en / rd_en| RESP REQ -->|byte_en| RESP REQ -->|wdata| RESP RESP -->|rdata| REQ

Now the timing agreement. There are two questions, and the answers to both are design choices the contract must state explicitly.

Question one: when is read data valid? Two common answers. In a combinational-read interface, the responder drives rdata as a pure function of the current addr and control inputs — the moment the address settles, the data is there, in the same delta cycle, no clock edge required. This is how an asynchronous SRAM read port behaves and how a register file's read port usually behaves. In a registered-read interface, the responder captures the address on a clock edge and drives rdata one cycle later. This is how a synchronous SRAM behaves, and it is what you get from on-chip SRAM macros and most real memory controllers. The difference is exactly one clock of latency, and it is load-bearing: a requester pipeline built for one read-latency will be off by a cycle on the other. The contract must say which it is, and the requester must obey.

Question two: is the responder always ready, or can it stall? The simplest contract is fixed latency: the responder is always ready, every transaction completes in a known number of cycles. A more general contract adds a ready/valid handshake: the requester raises valid to say "I have a request," the responder raises ready to say "I can take it," and the transaction happens only on the cycle both are high. Handshakes let a slow responder apply back-pressure — exactly what you need when a memory has wait states or a FIFO is full. We will model a single-cycle handshake at RTL level in this post; the full transaction-level abstraction of this idea (sockets, blocking/non-blocking transport) is where Section 3 picks it up.

In SystemC, the storage is a C++ array and the interface is the ports plus the processes that honor the contract:

  1. The read process drives rdata. If the contract is combinational-read, this is an SC_METHOD sensitive to addr, the read-enable, and any width/byte-enable inputs it reads — it recomputes rdata whenever its inputs change, in the same delta. If the contract is registered-read, the read process is instead clocked (sensitive << clk.pos()), and it captures the addressed data into rdata on the edge, so the value appears one cycle later via the sc_signal update-phase delay.
  1. The write process mutates the storage array. It is clocked — sensitive << clk.pos(), with dont_initialize() — so writes land on rising edges. It reads wr_en and byte_en and writes only the enabled lanes: if (byte_en & 1) mem[base+0] = wdata bits 7..0; and so on. Because mem[] is a plain C++ array, the write takes effect immediately when the process runs — there is no delta delay on the array itself; the delta delay lives only on the sc_signal ports.
  1. Address decode is logic — usually inside whichever process reads the address — that compares the high bits of addr against a memory map to decide which responder (or which sub-region) owns the transaction, then indexes the chosen array with the low bits.

That is the whole mental model. A memory interface is a contract (signals + timing) honored by a small set of processes around a C++ storage array. Combinational-vs-registered read is a choice about which process drives rdata and what it is sensitive to. Byte enables are a mask the write process applies lane by lane. Address decode is a comparison on the high bits. Alignment and sign/zero extension — which dominate the RV32I worked example later — are shaping operations the responder applies to fit the contract's width and signedness rules. Everything else is filling in the details.

Beginner: First Principles

The simplest interesting memory interface is a small register block on a generic bus. We will build a 16-byte block — call it a peripheral, because in a real SoC it might be a UART's or a timer's register file — exposed through the contract signals from the mental model: addr, wr_en, rd_en, byte_en, wdata, rdata. We will make the read combinational (data valid the same cycle the address is presented) and the write synchronous (lands on the rising clock edge). No RISC-V, no funct3, no sign extension — just the bare contract, so you see the skeleton before any flesh.

The block holds 16 bytes, addressed 0x00–0x0F. The data bus is 32 bits wide, so a transaction touches an aligned word (4 bytes) and byte_en[3:0] selects which of those four bytes participate. The requester drives a word-aligned addr (bits [1:0] zero) and a byte_en mask; the responder reads or writes exactly the masked lanes.

// file: peripheral_block.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            peripheral_block.cpp -o peripheral_block -lsystemc
// Run:   LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./peripheral_block

#include <systemc.h>
#include <cstdint>
#include <cstring>
#include <iostream>
#include <iomanip>

// A 16-byte register block on a 32-bit bus.
//   - Combinational read: rdata valid the same delta the address settles.
//   - Synchronous write:  enabled bytes land on the rising clock edge.
SC_MODULE(PeripheralBlock) {
  sc_in<bool>          clk;
  sc_in<sc_uint<32>>   addr;       // word-aligned byte address, 0x00..0x0C
  sc_in<bool>          wr_en;      // request a write this cycle
  sc_in<bool>          rd_en;      // request a read this cycle
  sc_in<sc_uint<4>>    byte_en;    // one bit per byte lane
  sc_in<sc_uint<32>>   wdata;      // write data
  sc_out<sc_uint<32>>  rdata;      // read data (responder-driven)

  static const int BLOCK_BYTES = 16;
  uint8_t mem[BLOCK_BYTES];

  // --- Combinational read: pure function of addr + rd_en ---
  void read_proc() {
    if (!rd_en.read()) { rdata.write(0); return; }
    uint32_t a = (uint32_t)addr.read() & (BLOCK_BYTES - 1) & ~0x3u; // word base
    uint32_t w = (uint32_t)mem[a]
               | ((uint32_t)mem[a + 1] << 8)
               | ((uint32_t)mem[a + 2] << 16)
               | ((uint32_t)mem[a + 3] << 24);
    rdata.write(w);
  }

  // --- Synchronous write: enabled lanes land on the rising edge ---
  void write_proc() {
    if (!wr_en.read()) return;
    uint32_t a  = (uint32_t)addr.read() & (BLOCK_BYTES - 1) & ~0x3u;
    uint32_t d  = (uint32_t)wdata.read();
    uint32_t be = (uint32_t)byte_en.read();
    if (be & 0x1) mem[a + 0] = (uint8_t)(d & 0xFF);
    if (be & 0x2) mem[a + 1] = (uint8_t)((d >> 8) & 0xFF);
    if (be & 0x4) mem[a + 2] = (uint8_t)((d >> 16) & 0xFF);
    if (be & 0x8) mem[a + 3] = (uint8_t)((d >> 24) & 0xFF);
  }

  SC_CTOR(PeripheralBlock) {
    memset(mem, 0, sizeof(mem));

    SC_METHOD(read_proc);
    sensitive << addr << rd_en << byte_en;   // combinational: list every read input

    SC_METHOD(write_proc);
    sensitive << clk.pos();                  // synchronous: clocked
    dont_initialize();                       // do not run before the first edge
  }
};

SC_MODULE(Driver) {
  sc_in<bool>          clk;
  sc_out<sc_uint<32>>  addr;
  sc_out<bool>         wr_en;
  sc_out<bool>         rd_en;
  sc_out<sc_uint<4>>   byte_en;
  sc_out<sc_uint<32>>  wdata;
  sc_in<sc_uint<32>>   rdata;

  void stim() {
    // Cycle 0: idle.
    addr.write(0); wr_en.write(false); rd_en.write(false);
    byte_en.write(0); wdata.write(0);
    wait();

    // Cycle 1: write full word 0xDEADBEEF to offset 0x00 (be = 1111).
    addr.write(0x00); wdata.write(0xDEADBEEF); byte_en.write(0xF);
    wr_en.write(true); rd_en.write(false);
    wait();

    // Cycle 2: read it back (combinational — rdata valid this cycle).
    addr.write(0x00); wr_en.write(false); rd_en.write(true);
    wait(SC_ZERO_TIME); wait(SC_ZERO_TIME);  // let the combinational read settle
    std::cout << "[" << sc_time_stamp() << "] read 0x00 -> 0x"
              << std::hex << std::setw(8) << std::setfill('0')
              << (uint32_t)rdata.read() << std::dec << "\n";
    wait();

    // Cycle 3: write only byte lane 2 (be = 0100) with 0xFF.
    addr.write(0x00); wdata.write(0x00FF0000); byte_en.write(0x4);
    wr_en.write(true); rd_en.write(false);
    wait();

    // Cycle 4: read back — only lane 2 changed (0xDEADBEEF -> 0xDEFFBEEF).
    addr.write(0x00); wr_en.write(false); rd_en.write(true);
    wait(SC_ZERO_TIME); wait(SC_ZERO_TIME);
    std::cout << "[" << sc_time_stamp() << "] read 0x00 -> 0x"
              << std::hex << std::setw(8) << std::setfill('0')
              << (uint32_t)rdata.read() << std::dec << "\n";
    wait();

    rd_en.write(false);
    wait();
    sc_stop();
  }

  SC_CTOR(Driver) {
    SC_THREAD(stim);
    sensitive << clk.pos();
  }
};

int sc_main(int, char*[]) {
  sc_clock           clk("clk", 10, SC_NS);
  sc_signal<sc_uint<32>> addr, wdata, rdata;
  sc_signal<bool>        wr_en, rd_en;
  sc_signal<sc_uint<4>>  byte_en;

  PeripheralBlock dut("dut");
  dut.clk(clk); dut.addr(addr); dut.wr_en(wr_en); dut.rd_en(rd_en);
  dut.byte_en(byte_en); dut.wdata(wdata); dut.rdata(rdata);

  Driver drv("drv");
  drv.clk(clk); drv.addr(addr); drv.wr_en(wr_en); drv.rd_en(rd_en);
  drv.byte_en(byte_en); drv.wdata(wdata); drv.rdata(rdata);

  sc_start();
  return 0;
}

Compile and run. Try to predict the two printed lines before you read the expected output.

Expected output:

[10 ns] read 0x00 -> 0xdeadbeef
[30 ns] read 0x00 -> 0xdeffbeef

Walk through it. The clock period is 10 ns, so rising edges are at 10, 20, 30, … ns. The driver is an SC_THREAD sensitive to clk.pos(), so each plain wait() parks it until the next rising edge. At the cycle-1 edge (t=10 ns) the driver asserts a full-word write of 0xDEADBEEF with byte_en = 1111. The write process fires on that same edge, sees wr_en high and all four lane bits set, and mutates mem[0..3] to 0xEF, 0xBE, 0xAD, 0xDE (little-endian: least-significant byte at the lowest address). Because mem[] is a C++ array, that write is immediate — there is no delta delay on the storage.

Still at t=10 ns, the driver raises rd_en and keeps addr = 0x00. The two wait(SC_ZERO_TIME) calls let the combinational read_proc evaluate and let its buffered rdata write commit through the update phase: read_proc is sensitive to rd_en, which just changed, so it wakes, reads mem[0..3], assembles 0xDEADBEEF, and writes rdata; the second zero-time wait lets that write commit so the thread reads the fresh value. The driver prints 0xdeadbeef at t=10 ns — the same cycle the read was requested. That is the combinational read-timing contract: data available the same cycle, no extra clock of latency. (One wait(SC_ZERO_TIME) is not enough — it samples rdata before the buffered write commits, returning the stale value. Two zero-time waits step past the update phase. This is the sc_signal request-update delay from Part 3 in action.)

At the cycle-3 edge (t=20 ns) the driver does a partial write: byte_en = 0100 (only lane 2), wdata = 0x00FF0000. The write process writes only mem[2] = 0xFF; lanes 0, 1, and 3 are untouched. This is the whole point of byte enables — a sub-word write that leaves its neighbors alone. The previous word 0xDEADBEEF becomes 0xDEFFBEEF: byte 2 (which was 0xAD) is now 0xFF, everything else unchanged. The read at t=30 ns confirms it.

Two operational details deserve a beat.

First, dont_initialize() on the write process. Without it, the write process would run once during the initialization phase — before the first clock edge — and if wr_en happened to be high at construction it would corrupt memory before the simulation properly starts. A clocked process that mutates state must always call dont_initialize(). There is no exception.

Second, the absence of dont_initialize() on the read process is deliberate. Letting the combinational read run once at initialization establishes a defined rdata (here, 0, because rd_en starts low) before any clock edge, so any downstream consumer has a defined value at t=0 rather than an indeterminate one. This mirrors exactly the combinational-process convention from the FSM and register-file posts.

Note The read process lists every input it reads in its sensitivity list: sensitive << addr << rd_en << byte_en. SystemC has no always @(*) auto-sensitivity. If you read an input but forget to list it, the read will not re-evaluate when that input changes, and rdata will go stale. Every read inside a combinational SC_METHOD must correspond to one entry in the sensitive << chain.

A natural question at this point: "why is the read combinational but the write clocked — why the asymmetry?" Because that is the contract this particular block commits to, and it is a common one: it models an asynchronous-read, synchronous-write SRAM, the same shape as the register file from Part 12. The read has no clock because the consumer wants the data now; the write is clocked because storage must change deterministically on a defined edge, not whenever an input wiggles. The next section shows the alternative — a registered read — and exactly how the timing contract changes when you pick it.

Intermediate: How It Really Works

The Beginner block committed to one read-timing contract (combinational) and lived at one fixed address. Real interfaces force two more decisions: which read-timing contract, and how the address selects among several responders. This section builds both, with traces that show the cycle-by-cycle difference.

The read-timing contract: combinational vs registered

The single most consequential interface decision is when read data is valid. Below is the same 16-byte block, but exposed two ways behind a switch. In the combinational variant the read port is an SC_METHOD sensitive to its inputs — rdata is valid the same delta the address settles. In the registered variant the read port is clocked — it captures the addressed word on the rising edge into rdata, which (because rdata is an sc_signal) becomes visible one delta after the edge, i.e. one full cycle after the address was presented.

// file: read_timing.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            read_timing.cpp -o read_timing -lsystemc
//
// Demonstrates the two read-timing contracts on the same storage:
//   CombMem      : rdata valid the same cycle the address is presented.
//   RegisteredMem: rdata valid one cycle later (synchronous read).

#include <systemc.h>
#include <cstdint>
#include <cstring>
#include <iostream>
#include <iomanip>

// Combinational-read memory: rdata is a pure function of addr + rd_en.
SC_MODULE(CombMem) {
  sc_in<bool>          clk;
  sc_in<sc_uint<32>>   addr;
  sc_in<bool>          rd_en;
  sc_out<sc_uint<32>>  rdata;

  static const int N = 16;
  uint8_t mem[N];

  void read_proc() {
    if (!rd_en.read()) { rdata.write(0); return; }
    uint32_t a = (uint32_t)addr.read() & (N - 1) & ~0x3u;
    rdata.write((uint32_t)mem[a] | ((uint32_t)mem[a+1] << 8)
              | ((uint32_t)mem[a+2] << 16) | ((uint32_t)mem[a+3] << 24));
  }

  SC_CTOR(CombMem) {
    memset(mem, 0, sizeof(mem));
    // Preload word 0x11223344 at offset 0.
    mem[0] = 0x44; mem[1] = 0x33; mem[2] = 0x22; mem[3] = 0x11;
    SC_METHOD(read_proc);
    sensitive << addr << rd_en;      // combinational
  }
};

// Registered-read memory: address captured on the edge, rdata one cycle later.
SC_MODULE(RegisteredMem) {
  sc_in<bool>          clk;
  sc_in<sc_uint<32>>   addr;
  sc_in<bool>          rd_en;
  sc_out<sc_uint<32>>  rdata;

  static const int N = 16;
  uint8_t mem[N];

  void read_proc() {                 // clocked
    if (!rd_en.read()) { rdata.write(0); return; }
    uint32_t a = (uint32_t)addr.read() & (N - 1) & ~0x3u;
    rdata.write((uint32_t)mem[a] | ((uint32_t)mem[a+1] << 8)
              | ((uint32_t)mem[a+2] << 16) | ((uint32_t)mem[a+3] << 24));
  }

  SC_CTOR(RegisteredMem) {
    memset(mem, 0, sizeof(mem));
    mem[0] = 0x44; mem[1] = 0x33; mem[2] = 0x22; mem[3] = 0x11;
    SC_METHOD(read_proc);
    sensitive << clk.pos();          // registered: one-cycle read latency
    dont_initialize();
  }
};

SC_MODULE(Probe) {
  sc_in<bool>          clk;
  sc_out<sc_uint<32>>  addr;
  sc_out<bool>         rd_en;
  sc_in<sc_uint<32>>   comb_rdata;
  sc_in<sc_uint<32>>   reg_rdata;

  void run() {
    // Present the request just after a posedge, in the cycle [10,20) ns, so the
    // registered memory has NOT yet seen a capturing edge with the request high.
    // Sample on negedges (mid-cycle), when all posedge activity has settled.
    addr.write(0); rd_en.write(false);
    wait(clk.posedge_event());                // t=10 ns
    addr.write(0x00); rd_en.write(true);      // request valid during [10,20) ns

    wait(clk.negedge_event());                // t=15 ns: same cycle as the request
    std::cout << "[" << sc_time_stamp() << "] cycle of the request; "
              << "comb=0x" << std::hex << std::setw(8) << std::setfill('0')
              << (uint32_t)comb_rdata.read() << " reg=0x"
              << std::setw(8) << std::setfill('0')
              << (uint32_t)reg_rdata.read() << std::dec << "\n";

    wait(clk.negedge_event());                // t=25 ns: one cycle later
    std::cout << "[" << sc_time_stamp() << "] one cycle later;    "
              << "comb=0x" << std::hex << std::setw(8) << std::setfill('0')
              << (uint32_t)comb_rdata.read() << " reg=0x"
              << std::setw(8) << std::setfill('0')
              << (uint32_t)reg_rdata.read() << std::dec << "\n";
    rd_en.write(false);
    wait(clk.posedge_event());
    sc_stop();
  }

  SC_CTOR(Probe) { SC_THREAD(run); }
};

int sc_main(int, char*[]) {
  sc_clock           clk("clk", 10, SC_NS);
  sc_signal<sc_uint<32>> addr, comb_rdata, reg_rdata;
  sc_signal<bool>        rd_en;

  CombMem cm("cm");
  cm.clk(clk); cm.addr(addr); cm.rd_en(rd_en); cm.rdata(comb_rdata);

  RegisteredMem rm("rm");
  rm.clk(clk); rm.addr(addr); rm.rd_en(rd_en); rm.rdata(reg_rdata);

  Probe p("p");
  p.clk(clk); p.addr(addr); p.rd_en(rd_en);
  p.comb_rdata(comb_rdata); p.reg_rdata(reg_rdata);

  sc_start();
  return 0;
}

Expected output:

[5 ns] cycle of the request; comb=0x11223344 reg=0x00000000
[15 ns] one cycle later;    comb=0x11223344 reg=0x11223344

Read the trace carefully — it is the whole point of the section. The probe asserts rd_en and addr = 0x00 just after a rising edge, so the request is valid throughout that clock cycle, and we sample on the following negedge (mid-cycle, after all edge activity has settled). In the cycle of the request, the combinational memory's read process — sensitive to rd_en and addr, both of which just changed — has already fired in the same delta and driven comb_rdata = 0x11223344. The registered memory's read process is sensitive to clk.pos(); it does not fire from the input change, and on the edge that began this cycle the request was not yet asserted, so it captured nothing — reg_rdata is still its initial 0. One cycle later, the registered memory's clocked read process fires on the next rising edge (with the request now high), captures the addressed word, and one delta after drives reg_rdata = 0x11223344. The absolute timestamps (5 ns, 15 ns) depend on the clock's start phase; what matters is the relationship — the registered read lands exactly one clock period after the combinational read.

That one-cycle gap is the difference between the two contracts, and it is not cosmetic. If you build a CPU pipeline that presents a load address in cycle N and expects the data in cycle N — that is, you assumed combinational read — and you wire it to a registered-read SRAM, every load reads stale data and the CPU computes wrong answers, with no error message anywhere. If you assumed registered read and wired it to a combinational memory, you waste a pipeline stage waiting for data that was already there. The contract is a promise the requester and responder must agree on in writing.

When do you pick which? Combinational read models register files, small look-up ROMs, and asynchronous SRAM — anywhere the consumer needs the data within the same cycle and the array is small enough to read in one combinational path. Registered read models synchronous SRAM macros, on-chip memory blocks, and anything large enough that a same-cycle read would blow the timing budget — you trade one cycle of latency for a clean, registered output and a relaxed timing path. Real CPUs almost always use registered-read data memory and absorb the cycle in the pipeline; the RV32I single-cycle CPU in this series uses a combinational read precisely because it has no pipeline to absorb a cycle in. That is a deliberate simplification we will revisit when the CPU gains a pipeline in a later section.

Address decoding and the memory map

A real bus has more than one responder hanging off it. The requester drives a single address; address decode is the logic that examines the high bits of that address and selects which responder owns the transaction. The set of "which address range belongs to which device" rules is the memory map.

The example below puts two responders behind one address: a 16-byte RAM region at base 0x0000, and a single read-only 32-bit STATUS register at 0x1000. A read to anything in 0x0000–0x000F hits the RAM; a read to 0x1000 returns the status word; a read to any unmapped address returns a defined sentinel 0xBADADD12 (so that bus errors are visible rather than silently returning zero). This "decode high bits, index with low bits, default for holes" structure is exactly what every bus fabric does, scaled down.

// file: address_decode.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            address_decode.cpp -o address_decode -lsystemc
//
// One bus, two responders selected by address:
//   0x0000..0x000F : 16-byte RAM
//   0x1000         : read-only STATUS register
//   anything else  : unmapped -> sentinel 0xBADADD12

#include <systemc.h>
#include <cstdint>
#include <cstring>
#include <iostream>
#include <iomanip>

SC_MODULE(BusResponder) {
  sc_in<bool>          clk;
  sc_in<sc_uint<32>>   addr;
  sc_in<bool>          wr_en;
  sc_in<bool>          rd_en;
  sc_in<sc_uint<4>>    byte_en;
  sc_in<sc_uint<32>>   wdata;
  sc_out<sc_uint<32>>  rdata;

  static const uint32_t RAM_BASE   = 0x0000;
  static const uint32_t RAM_BYTES  = 16;
  static const uint32_t STATUS_ADDR = 0x1000;
  static const uint32_t UNMAPPED   = 0xBADADD12;

  uint8_t  ram[RAM_BYTES];
  uint32_t status;                 // read-only status register

  bool in_ram(uint32_t a) const {
    return a >= RAM_BASE && a < (RAM_BASE + RAM_BYTES);
  }

  void read_proc() {
    if (!rd_en.read()) { rdata.write(0); return; }
    uint32_t a = (uint32_t)addr.read();
    if (in_ram(a)) {
      uint32_t base = (a - RAM_BASE) & ~0x3u;
      rdata.write((uint32_t)ram[base] | ((uint32_t)ram[base+1] << 8)
                | ((uint32_t)ram[base+2] << 16) | ((uint32_t)ram[base+3] << 24));
    } else if (a == STATUS_ADDR) {
      rdata.write(status);
    } else {
      rdata.write(UNMAPPED);       // defined value for a hole in the map
    }
  }

  void write_proc() {
    if (!wr_en.read()) return;
    uint32_t a = (uint32_t)addr.read();
    if (in_ram(a)) {
      uint32_t base = (a - RAM_BASE) & ~0x3u;
      uint32_t d = (uint32_t)wdata.read();
      uint32_t be = (uint32_t)byte_en.read();
      if (be & 0x1) ram[base+0] = (uint8_t)(d & 0xFF);
      if (be & 0x2) ram[base+1] = (uint8_t)((d >> 8) & 0xFF);
      if (be & 0x4) ram[base+2] = (uint8_t)((d >> 16) & 0xFF);
      if (be & 0x8) ram[base+3] = (uint8_t)((d >> 24) & 0xFF);
    }
    // Writes to STATUS_ADDR or unmapped space are silently dropped (read-only).
  }

  SC_CTOR(BusResponder) {
    memset(ram, 0, sizeof(ram));
    status = 0x0000C0DE;           // some fixed status value
    SC_METHOD(read_proc);
    sensitive << addr << rd_en;
    SC_METHOD(write_proc);
    sensitive << clk.pos();
    dont_initialize();
  }
};

SC_MODULE(Probe) {
  sc_in<bool>          clk;
  sc_out<sc_uint<32>>  addr;
  sc_out<bool>         wr_en;
  sc_out<bool>         rd_en;
  sc_out<sc_uint<4>>   byte_en;
  sc_out<sc_uint<32>>  wdata;
  sc_in<sc_uint<32>>   rdata;

  void read_at(uint32_t a, const char* label) {
    addr.write(a); rd_en.write(true); wr_en.write(false);
    wait(SC_ZERO_TIME); wait(SC_ZERO_TIME);   // settle combinational read + commit
    std::cout << "[" << sc_time_stamp() << "] read " << label
              << " (0x" << std::hex << a << ") -> 0x"
              << std::setw(8) << std::setfill('0') << (uint32_t)rdata.read()
              << std::dec << "\n";
    rd_en.write(false);
  }

  void run() {
    addr.write(0); wr_en.write(false); rd_en.write(false); byte_en.write(0);
    wait();
    // Write a word into RAM, then read all three regions.
    addr.write(0x0000); wdata.write(0xCAFEF00D); byte_en.write(0xF);
    wr_en.write(true); rd_en.write(false);
    wait();
    wr_en.write(false);
    read_at(0x0000, "RAM");
    wait();
    read_at(0x1000, "STATUS");
    wait();
    read_at(0x2000, "UNMAPPED");
    wait();
    sc_stop();
  }

  SC_CTOR(Probe) { SC_THREAD(run); sensitive << clk.pos(); }
};

int sc_main(int, char*[]) {
  sc_clock           clk("clk", 10, SC_NS);
  sc_signal<sc_uint<32>> addr, wdata, rdata;
  sc_signal<bool>        wr_en, rd_en;
  sc_signal<sc_uint<4>>  byte_en;

  BusResponder dut("dut");
  dut.clk(clk); dut.addr(addr); dut.wr_en(wr_en); dut.rd_en(rd_en);
  dut.byte_en(byte_en); dut.wdata(wdata); dut.rdata(rdata);

  Probe p("p");
  p.clk(clk); p.addr(addr); p.wr_en(wr_en); p.rd_en(rd_en);
  p.byte_en(byte_en); p.wdata(wdata); p.rdata(rdata);

  sc_start();
  return 0;
}

Expected output:

[20 ns] read RAM (0x0) -> 0xcafef00d
[30 ns] read STATUS (0x1000) -> 0x0000c0de
[40 ns] read UNMAPPED (0x2000) -> 0xbadadd12

The decode is the three-way if/else if/else in read_proc: high-bit comparison selects the region, the low bits index within it, and the final else gives a defined value for every address that belongs to no device. That last clause matters more than it looks. A decode that returns 0 for unmapped space hides bugs — a typo in a base address reads zero, looks like an uninitialized-but-plausible value, and the real fault surfaces three modules downstream. Returning a loud sentinel (or, in a stricter model, asserting a bus-error response) turns a silent wrong-address bug into an obvious one. Real fabrics issue a DECERR/SLVERR response for exactly this reason.

Note Address decode and array indexing are different layers. Decode answers "which device does this address belong to?" by comparing high bits against the memory map. Indexing answers "which location inside that device?" by using the low bits. Conflating them — for example, masking the address to the RAM size before checking which region it is in — makes a STATUS-register read alias onto RAM. Decode first, then index.

Foreshadowing the handshake

Everything above used a fixed-latency contract: the responder is always ready, so there is no "can you take this now?" negotiation. The general case adds a ready/valid handshake — the requester raises valid, the responder raises ready, and the transfer happens only on a cycle where both are high — which lets a slow responder apply back-pressure. At RTL you would add valid/ready ports and gate the write/read on valid && ready. The transaction-level form of this same idea — where a single function call carries an entire transaction and the responder can model arbitrary latency without per-cycle signaling — is the subject of Section 3.

Advanced: Edge Cases & LRM Corners

Now the worked example. The RV32I data memory is one concrete instance of the contract from the Mental Model: a combinational-read, synchronous-write byte-addressable memory whose width-and-signedness code happens to be the instruction's funct3 field. Everything specific to RISC-V — eight access widths, byte-enable generation, alignment, sign versus zero extension — is shaping applied on top of the generic interface you already understand.

The RV32I width/sign contract

RISC-V encodes access width and signedness entirely in the 3-bit funct3 field. The decoder (covered in an earlier part) extracts it; the data memory consumes it. The same funct3 value means different things for loads versus stores, with the read/write direction supplied by separate mem_read/mem_write control bits.

funct3 Load Store Width Extension
000 LB SB byte (8-bit) sign-extend (load)
001 LH SH half (16-bit) sign-extend (load)
010 LW SW word (32-bit) none
100 LBU — byte (8-bit) zero-extend
101 LHU — half (16-bit) zero-extend

Two structural facts fall straight out of the generic interface. First, the byte-enable the store needs is a function of the access width and the byte offset within the aligned word — exactly the byte_en mask from the Beginner block, now generated from funct3 and addr[1:0] instead of being driven directly. Second, sign/zero extension is the responder shaping read data to the contract: the array holds raw bytes; the load shapes them into a 32-bit value whose upper bits follow the signedness rule.

Byte-enable generation

The store byte-enable mask is be = width_mask << addr[1:0], where the width mask is 0b0001 for a byte, 0b0011 for a halfword, and 0b1111 for a word. This is the single most error-prone line in the whole module, so look at it closely:

Access addr[1:0] Width mask be = mask << offset
SB 00 0001 0001
SB 01 0001 0010
SB 10 0001 0100
SB 11 0001 1000
SH 00 0011 0011
SH 10 0011 1100
SW 00 1111 1111

The trap is SH at offset 2: the byte enable is 1100, not 0011. An engineer who thinks "halfword means the low two bytes" hard-codes 0011 and silently corrupts bytes 0 and 1 while leaving bytes 2 and 3 — the ones the program actually wanted to write — untouched. The shift-by-offset formula gets it right for every legal offset.

The RV32I data memory, complete

Below is a self-contained, runnable RV32I data memory using the combinational-read / synchronous-write contract. The read process shapes load data per funct3; the write process generates byte enables per funct3 and addr[1:0] and writes only the enabled lanes. An alignment check warns (it does not abort) on a misaligned halfword or word — the base ISA requires natural alignment, and a real core would raise an exception, which we foreshadow with a warning since exception plumbing belongs to the control unit.

// file: rv32i_dmem.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            rv32i_dmem.cpp -o rv32i_dmem -lsystemc
// Run:   LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./rv32i_dmem
//
// RV32I data memory: combinational load, synchronous store.
//   Loads : LB(000) LH(001) LW(010) LBU(100) LHU(101)  -- sign/zero extension
//   Stores: SB(000) SH(001) SW(010)                    -- byte-enable generation
// Byte-addressable, little-endian. SystemC 2.3.x.

#include <systemc.h>
#include <cstdint>
#include <cstring>
#include <iostream>
#include <iomanip>

enum funct3_t : uint32_t {
  F3_LB  = 0x0,  // also SB
  F3_LH  = 0x1,  // also SH
  F3_LW  = 0x2,  // also SW
  F3_LBU = 0x4,
  F3_LHU = 0x5,
};

SC_MODULE(DataMemory) {
  sc_in<bool>          clk;
  sc_in<sc_uint<32>>   addr;       // byte address
  sc_in<sc_uint<32>>   wdata;      // store data
  sc_out<sc_uint<32>>  rdata;      // load data (responder-driven)
  sc_in<bool>          mem_read;   // perform a load this cycle
  sc_in<bool>          mem_write;  // perform a store this cycle
  sc_in<sc_uint<3>>    funct3;     // width + sign selector

  static const int MEM_BYTES = 4096;
  uint8_t mem[MEM_BYTES];

  // Warn (do not abort) on a misaligned access.
  void check_align(uint32_t a, uint32_t width, const char* op) const {
    if ((a & (width - 1)) != 0)
      std::cerr << "[dmem] ALIGN WARNING: " << op << " at 0x"
                << std::hex << a << " not " << std::dec << width
                << "-byte aligned\n";
  }

  // --- Combinational load: shape raw bytes per funct3 ---
  void read_proc() {
    if (!mem_read.read()) { rdata.write(0); return; }
    uint32_t a = (uint32_t)addr.read();
    uint32_t f = (uint32_t)funct3.read();
    uint32_t result = 0;
    switch (f) {
      case F3_LB: {                                  // sign-extend byte
        uint8_t raw = mem[a];
        result = (uint32_t)(int32_t)(int8_t)raw;     // int8_t -> int32_t sign-extends
        break;
      }
      case F3_LH: {                                  // sign-extend halfword
        check_align(a, 2, "LH");
        uint16_t raw = (uint16_t)mem[a] | ((uint16_t)mem[a+1] << 8);
        result = (uint32_t)(int32_t)(int16_t)raw;
        break;
      }
      case F3_LW: {                                  // full word, no extension
        check_align(a, 4, "LW");
        result = (uint32_t)mem[a] | ((uint32_t)mem[a+1] << 8)
               | ((uint32_t)mem[a+2] << 16) | ((uint32_t)mem[a+3] << 24);
        break;
      }
      case F3_LBU: {                                 // zero-extend byte
        result = (uint32_t)mem[a];                   // uint8_t -> uint32_t zero-extends
        break;
      }
      case F3_LHU: {                                 // zero-extend halfword
        check_align(a, 2, "LHU");
        result = (uint32_t)mem[a] | ((uint32_t)mem[a+1] << 8);
        break;
      }
      default:
        result = 0;
        break;
    }
    rdata.write(result);
  }

  // --- Synchronous store: generate byte enables, write enabled lanes ---
  void write_proc() {
    if (!mem_write.read()) return;
    uint32_t a    = (uint32_t)addr.read();
    uint32_t data = (uint32_t)wdata.read();
    uint32_t f    = (uint32_t)funct3.read();

    uint32_t offset = a & 0x3;                 // addr[1:0]
    uint32_t width_mask;                        // 0b0001 / 0b0011 / 0b1111
    uint32_t width;
    switch (f) {
      case F3_LB: width_mask = 0x1; width = 1; break;            // SB
      case F3_LH: width_mask = 0x3; width = 2; check_align(a,2,"SH"); break; // SH
      case F3_LW: width_mask = 0xF; width = 4; check_align(a,4,"SW"); break; // SW
      default:    width_mask = 0x0; width = 0; break;
    }
    uint32_t be = (width_mask << offset) & 0xF;   // byte-enable mask

    uint32_t word_base = a & ~0x3u;               // aligned word base
    // Lane the data so byte k of the word carries data byte (k - offset).
    for (uint32_t lane = 0; lane < 4; ++lane) {
      if (be & (1u << lane)) {
        uint32_t data_byte_index = lane - offset; // which source byte feeds this lane
        mem[word_base + lane] = (uint8_t)((data >> (8 * data_byte_index)) & 0xFF);
      }
    }
    (void)width; // width is the alignment-check width; kept for clarity
  }

  SC_CTOR(DataMemory) {
    memset(mem, 0, sizeof(mem));
    SC_METHOD(read_proc);
    sensitive << addr << mem_read << funct3;
    SC_METHOD(write_proc);
    sensitive << clk.pos();
    dont_initialize();
  }
};

static int passes = 0, fails = 0;
static void check(const char* name, uint32_t got, uint32_t exp) {
  if (got == exp) { std::cout << "  PASS  " << name << " = 0x" << std::hex
                              << std::setw(8) << std::setfill('0') << got
                              << std::dec << "\n"; passes++; }
  else { std::cout << "  FAIL  " << name << " got=0x" << std::hex
                   << std::setw(8) << std::setfill('0') << got << " exp=0x"
                   << std::setw(8) << std::setfill('0') << exp << std::dec
                   << "\n"; fails++; }
}

SC_MODULE(TB) {
  sc_clock           clk{"clk", 10, SC_NS};
  sc_signal<sc_uint<32>> addr, wdata, rdata;
  sc_signal<bool>        mem_read, mem_write;
  sc_signal<sc_uint<3>>  funct3;

  DataMemory dut{"dut"};

  void store(uint32_t a, uint32_t d, uint32_t f) {
    addr.write(a); wdata.write(d); funct3.write(f);
    mem_write.write(true); mem_read.write(false);
    wait(clk.posedge_event()); wait(SC_ZERO_TIME);
    mem_write.write(false);
  }
  uint32_t load(uint32_t a, uint32_t f) {
    addr.write(a); funct3.write(f);
    mem_read.write(true); mem_write.write(false);
    wait(SC_ZERO_TIME); wait(SC_ZERO_TIME);   // settle read + commit buffered rdata
    uint32_t r = (uint32_t)rdata.read();
    mem_read.write(false);
    return r;
  }

  void run() {
    mem_read.write(false); mem_write.write(false);
    addr.write(0); wdata.write(0); funct3.write(0);
    wait(clk.posedge_event());

    std::cout << "\n=== SW / LW (word) ===\n";
    store(0x00, 0xDEADBEEF, F3_LW);
    check("LW @0x00", load(0x00, F3_LW), 0xDEADBEEF);

    std::cout << "\n=== SB at all four byte lanes ===\n";
    store(0x10, 0x00000000, F3_LW);    // clear the word
    store(0x10, 0xAA, F3_LB);          // lane 0
    store(0x11, 0xBB, F3_LB);          // lane 1
    store(0x12, 0xCC, F3_LB);          // lane 2
    store(0x13, 0xDD, F3_LB);          // lane 3
    check("LW @0x10 after 4x SB", load(0x10, F3_LW), 0xDDCCBBAA);

    std::cout << "\n=== SH at offset 0 and offset 2 (byte-enable corner) ===\n";
    store(0x20, 0x00000000, F3_LW);
    store(0x20, 0xABCD, F3_LH);        // low halfword,  be = 0011
    store(0x22, 0x1234, F3_LH);        // high halfword, be = 1100  (NOT 0011)
    check("LW @0x20 after SH lo+hi", load(0x20, F3_LW), 0x1234ABCD);

    std::cout << "\n=== Sign vs zero extension (LB/LBU, LH/LHU) ===\n";
    store(0x30, 0x80, F3_LB);          // store byte 0x80
    check("LB  0x80 (sign)", load(0x30, F3_LB),  0xFFFFFF80);
    check("LBU 0x80 (zero)", load(0x30, F3_LBU), 0x00000080);
    store(0x34, 0x8000, F3_LH);        // store halfword 0x8000
    check("LH  0x8000 (sign)", load(0x34, F3_LH),  0xFFFF8000);
    check("LHU 0x8000 (zero)", load(0x34, F3_LHU), 0x00008000);
    store(0x38, 0x7F, F3_LB);          // positive byte must NOT sign-extend
    check("LB  0x7F (positive)", load(0x38, F3_LB), 0x0000007F);

    std::cout << "\n=== Partial overwrite (SB inside a word) ===\n";
    store(0x40, 0x11223344, F3_LW);
    store(0x42, 0xFF, F3_LB);          // overwrite only lane 2
    check("LW @0x40 after SB lane2", load(0x40, F3_LW), 0x11FF3344);

    std::cout << "\n=== Misaligned LH (expect ALIGN WARNING, no crash) ===\n";
    store(0x50, 0xAABBCCDD, F3_LW);
    (void)load(0x51, F3_LH);           // misaligned; prints a warning
    std::cout << "  [continued past misaligned access]\n";
    passes++;

    std::cout << "\n========================================\n";
    std::cout << "  PASS: " << passes << "   FAIL: " << fails << "\n";
    std::cout << "========================================\n";
    sc_stop();
  }

  SC_CTOR(TB) {
    dut.clk(clk); dut.addr(addr); dut.wdata(wdata); dut.rdata(rdata);
    dut.mem_read(mem_read); dut.mem_write(mem_write); dut.funct3(funct3);
    SC_THREAD(run);
  }
};

int sc_main(int, char*[]) {
  TB tb{"tb"};
  sc_start();
  return 0;
}

Expected output:

=== SW / LW (word) ===
  PASS  LW @0x00 = 0xdeadbeef

=== SB at all four byte lanes ===
  PASS  LW @0x10 after 4x SB = 0xddccbbaa

=== SH at offset 0 and offset 2 (byte-enable corner) ===
  PASS  LW @0x20 after SH lo+hi = 0x1234abcd

=== Sign vs zero extension (LB/LBU, LH/LHU) ===
  PASS  LB  0x80 (sign) = 0xffffff80
  PASS  LBU 0x80 (zero) = 0x00000080
  PASS  LH  0x8000 (sign) = 0xffff8000
  PASS  LHU 0x8000 (zero) = 0x00008000
  PASS  LB  0x7F (positive) = 0x0000007f

=== Partial overwrite (SB inside a word) ===
  PASS  LW @0x40 after SB lane2 = 0x11ff3344

=== Misaligned LH (expect ALIGN WARNING, no crash) ===
[dmem] ALIGN WARNING: LH at 0x51 not 2-byte aligned
  [continued past misaligned access]

========================================
  PASS: 10   FAIL: 0
========================================

(The ALIGN WARNING line is written to stderr; the rest is stdout. Depending on how your terminal interleaves the two streams it may appear at a slightly different position, but it always prints exactly once, for the one misaligned LH.)

Reading the load/store trace

Three results in that trace are worth dwelling on because each one is a bug magnet.

The four-SB word. Storing 0xAA, 0xBB, 0xCC, 0xDD to byte addresses 0x10, 0x11, 0x12, 0x13 and then reading the word back yields 0xDDCCBBAA, not 0xAABBCCDD. This is little-endian laid bare: the byte at the lowest address (0xAA at 0x10) is the least-significant byte of the word, so it sits in bits [7:0]. The byte-enable generation placed each store in its correct lane: SB at offset 0 set be = 0001, at offset 1 be = 0010, at offset 2 be = 0100, at offset 3 be = 1000, and the data-lane shift fed each store byte to the matching lane.

The SH-at-offset-2 corner. Storing 0xABCD at offset 0 (be = 0011) and 0x1234 at offset 2 (be = 1100) assembles 0x1234ABCD. Had the byte enable for the offset-2 store been hard-coded to 0011, the high store would have clobbered 0xABCD and left the high half of the word zero — a classic, hard-to-spot corruption. The width_mask << offset formula produced 1100 and routed 0x1234 into bytes 2 and 3.

Sign versus zero extension. The byte 0x80 loaded as LB becomes 0xFFFFFF80 (sign-extended: bit 7 is set, so the upper 24 bits fill with ones), but loaded as LBU becomes 0x00000080 (zero-extended). The bytes in memory are identical; only the load's signedness rule differs. The C++ cast chain (int32_t)(int8_t)raw does the sign extension because the standard guarantees that widening a signed type sign-extends; (uint32_t)raw zero-extends because widening an unsigned type fills with zeros. Get the intermediate cast wrong — write (int32_t)(uint8_t)raw for LB — and you zero-extend a signed load: correct for 0x00–0x7F, wrong for 0x80–0xFF, and invisible in any test whose data happens to be positive.

LRM corner: the array write is not a signal write

One subtlety from IEEE 1666-2011 §6.4 underpins the whole module. The store does mem[a] = ..., a plain C++ array assignment, which takes effect immediately — there is no request-update buffering, no delta delay. The sc_signal ports (addr, rdata, …) do buffer and commit in the update phase. This split is exactly what you want: the storage mutates synchronously on the clock edge inside the write process, and the combinational read sees the new bytes the next time it evaluates, while the ports still obey signal timing. If you mistakenly modeled each memory byte as an sc_signal, every store would incur an extra delta of latency and the model would be both slower and harder to reason about. Keep storage in a C++ array; keep only the interface in signals.

LRM corner: combinational read needs every input in its sensitivity list

The read process is sensitive << addr << mem_read << funct3 — all three of its inputs. Drop funct3 and a load that changes only the width code (say LB then LBU at the same address) will not re-evaluate, so rdata keeps the previous width's value and the sign/zero distinction silently fails. SystemC has no always @(*); the list is manual and must be complete. This is the same discipline the FSM post hammered on, applied to a memory's read port.

Hands-on exercise

Build a memory-mapped GPIO peripheral and exercise it through the generic contract, then layer the RV32I shaping rules onto a second memory.

Part A — the peripheral. Model a GPIO block with three 32-bit registers behind a bus interface (addr, wr_en, rd_en, byte_en, wdata, rdata), combinational read, synchronous write:

  • DIR at offset 0x00 — read/write direction register (1 = output, 0 = input).
  • OUT at offset 0x04 — read/write output data register.
  • IN at offset 0x08 — read-only input data register; writes to it are silently dropped.

Decode the offset to pick the register, honor byte_en on writes to DIR and OUT, and return a sentinel (0xDEADC0DE) for any offset outside 0x00–0x08. Drive it with a probe that writes DIR = 0x000000FF, writes OUT = 0x000000A5, reads both back, attempts a write to IN (verify it does not change), and reads an unmapped offset (verify the sentinel). Predict every printed value before you run.

Part B — registered-read variant. Re-expose the same three registers behind a registered-read port (clocked read process, dont_initialize()). Drive the identical stimulus and watch every read appear one cycle later than in Part A. Confirm that a requester that reads on the same cycle it presents the address now sees stale data — this is the contract mismatch from the Intermediate section, reproduced by your own hand.

Part C — RV32I shaping on top. Take the DataMemory module from the Advanced section and add a tenth check to its testbench: store 0xFFFFFFFF as a word at 0x60, then load byte 0 as LB (expect 0xFFFFFFFF) and as LBU (expect 0x000000FF), and load the low halfword as LH (expect 0xFFFFFFFF) and as LHU (expect 0x0000FFFF). Then store a halfword 0xBEEF at offset 2 of a cleared word at 0x64 and confirm the word reads back 0xBEEF0000, proving your byte enable was 1100. No solution is provided.

Hints

  • Reuse the read_proc / write_proc two-process split from the Beginner block verbatim; only the decode and the register set change.
  • For the registered-read variant, the only change is moving the read process's sensitivity from addr/rd_en to clk.pos() and adding dont_initialize(). The body is identical. That one edit is the entire combinational-to-registered conversion.
  • Generate the GPIO byte enables exactly as the RV32I store does: be = width_mask << (addr & 0x3). Even though the registers are word-aligned, writing the mask out makes the sub-word write path explicit.
  • For Part C, remember that LB of 0xFF is 0xFFFFFFFF (sign bit set) but LBU of the same byte is 0x000000FF. If they come out equal, your intermediate cast is wrong.
  • Drive stimulus from an SC_THREAD that wait()s on clk.pos(); avoid driving inputs at non-clock-edge times unless you specifically want to study glitching.

Common mistakes

  • Hard-coding the halfword byte enable to 0011. A store-halfword at byte offset 2 needs be = 1100, not 0011. The correct formula is be = 0b0011 << addr[1:0], which yields 1100 at offset 2. Hard-coding 0011 writes the wrong two bytes — it corrupts bytes 0 and 1 and leaves the intended bytes 2 and 3 untouched. The bug is invisible whenever halfwords happen to be word-aligned (offset 0), so it survives casual testing and only bites on offset-2 stores. Fix: always shift the width mask by the byte offset.
  • Sign-extending an unsigned load. LBU/LHU zero-extend; LB/LH sign-extend. The C++ idiom for sign extension is (int32_t)(int8_t)raw (signed intermediate); for zero extension it is (uint32_t)raw (unsigned widening). Writing (int32_t)(int8_t) for LBU, or forgetting the int8_t intermediate for LB, produces results that are correct for values 0x00–0x7F and wrong for 0x80–0xFF. Because most test data is small and positive, the bug hides until a value with the top bit set shows up. Fix: match the cast to the instruction's signedness, and always test a value with the high bit set (0x80, 0xFF, 0x8000).
  • Assuming combinational read when the memory is registered (or vice versa). A combinational-read interface returns data the same cycle the address is presented; a registered-read interface returns it one cycle later. Wire a requester built for one contract to a responder honoring the other and every read is off by a cycle — stale data on a combinational consumer of a registered memory, or a wasted cycle the other way. There is no error message. Fix: write the read-latency into the interface contract and make both sides obey it; when in doubt, add an assertion that the consumer samples data on the cycle the contract specifies.
  • Letting an unaligned access silently span two words. A word load at byte address 0x1001 is not naturally aligned; in real hardware it either faults (base RV32I) or is split into two bus beats. A model that indexes mem[a..a+3] from an unaligned a reads across a word boundary and returns a value no real aligned-only core would ever produce, masking a program bug. Fix: check alignment ((a & (width-1)) == 0) and either warn, fault, or explicitly model the split — but never let it pass silently as if it were a normal access.
  • Modeling every memory byte as an sc_signal. The storage array should be a plain C++ array (uint8_t mem[]); only the ports are signals. Wrapping each byte in an sc_signal adds a delta of latency to every store, slows simulation by orders of magnitude on a large memory, and buys nothing — the C++ array already gives correct synchronous-write/combinational-read behavior when the write happens inside a clocked process. Fix: storage in a C++ array, interface in signals.
  • Returning 0 for unmapped addresses. A decode whose default clause returns 0 hides addressing bugs: a wrong base address reads a plausible-looking zero, and the real fault surfaces far downstream. Fix: return a loud sentinel or assert a bus-error response so a hole in the memory map is immediately visible.

Recap

After working through this post you can now:

  • State a memory interface as a contract — a bundle of signals (addr, direction, byte_en, wdata, rdata) plus a timing agreement — and say who drives each signal.
  • Explain the difference between a combinational-read and a registered-read interface, predict the exact cycle read data is valid in each, and name the bug that results from mismatching the two.
  • Decode an address against a memory map: compare high bits to select a responder, index with low bits, and return a defined value for unmapped holes.
  • Generate a correct byte-enable mask for any sub-word access at any offset using be = width_mask << addr[1:0], and explain why SH at offset 2 needs 1100.
  • Implement the RV32I data memory's eight load/store widths, including little-endian byte assembly and the sign-versus-zero extension rule, as a worked instance of the generic contract.
  • Diagnose the classic memory-interface bugs on sight: the wrong halfword byte enable, the sign-extended unsigned load, the silent unaligned access, the over-modeled sc_signal storage, and the silent unmapped read.
  • Recognize where a ready/valid handshake would slot in to add back-pressure, and why the transaction-level form of that idea is deferred to Section 3.

Further reading

Standards

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §4.2.1.3 (update phase and delta cycles), §5.2.16 (SC_METHOD semantics), §5.2.18 (sensitivity lists and dont_initialize), §6.4 (sc_signal request-update write semantics).
  • The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA, §2.6 — load and store instructions, natural-alignment requirement, little-endian default for the base ISA.

Vendor and consortium documents

  • ARM, AMBA AHB-Lite Protocol Specification (IHI0033) — HSIZE/HADDR width-and-alignment contract and byte-strobe semantics, the industry analogue of the funct3/byte_en interface modeled here.
  • ARM, AMBA AXI Protocol Specification (AXI4) — WSTRB write-strobe (byte-enable) lane model.
  • Accellera Systems Initiative, SystemC 2.3.x User Guide — channel and process mechanics relevant to memory modeling.

Textbooks

  • Bhasker, A SystemC Primer (2nd ed.), ch. 6 — modeling memories, ports, and channels.
  • Grötker, Liao, Martin, and Swan, System Design with SystemC, ch. 4–5 — the request-update model and channel design.

Real-world references

  • Ibex ibex_load_store_unit.sv (public source, lowRISC) — production RV32I load/store unit: byte-enable generation and load data extension in synthesizable RTL.
  • PicoRV32 (public source) — compact RV32I core whose memory interface shows the same width/alignment handling at the leaf level.

Next in this section

→ Part 7: Structural Composition & Hierarchy — composing the register file, decoder, PC, instruction memory, and the data memory built here into a working single-cycle RV32I datapath plus control unit, with correct port binding, sensitivity, and reset across a module hierarchy. The memory interface from this post becomes one block in that top-level composition.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment