23. SystemC Tutorial - Capstone: Build a Virtual Platform

Why this matters

Written 2026-06-05 for the concept-first series.

You've learned the pieces. Now build the platform.

Across the previous seven parts of this section you learned every primitive that a transaction-level virtual platform is made of — one at a time, each from first principles, each compiled and run against a real SystemC install. You learned why transaction-level modeling exists at all: that simulating a system signal-by-signal, delta-cycle by delta-cycle, does not scale to running real software, and that trading pin-and-cycle accuracy for transactions buys you orders of magnitude of speed. You learned the tlm_generic_payload field by field and the b_transport call that carries it, including the response-status discipline that separates correct models from silently broken ones. You learned what sockets are and what the convenience sockets wrap. You learned the non-blocking, approximately-timed phase protocol for when loosely-timed is not enough. You learned the quantum keeper and temporal decoupling — how a model runs ahead of simulated time to go fast. You built a bus that routes a transaction by decoding its address. And you built a memory-mapped peripheral with real register side effects, plus the debug-transport backdoor that inspects state without disturbing it.

Every one of those posts ended the same way: here is the worked example, here is the trace, here is the bug to avoid. This one is different. There is no worked solution here. The capstone hands you a runnable virtual platform — a CPU wrapper, a bus, a RAM, and a UART, all wired and all compiling — with one deliberate gap: the glue that turns the processor's load/store port into TLM transactions. Three exercises of increasing difficulty ask you to fill that gap, make the platform faster, and re-architect how its program is loaded. This is exactly how a real modeling team onboards an engineer onto a virtual platform: the platform is handed to you complete, and wiring the device-under-model into it — and growing it — is the job. The single most valuable thing you can do at the end of a tutorial section is close the book and rebuild the result from memory, discovering for yourself which details were load-bearing. The pieces are yours. Assemble them.

Prerequisites

This capstone assumes you have worked through — not skimmed — the seven TLM posts that precede it in this section, plus the RV32I CPU you built in Section 2. Each link below is a building block you will reach for directly while solving the exercises:

If any of these is hazy, re-read it before starting. The exercises will not re-teach the primitives; they assume you can reach for each one.

What you have at this point

Concretely, here is the inventory of finished ideas you carry into this capstone. Each is something you built and ran in an earlier part; the capstone composes them into one platform.

  • The generic payload and b_transport (Part 2) — a transaction is a function call carrying a struct by reference; the initiator owns the request fields, the target sets the response status and (for a read) the data bytes; latency is annotated onto the sc_time& delay argument, not consumed by a clock.
  • Sockets (Part 3) — simple_initiator_socket and simple_target_socket from tlm_utils, plus the multi_passthrough_initiator_socket a bus uses to fan out to many targets.
  • Approximately-timed transport (Part 4) — the non-blocking phase protocol, in your back pocket for when LT is not enough.
  • Temporal decoupling (Part 5) — the quantum keeper that lets an initiator run ahead of simulated time and synchronise only occasionally, the single biggest LT speed lever.
  • The bus fabric (Part 6) — an address-indexed router: decode the address against a memory map, localize (subtract the base), forward the same payload, restore the global address on the way back.
  • The memory-mapped peripheral and debug transport (Part 7) — a target whose b_transport decodes an offset to a register and runs that register's side-effecting behavior, plus the transport_dbg backdoor that observes state with side effects off and zero time consumed.
  • The RV32I CPU from Section 2 — a processor that fetches, decodes, and executes the base integer instruction set, with a load/store data-memory port that, until now, talked to a signal-level data memory.

That is a complete parts bin for a virtual platform. The capstone asks you to assemble it — and to be the one who connects the CPU's data port to the bus.

What you're going to build

The deliverable is a small virtual platform: an RV32I CPU, wrapped as a TLM-2.0 initiator, talking through a bus to two targets — a RAM model and a memory-mapped UART — and running a toy program that prints HELLO to your console by storing bytes to the UART's DATA register. When it works, you will have watched a processor model execute software whose only output reaches the world as transactions routed across a fabric to a peripheral. That is, in miniature, exactly what every commercial virtual platform does when it boots an operating system: software runs on a CPU model, and every device interaction is a transaction.

Here is the topology. The CPU wrapper owns a behavioural RV32I core and exposes a single TLM initiator socket. Every time the core executes a load or a store, it does not poke a signal-level memory; it calls out to the wrapper, which issues a b_transport transaction onto the bus. The bus decodes the transaction's address against a two-entry memory map — RAM at 0x0000_0000, UART at 0x1000_0000, the section's standard map — localizes the address, forwards the same payload to the matching target, and restores the address on the return path. The RAM is a plain memory target (Part 2's pattern, with a transport_dbg backdoor added for Exercise 3). The UART is the peripheral from Part 7: a write to its DATA register emits the low byte to std::cout; a read of its STATUS register returns a tx-ready flag.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b'}}}%%
flowchart LR
    CORE["RV32I core"] -->|load/store callback| WRAP["CpuWrapper
(initiator socket)"] WRAP -->|b_transport| BUS["Bus
(address decode)"] BUS -->|0x0000_0000| RAM["Ram
(memory target)"] BUS -->|0x1000_0000| UART["Uart
(peripheral)"] UART -.->|DATA write| OUT["console: HELLO"]

The toy program is a hand-assembled sequence of RV32I instructions placed in the core's instruction ROM. It loads the UART base address into a register, then for each character of "HELLO" it materialises the character in a second register and stores it to the UART DATA register, and finally halts on ebreak:

lui   x1, 0x10000        # x1 = UART base 0x1000_0000
addi  x2, x0, 'H'        # x2 = 0x48
sw    x2, 0(x1)          # store -> UART DATA (emits 'H')
addi  x2, x0, 'E'        # ... repeat for E, L, L, O
sw    x2, 0(x1)
...
ebreak                   # halt

Each sw is where the magic happens: it reaches the core's data-memory port, which — once you have wired it — becomes a TLM write transaction. The bus decodes 0x1000_0000 to the UART, localizes it to offset 0x0, and the UART's DATA-write side effect prints the character. Five stores, five characters, then the CPU halts.

You are not building the bus, the RAM, the UART, or the core. Those are provided below, complete and compiling. You are building the one connection that turns the core's load/store port into bus traffic — the CpuWrapper::mem_access glue — and then, in the later exercises, making it faster and changing how the program loads.

Scaffold provided

Below is the complete scaffold — one self-contained C++17 file. Everything in it compiles and runs as given. The RV32I core, the bus, the RAM, and the UART are complete and you should not need to touch them. The one deliberate gap is the body of CpuWrapper::mem_access, marked // EXERCISE 1: the function that should translate the core's load/store callbacks into TLM transactions. As shipped, that stub issues no transaction — so the platform boots, the CPU runs, decodes the toy program, reaches its first store, and fails loudly at exactly the data port that is not yet wired. That is intentional: a runnable scaffold that fails its own job at a single, clearly named point is the starting line, not the finish.

Note Build with a C++17 compiler against your SystemC install. On the reference machine for this post (SystemC 3.0.x on Apple silicon): g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API -I/usr/local/systemc-3.0.0/include vp.cpp -L/usr/local/systemc-3.0.0/lib-macosarm64 -lsystemc -o vp, then run with DYLD_LIBRARY_PATH=/usr/local/systemc-3.0.0/lib-macosarm64 ./vp. Substitute your own SystemC path and library-directory architecture suffix (lib-linux64, lib-macosarm64, etc.); on Linux the loader variable is LD_LIBRARY_PATH. The -DSC_ALLOW_DEPRECATED_IEEE_API flag silences 3.0.x deprecation warnings for the classic API the examples use.
// file: vp.cpp
// Capstone virtual platform scaffold: RV32I CPU (Section 2) re-fronted with a
// TLM-2.0 initiator socket -> bus -> RAM + UART peripheral.
//
// Build:
//   g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
//       -I/usr/local/systemc-3.0.0/include vp.cpp \
//       -L/usr/local/systemc-3.0.0/lib-macosarm64 -lsystemc -o vp
// Run:
//   DYLD_LIBRARY_PATH=/usr/local/systemc-3.0.0/lib-macosarm64 ./vp
//
// As shipped this COMPILES AND RUNS, but the CPU-wrapper's data-side TLM glue
// (CpuWrapper::mem_access, search for "EXERCISE 1") is stubbed. The platform
// boots, the CPU fetches and decodes, and then FAILS LOUDLY at the first store
// to the UART, because the stub never issues a transaction and reports the gap.
// Turning that loud failure into a printed "HELLO" is Exercise 1.

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <tlm_utils/multi_passthrough_initiator_socket.h>
#include <iostream>
#include <iomanip>
#include <sstream>
#include <functional>
#include <cstdint>

// --- Memory map (the section convention) -------------------------------------
static const sc_dt::uint64 RAM_BASE  = 0x00000000ull;
static const sc_dt::uint64 RAM_SIZE  = 0x00010000ull;   // 64 KB
static const sc_dt::uint64 UART_BASE = 0x10000000ull;
static const sc_dt::uint64 UART_SIZE = 0x00001000ull;   //  4 KB

// --- ALU operation encoding (Section 2 Part 1) -------------------------------
enum alu_op_t {
  ALU_ADD = 0, ALU_SUB = 1, ALU_AND = 2, ALU_OR  = 3, ALU_XOR = 4,
  ALU_SLT = 5, ALU_SLTU = 6, ALU_SLL = 7, ALU_SRL = 8, ALU_SRA = 9
};

// --- RV32I single-cycle CPU core (Section 2 capstone, datapath complete) -----
// This is a behavioural model of the processor you built in Section 2, collapsed
// into one SC_THREAD so the capstone can focus on the *system* around it. It
// fetches from a private instruction ROM, executes one instruction per step,
// and reaches OUT of the core for every load/store through two callbacks the
// wrapper supplies: load_cb(addr) and store_cb(addr, data). Those callbacks are
// the data-memory port we are about to re-front with TLM.
struct Rv32iCore {
  uint32_t regs[32] = {};
  uint32_t pc       = 0;
  uint32_t imem[256] = {};            // instruction ROM, word-indexed
  bool     halted   = false;

  // Data-memory port: the wrapper installs these. A null callback is a bug.
  std::function<uint32_t(uint32_t)>          load_cb;
  std::function<void(uint32_t, uint32_t)>    store_cb;

  uint32_t rd_reg(int r) const { return r == 0 ? 0u : regs[r]; }
  void     wr_reg(int r, uint32_t v) { if (r != 0) regs[r] = v; }

  // Sign-extend a 12-bit immediate.
  static int32_t sext12(uint32_t imm12) {
    return (int32_t)(imm12 << 20) >> 20;
  }

  // Execute exactly one instruction. Returns false once halted.
  bool step() {
    if (halted) return false;
    uint32_t i      = imem[(pc >> 2) & 0xFF];
    uint32_t opcode = i & 0x7F;
    uint32_t rd     = (i >> 7)  & 0x1F;
    uint32_t f3     = (i >> 12) & 0x07;
    uint32_t rs1    = (i >> 15) & 0x1F;
    uint32_t rs2    = (i >> 20) & 0x1F;
    uint32_t f7     = (i >> 25) & 0x7F;
    uint32_t next   = pc + 4;

    switch (opcode) {
      case 0x33: {                                  // R-type ADD/SUB/AND/OR
        uint32_t a = rd_reg(rs1), b = rd_reg(rs2), res = 0;
        if      (f3 == 0) res = (f7 & 0x20) ? (a - b) : (a + b);
        else if (f3 == 7) res = a & b;
        else if (f3 == 6) res = a | b;
        wr_reg(rd, res);
        break;
      }
      case 0x13: {                                  // I-type ADDI
        int32_t imm = sext12((i >> 20) & 0xFFF);
        if (f3 == 0) wr_reg(rd, rd_reg(rs1) + imm);
        break;
      }
      case 0x37:                                    // LUI
        wr_reg(rd, i & 0xFFFFF000);
        break;
      case 0x23: {                                  // S-type SW (store word)
        int32_t imm = sext12(((i >> 25) << 5) | ((i >> 7) & 0x1F));
        uint32_t addr = rd_reg(rs1) + imm;
        if (!store_cb) { std::cerr << "core: null store_cb\n"; sc_stop(); return false; }
        store_cb(addr, rd_reg(rs2));                // <- data-memory port (store)
        break;
      }
      case 0x03: {                                  // I-type LW (load word)
        int32_t imm = sext12((i >> 20) & 0xFFF);
        uint32_t addr = rd_reg(rs1) + imm;
        if (!load_cb) { std::cerr << "core: null load_cb\n"; sc_stop(); return false; }
        wr_reg(rd, load_cb(addr));                  // <- data-memory port (load)
        break;
      }
      case 0x63: {                                  // B-type BNE (branch if !=)
        int32_t imm = sext12(((i >> 31) << 12) | (((i >> 7) & 1) << 11)
                            | (((i >> 25) & 0x3F) << 5) | (((i >> 8) & 0xF) << 1));
        if (f3 == 1 && rd_reg(rs1) != rd_reg(rs2)) next = pc + imm;
        break;
      }
      case 0x73:                                    // SYSTEM: EBREAK halts
        if (((i >> 20) & 0xFFF) == 1) { halted = true; return false; }
        break;
      default: break;                               // unknown decodes as NOP
    }
    pc = next;
    return true;
  }
};

// --- RAM: a plain TLM memory target (Part 2 pattern) -------------------------
SC_MODULE(Ram) {
  tlm_utils::simple_target_socket<Ram> socket;
  unsigned char mem[RAM_SIZE];

  SC_CTOR(Ram) : socket("socket") {
    for (sc_dt::uint64 i = 0; i < RAM_SIZE; ++i) mem[i] = 0;
    socket.register_b_transport(this, &Ram::b_transport);
    socket.register_transport_dbg(this, &Ram::transport_dbg);
  }

  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    sc_dt::uint64  addr = trans.get_address();
    unsigned int   len  = trans.get_data_length();
    unsigned char* ptr  = trans.get_data_ptr();
    if (addr + len > RAM_SIZE) {
      trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE); return;
    }
    if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
      for (unsigned i = 0; i < len; ++i) mem[addr + i] = ptr[i];
    else
      for (unsigned i = 0; i < len; ++i) ptr[i] = mem[addr + i];
    delay += sc_time(len, SC_NS);
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }

  // Backdoor: no side effects, no time. Used by the debug loader in Exercise 3.
  unsigned int transport_dbg(tlm::tlm_generic_payload& trans) {
    sc_dt::uint64  addr = trans.get_address();
    unsigned int   len  = trans.get_data_length();
    unsigned char* ptr  = trans.get_data_ptr();
    if (addr + len > RAM_SIZE) len = (addr < RAM_SIZE) ? (unsigned)(RAM_SIZE - addr) : 0;
    if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
      for (unsigned i = 0; i < len; ++i) mem[addr + i] = ptr[i];
    else
      for (unsigned i = 0; i < len; ++i) ptr[i] = mem[addr + i];
    return len;
  }
};

// --- UART: a memory-mapped peripheral (Part 7 pattern) -----------------------
// DATA(0x0): write emits the low byte; STATUS(0x4): read returns tx-ready=1.
SC_MODULE(Uart) {
  tlm_utils::simple_target_socket<Uart> socket;

  static const sc_dt::uint64 REG_DATA   = 0x0;
  static const sc_dt::uint64 REG_STATUS = 0x4;

  SC_CTOR(Uart) : socket("socket") {
    socket.register_b_transport(this, &Uart::b_transport);
  }

  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    sc_dt::uint64  off = trans.get_address();        // bus localizes for us
    unsigned int   len = trans.get_data_length();
    unsigned char* ptr = trans.get_data_ptr();
    if (len != 4) { trans.set_response_status(tlm::TLM_BURST_ERROR_RESPONSE); return; }
    uint32_t v = 0;
    if (trans.get_command() == tlm::TLM_WRITE_COMMAND) {
      v = ptr[0] | (ptr[1] << 8) | (ptr[2] << 16) | (ptr[3] << 24);
      if (off == REG_DATA) { std::cout << (char)(v & 0xFF) << std::flush; }
      // STATUS ignores writes.
    } else {
      if (off == REG_STATUS) v = 0x1;                // tx-ready
      else                   v = 0;
      ptr[0] = v & 0xFF; ptr[1] = (v >> 8) & 0xFF;
      ptr[2] = (v >> 16) & 0xFF; ptr[3] = (v >> 24) & 0xFF;
    }
    delay += sc_time(5, SC_NS);
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

// --- Bus: address-indexed router (Part 6 pattern) ----------------------------
SC_MODULE(Bus) {
  tlm_utils::simple_target_socket<Bus>                  in;    // from CPU wrapper
  tlm_utils::multi_passthrough_initiator_socket<Bus>    out;   // to RAM, UART

  struct Region { sc_dt::uint64 base, size; int idx; const char* name; };
  Region map_[2] = {
    { RAM_BASE,  RAM_SIZE,  0, "RAM"  },
    { UART_BASE, UART_SIZE, 1, "UART" },
  };

  SC_CTOR(Bus) : in("in"), out("out") {
    in.register_b_transport(this, &Bus::b_transport);
    in.register_transport_dbg(this, &Bus::transport_dbg);
  }

  bool decode(sc_dt::uint64 g, sc_dt::uint64& local, int& idx) {
    for (auto& r : map_) {
      if (g >= r.base && g < r.base + r.size) { local = g - r.base; idx = r.idx; return true; }
    }
    return false;
  }

  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    sc_dt::uint64 g = trans.get_address(), local; int idx;
    if (!decode(g, local, idx)) {
      trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE); return;
    }
    trans.set_address(local);                        // localize
    out[idx]->b_transport(trans, delay);
    trans.set_address(g);                            // restore
  }

  unsigned int transport_dbg(tlm::tlm_generic_payload& trans) {
    sc_dt::uint64 g = trans.get_address(), local; int idx;
    if (!decode(g, local, idx)) return 0;
    trans.set_address(local);
    unsigned int n = out[idx]->transport_dbg(trans);
    trans.set_address(g);
    return n;
  }
};

// --- CPU wrapper: re-fronts the core's data port with a TLM initiator socket -
// The wrapper owns the core and a single initiator socket onto the bus. It
// drives the core in an SC_THREAD and services the core's load/store callbacks
// by issuing TLM transactions through `socket`. THAT glue is Exercise 1.
SC_MODULE(CpuWrapper) {
  tlm_utils::simple_initiator_socket<CpuWrapper> socket;
  Rv32iCore core;

  SC_CTOR(CpuWrapper) : socket("socket") {
    // Install the data-memory port: route every core load/store into mem_access.
    core.load_cb  = [this](uint32_t a)            { return mem_access(false, a, 0); };
    core.store_cb = [this](uint32_t a, uint32_t d){        mem_access(true,  a, d); };
    SC_THREAD(run);
  }

  bool glue_missing = false;                         // set by the stub, read by run()

  // --- EXERCISE 1: the data-side TLM glue. -----------------------------------
  // This function is the bridge between the core's load/store port and the bus.
  // For a store (is_write true) it must build a TLM_WRITE_COMMAND payload at
  // global address `addr`, point it at a 4-byte buffer holding `data`, set
  // length 4 / streaming width 4 / no byte enables / status INCOMPLETE, call
  // socket->b_transport(trans, delay), pay the delay, and check the response.
  // For a load it must do the symmetric TLM_READ_COMMAND and return the word.
  //
  // As shipped, this stub issues NO transaction. It returns a poison value for
  // loads and, for stores, reports the missing glue and halts the platform so
  // the failure is impossible to miss. Replace the body below.
  uint32_t mem_access(bool is_write, uint32_t addr, uint32_t data) {
    // EXERCISE 1 stub -- no transaction issued (wrong on purpose).
    std::ostringstream dh; dh << std::hex << std::setw(8) << std::setfill('0') << data;
    std::cout << "\n*** EXERCISE 1 NOT DONE: CpuWrapper::mem_access has no TLM glue ***\n"
              << "    core requested a " << (is_write ? "STORE" : "LOAD")
              << " at 0x" << std::hex << std::setw(8) << std::setfill('0') << addr
              << (is_write ? (" data=0x" + dh.str()) : std::string())
              << std::dec << std::setfill(' ') << "\n"
              << "    the platform booted and the CPU ran, but the data port\n"
                 "    never reaches the bus. Implement the b_transport glue.\n";
    glue_missing = true;                             // stop the run loop after this
    return 0xDEADBEEFu;
    // -- END EXERCISE 1 --
  }

  void run() {
    std::cout << "=== virtual platform: RV32I CPU -> bus -> RAM + UART ===\n";
    int budget = 200;                                // safety cap on instructions
    while (!core.halted && !glue_missing && budget-- > 0) {
      core.step();
    }
    if      (glue_missing) std::cout << "platform halted: data port not wired (Exercise 1)\n";
    else if (core.halted)  std::cout << "\n[" << sc_time_stamp() << "] CPU halted (ebreak)\n";
    sc_stop();
  }
};

// --- Top: instantiate the platform, load the toy program, bind everything ----
//
// Toy program (RV32I): write "HELLO" to the UART DATA register one byte at a
// time, then ebreak. x1 holds the UART base; each character is materialised in
// x2 (lui|addi) and stored with `sw x2, 0(x1)`.
SC_MODULE(Top) {
  CpuWrapper cpu;
  Bus        bus;
  Ram        ram;
  Uart       uart;

  // Encode helpers (hand-assembler for the toy program).
  static uint32_t enc_lui (uint32_t rd, uint32_t imm20) { return (imm20 << 12) | (rd << 7) | 0x37; }
  static uint32_t enc_addi(uint32_t rd, uint32_t rs1, int32_t imm) {
    return ((imm & 0xFFF) << 20) | (rs1 << 15) | (0 << 12) | (rd << 7) | 0x13;
  }
  static uint32_t enc_sw  (uint32_t rs2, uint32_t rs1, int32_t imm) {
    uint32_t hi = (imm >> 5) & 0x7F, lo = imm & 0x1F;
    return (hi << 25) | (rs2 << 20) | (rs1 << 15) | (0x2 << 12) | (lo << 7) | 0x23;
  }
  static uint32_t enc_ebreak() { return (1u << 20) | 0x73; }

  SC_CTOR(Top) : cpu("cpu"), bus("bus"), ram("ram"), uart("uart") {
    cpu.socket.bind(bus.in);
    bus.out.bind(ram.socket);                        // index 0 = RAM
    bus.out.bind(uart.socket);                       // index 1 = UART

    // Build the program: x1 = UART_BASE, then store 'H','E','L','L','O'.
    auto& im = cpu.core.imem;
    int n = 0;
    auto emit = [&](uint32_t w){ im[n++] = w; };
    emit(enc_lui (1, UART_BASE >> 12));              // x1 = 0x10000000
    const char* msg = "HELLO";
    for (const char* p = msg; *p; ++p) {
      emit(enc_addi(2, 0, (unsigned char)*p));       // x2 = char
      emit(enc_sw  (2, 1, 0));                        // [x1] = x2  -> UART DATA
    }
    emit(enc_ebreak());
  }
};

int sc_main(int, char*[]) {
  Top top("top");
  sc_start();
  return 0;
}

Build it and run it exactly as shipped. Because mem_access issues no transaction, the very first sw — the store of 'H' (0x48) to the UART — has nowhere to go. The stub prints a precise diagnostic and stops the platform cleanly:

Scaffold output (as shipped, before you do anything):

=== virtual platform: RV32I CPU -> bus -> RAM + UART ===

*** EXERCISE 1 NOT DONE: CpuWrapper::mem_access has no TLM glue ***
    core requested a STORE at 0x10000000 data=0x00000048
    the platform booted and the CPU ran, but the data port
    never reaches the bus. Implement the b_transport glue.
platform halted: data port not wired (Exercise 1)

Info: /OSCI/SystemC: Simulation stopped by user.

Read that trace carefully, because it tells you the rest of the platform already works. The banner printed, so elaboration and binding succeeded — the CPU wrapper's initiator socket is bound to the bus, the bus's two initiator sockets are bound to RAM and UART, and there is no (E109) complete-binding failure. The CPU then ran: it executed the lui and the first addi (those touch only registers, no memory), reached the first sw, and called into mem_access. The diagnostic names exactly what the core asked for — a store at 0x10000000 of 0x48, which is the ASCII 'H' headed for the UART DATA register — and exactly what is missing: the glue that would turn that request into a b_transport. The failure is loud, single, and localized to one function. Your job in Exercise 1 is to turn that diagnostic into a printed HELLO.

Exercise 1: Implement the CPU-wrapper glue

Goal. Make the scaffold print HELLO and halt, by implementing the body of CpuWrapper::mem_access — and nothing else.

What to do. mem_access(is_write, addr, data) is the single bridge between the core's data-memory port and the bus. The core calls it for every load (via load_cb) and every store (via store_cb); the wrapper already installs those callbacks and already owns an initiator socket bound to the bus. You are filling in one function body, no more. As shipped it issues no transaction and reports the gap; replace that stub with the loosely-timed initiator loop you wrote in Part 2.

Specification, step by step. Inside mem_access you must:

  • Construct a tlm::tlm_generic_payload and an sc_time delay = SC_ZERO_TIME.
  • Set the command from is_write: TLM_WRITE_COMMAND for a store, TLM_READ_COMMAND for a load.
  • Set the address to the global addr the core computed. The bus localizes it; the wrapper hands over the global address exactly as the initiator did in Part 6.
  • Point the data pointer at a local 4-byte buffer. For a store, the buffer holds data before the call; for a load, the target writes the loaded word into it. A uint32_t buf and reinterpret_cast<unsigned char*>(&buf) is the idiom from Part 2.
  • Set the data length to 4 and the streaming width to 4 (a contiguous 32-bit access — bytes, not words), clear byte enables with set_byte_enable_ptr(nullptr), and set the response status to TLM_INCOMPLETE_RESPONSE as the "not done yet" sentinel.
  • Call socket->b_transport(trans, delay).
  • Pay the annotated delay with wait(delay). This is legal because mem_access runs on the wrapper's SC_THREAD (run calls core.step(), which calls back into mem_access on the same thread). The UART and RAM annotate latency rather than blocking, exactly as the section taught — so you, the initiator, pay it.
  • Check the response: if trans.is_response_error(), the access failed — report it and stop. For a load, return the word from your buffer; for a store, the return value is unused.

Acceptance. Built and run as before, the platform now prints:

=== virtual platform: RV32I CPU -> bus -> RAM + UART ===
HELLO
[25 ns] CPU halted (ebreak)

The five characters appear because each sw reaches the UART DATA register through the bus and triggers its emit side effect. The final timestamp is 25 ns: five stores, each annotated 5 ns of UART byte-time, paid as the initiator goes. If you see no HELLO and the same EXERCISE-1 diagnostic, your stub is still in place. If you see a bus error instead, check that you set the global address (0x1000_0000, not a localized offset) — localization is the bus's job, not yours.

Do not modify the core, the bus, the RAM, the UART, the bindings, or the toy program. If you find yourself editing anything outside mem_access, step back — the exercise is solvable by filling in that one function.

Exercise 2: Add a quantum keeper

Goal. Replace the scaffold's pay-as-you-go timing — a wait(delay) after every transaction — with a quantum keeper, so the CPU wrapper runs ahead of simulated time and synchronises only at quantum boundaries. Then measure the difference.

Why. The Exercise-1 wrapper pays the annotated delay immediately after every access (idiom A from Part 2). That is correct but slow: each wait() is a context switch into the SystemC kernel and back. Temporal decoupling, the subject of Part 5, lets the wrapper accumulate delay locally and synchronise only when the accumulated time exceeds a quantum — turning many tiny wait()s into a few large ones. On a real platform booting an OS, this is the difference between a model that runs in seconds and one that runs in minutes.

What to build. Add a tlm_utils::tlm_quantumkeeper qk; member to CpuWrapper and rework the timing path:

  • Set the global quantum once (for example tlm::tlm_global_quantum::instance().set(sc_time(100, SC_NS));) and qk.reset() at the start of run.
  • In mem_access, pass the keeper's local time as the delay into b_transport — that is, seed delay = qk.get_local_time() before the call so the target annotates on top of the time already banked. After the call, fold the returned delay back in with qk.set(delay) (or qk.inc(...) over the annotated increment), and call qk.need_sync() / qk.sync() to advance real simulation time only when the quantum is exhausted.
  • Do not call wait(delay) per access anymore; the keeper owns synchronisation.

What to measure. Instrument two things and compare against the Exercise-1 baseline:

  • The final sc_time_stamp() at halt. With a 100 ns quantum and only five 5 ns accesses (25 ns total, under one quantum), the architectural result is unchanged but the wall-clock-to-simulation synchronisation happens once at the end instead of five times. Confirm HELLO still prints and the program still halts.
  • The number of sync()/wait() boundaries actually crossed. Add a counter that increments each time the keeper synchronises, and compare it to the five waits the Exercise-1 wrapper performed. Then shrink the quantum (say to 4 ns) and watch the sync count rise — proving the quantum is the knob that trades timing precision for speed.

Acceptance. HELLO still prints; the program halts; and your measurement shows the keeper crossing fewer synchronisation boundaries than the per-access wait() version for a quantum larger than the total access time, and more as you shrink the quantum. The exact final timestamp may differ from 25 ns depending on how you fold the delay — that is expected and is the point: temporal decoupling trades exact per-access timing for speed, and you are measuring that trade.

Exercise 3: Load the program via debug transport

Goal. Stop hand-loading the program into the core's private instruction array at elaboration. Instead, write the program words into RAM through the bus's transport_dbg backdoor at reset, and have the core fetch each instruction from RAM through a debug read — proving the platform still prints HELLO with the program living in a real memory model rather than a C++ array.

Why. The scaffold cheats on program loading: Top's constructor pokes instruction words straight into cpu.core.imem, a plain array with no model behind it. Real platforms do not work that way — the program lives in RAM (or flash), and a loader puts it there. The professional mechanism for seeding memory without burning functional cycles is debug transport (Part 7): transport_dbg writes target state with side effects off and zero simulation time. This exercise has you use the backdoor twice — once to load, once to fetch.

What to build, in two halves.

  1. Debug-load the program into RAM. In Top (or a small loader method), instead of writing cpu.core.imem, build a tlm_generic_payload with TLM_WRITE_COMMAND, point it at your buffer of program words, and call the bus's transport_dbg at a RAM address (say base 0x0000_0000). Because it goes through the bus, it decodes to RAM and writes the bytes with no time consumed. Confirm the call returns the number of bytes you asked it to write — that byte count, not a response status, is how transport_dbg reports success.
  1. Fetch from RAM via debug read. Give the core a fetch hook (a callback like its load_cb) that, for each instruction, issues a transport_dbg read of four bytes at the PC's RAM address through the wrapper's socket, and feeds the resulting word into the decoder. Fetch is an inspection of memory that must not consume time or perturb state, which is exactly what debug transport guarantees — so fetching this way costs zero simulated time, just as an instruction fetch from a backdoor-loaded ROM should.

Constraints. The debug path must not call wait() and must not change any device's functional state — that is the transport_dbg contract. Keep the data-side accesses (the sw to the UART) on the functional b_transport path from Exercise 1; only fetch and load move to the debug path. Mixing them up — fetching functionally, or storing via debug — defeats the purpose: a functional fetch would burn RAM latency every instruction, and a debug store to the UART would not trigger the emit side effect, so HELLO would never print.

Acceptance. The program no longer lives in core.imem; it is written into RAM by a debug transaction at reset and fetched from RAM by debug reads. The platform still prints HELLO and halts. As a check that you used the backdoor correctly, confirm that the debug load consumed zero simulation time (sc_time_stamp() is unchanged across it) and that a functional read of the UART STATUS register still returns tx-ready — proving the debug path left functional state untouched.

Hints

These are hints, not solutions. They point at the shape of the answer and the trap to avoid; they stop short of the code.

Hints for Exercise 1

  • The body you need already exists, in Part 2: the LT initiator that fills a payload, calls b_transport, pays the delay, and checks the response. Port that loop into mem_access, parameterised by is_write. If your function grows past a dozen lines, you are over-building it.
  • One buffer, one command switch. A uint32_t buf = data; works for both directions: the store reads it before the call, the load receives the word into it. Point the data pointer at &buf and set_data_length(4).
  • Hand over the global address. The core computed rd_reg(rs1) + imm, which for the toy program is 0x1000_0000. Do not subtract anything — localization is the bus's job, and subtracting here would route the access to the wrong target or off the end of one.
  • wait(delay) is legal here because mem_access is reached from the wrapper's SC_THREAD. If you get a "wait() called from a method process" error, you have moved the call out of the thread's call chain — keep it on the run -> core.step() -> callback path.
  • Reset the response status to TLM_INCOMPLETE_RESPONSE before every call, not just the first. If you reuse one payload across accesses, a stale TLM_OK_RESPONSE from a previous store could mask a real failure on the next.

Hints for Exercise 2

  • The keeper replaces the wait(delay), not the b_transport. Keep the payload fill and the transport call exactly as in Exercise 1; only the timing mechanism changes from immediate-wait to accumulate-and-sync.
  • Seed the delay you pass into b_transport with qk.get_local_time() so the target annotates on top of time already banked, then fold the new total back into the keeper. Forgetting to seed it makes every access annotate from zero and silently understates accumulated time.
  • need_sync() is your gate: call sync() only when it returns true. Calling sync() after every access defeats the whole optimization — you are back to pay-as-you-go with extra steps.
  • To see the speed lever, instrument the sync count and sweep the quantum. A quantum larger than your total access time syncs once; a tiny quantum syncs nearly every access. That curve is the measurement the exercise asks for.
  • The architectural output (HELLO, then halt) must be invariant under the quantum. If HELLO changes or disappears when you change the quantum, you have let timing leak into functionality — a sign you are syncing or ordering accesses incorrectly, not just decoupling their time.

Hints for Exercise 3

  • transport_dbg returns a byte count, not a tlm_response_status. Check the count to confirm the access landed; a return of 0 means the bus could not decode the address or the target has no debug callback. The RAM in the scaffold already registers one; the UART deliberately does not.
  • Route the debug load through the bus, not straight at the RAM socket. Going through the bus exercises the same decode and localize path the functional access uses, and is how a real loader addresses RAM by its global address. The bus's transport_dbg is already implemented in the scaffold.
  • Keep stores functional. The single most damaging Exercise-3 mistake is to send the UART write down the debug path "for consistency" — a debug write to DATA does not run the emit side effect, so HELLO silently never prints. Debug is for loading and fetching; the side-effecting store stays on b_transport.
  • Confirm zero time. Bracket the debug load with sc_time_stamp() reads; if time advanced, something on that path called wait() or you accidentally used b_transport. The debug path must be timeless by contract.
  • Fetch four bytes at a time at the word-aligned PC. The RAM is byte-addressed; a fetch is a 4-byte debug read at pc & ~3. Assemble the little-endian word from the bytes the read returns, exactly as the RAM stored them.

Common mistakes

These are drawn from the bug catalogs across Parts 1–7 — the same traps resurface the moment you assemble the pieces into a platform.

  • Forgetting the response-status check (or the reset on reuse). The only channel that tells an initiator whether a transaction worked is the response status: the initiator sets TLM_INCOMPLETE_RESPONSE before, the target sets TLM_OK_RESPONSE or an error, the initiator checks is_response_ok() / is_response_error() after (Part 2). In mem_access, skipping the check means a bus address error — a store to an unmapped address, say — passes silently and the bug surfaces far downstream. And if you reuse one payload across accesses without resetting the status, a stale TLM_OK_RESPONSE masks a real later failure. Fix: reset to incomplete before every call; check after every call.
  • Reusing a payload without clearing stale fields. Because the payload is passed by reference and the platform threads one object through initiator -> bus -> target and back, a leftover byte-enable pointer, data length, or data pointer from a previous access silently corrupts the next (Part 2). A stale byte-enable mask is the classic: it masks bytes of an access that wanted none. Fix: set every field you depend on (command, address, data pointer, length) on every access, and set_byte_enable_ptr(nullptr) when you do not want a mask. A fresh stack payload per access, as the scaffold uses, sidesteps most of this.
  • Address not localized — or localized twice. The bus's job is to subtract the region base so each target lives in a clean zero-based space, then restore the global address on the return path (Part 6). Two failure modes meet here. If you, the initiator, "helpfully" subtract the base in mem_access, the bus subtracts it again and the access lands at the wrong offset. If the bus forgets to restore the global address after the call, every monitor and scoreboard above it sees the local offset instead of the real address — data values look fine, addresses are wrong, and the bug hides. Fix: the initiator always uses global addresses; the bus owns the localize-then-restore contract and does it exactly once.
  • A debug read with side effects (or that consumes time). transport_dbg is a backdoor: by contract it must not change functional state and must not consume simulation time (Part 7). In Exercise 3, fetching through a functional read instead of a debug read would burn RAM latency on every instruction and, worse, a debug write to the UART DATA register would skip the emit side effect so HELLO never appears. Fix: use debug transport only for inspection and loading; keep side-effecting functional accesses on b_transport. Never call wait() inside transport_dbg.
  • Quantum sync forgotten — or done every access. Temporal decoupling banks delay locally and synchronises only at quantum boundaries (Part 5). Forget to seed b_transport's delay from qk.get_local_time() and the keeper loses track of accumulated time, understating it. Forget to ever sync() and the wrapper runs arbitrarily far ahead of every other process, breaking causality with anything it must observe. Sync after every access and you have thrown away the speedup. Fix: seed the local time in, fold the annotated delay back, and sync() only when need_sync() says the quantum is exhausted.
  • An unbound socket at elaboration. If you add a component in Exercise 3 (a loader, an extra target) and forget to bind one of its sockets, the kernel raises (E109) complete binding failed at sc_start() — not at the constructor where the binding lives (Part 3). The hierarchical name in the message points at the exact unbound socket. Fix: read the name in the (E109) and bind that socket; do not hunt at the call site of sc_start.

Where to go next

  • Section 4 — UVM-SystemC is where this platform stops being something you drive with a hand-written SC_THREAD and becomes something you verify with a real methodology. UVM-SystemC brings the Universal Verification Methodology — sequences, drivers, monitors, scoreboards, and a phased test hierarchy — to the SystemC/TLM world you have just built. The transactions you have been issuing by hand become sequence items generated, randomised, and scored by a reusable verification environment.
  • The bridge is exactly the work you did here: a UVM-SystemC driver is the same idea as your mem_access glue — it turns abstract sequence items into TLM transactions on a socket — and a monitor is a transport_dbg-style observer that samples without perturbing. The platform you built is precisely the kind of device-under-test a UVM-SystemC environment wraps.
  • If you want to deepen the platform itself first, the natural extensions are: add a second initiator (a DMA engine) and watch the bus arbitrate; give the UART a real receive FIFO and an interrupt; or move the data path to the approximately-timed protocol from Part 4 so overlapping transactions become visible.

Further reading

Section 3 (this section), Parts 1–7

Section 2 (the RV32I CPU this capstone re-fronts)

Standards and references

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC Language Reference Manual — §10 (generic payload, b_transport, timing annotation), §11 (debug transport interface and the transport_dbg contract), §12–13 (sockets and the base protocol), and the temporal-decoupling / quantum facilities in tlm_utils.
  • Aynsley (Doulos), OSCI TLM-2.0 Language Reference Manual (JA32) — the canonical narrative on LT coding style, the base protocol, sockets, and quantum-based temporal decoupling.
  • Accellera Systems Initiative, SystemC 3.0.0 distribution, include/tlm_utils/ — authoritative source for the convenience sockets (simple_initiator_socket, simple_target_socket, multi_passthrough_initiator_socket) and tlm_quantumkeeper.
  • The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA — the RV32I base encodings the toy program uses (LUI, ADDI, SW, LW, BNE, EBREAK).

Next in the series

→ Section 4 — UVM-SystemC opens the next stretch of this series. It picks up exactly where this capstone leaves off: you have a working virtual platform driven by hand; Section 4 wraps that platform in the Universal Verification Methodology — sequences, drivers, monitors, scoreboards, and phased tests — so you verify it the way real teams sign off real SoCs. No link yet — coming next in this series.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment