15. SystemC Tutorial - Capstone: Build the Rest Yourself

Why this matters

Rewritten 2026-06-05 with the concept-first approach.

You've learned the pieces. Now build the machine.

Across the previous seven parts of this section you built every leaf of a single-cycle RV32I processor — one at a time, each from first principles, each verified against its own little testbench. You built a combinational ALU that adds, subtracts, shifts, and compares. You built an instruction decoder that cracks a 32-bit word into register addresses, an immediate, and a bundle of control bits. You built a program counter that fetches in order and redirects on a branch. You learned the two-process FSM discipline that every controller in every chip is made of. You built a 32×32 register file with x0 hardwired to zero. You built a data memory with byte enables and sign-extending loads. And you watched all six leaves get wired into one container that ran a three-instruction program and printed PASS.

Every one of those posts ended the same way: here is the worked example, here is the trace, here is the bug to avoid. This one is different. There is no worked solution here. The capstone hands you a runnable testbench harness — clock, reset, trace, program loader, self-checking register comparison — with the processor itself left as a marked gap, and three exercises of increasing difficulty that ask you to fill it in, extend it, and re-architect it. This is exactly how a real CPU verification team onboards a new engineer: the environment is handed to you complete; wiring the device-under-test into it and growing it is the job. The single most valuable thing you can do at the end of a tutorial section is close the book and rebuild the result from memory, discovering for yourself which details were load-bearing. That is what this post is for. The pieces are yours. Assemble them.

Prerequisites

This capstone assumes you have worked through — not merely skimmed — the seven RTL-patterns posts that precede it, plus the Foundations material from Section 1. Each link below is a building block you will reach for directly while solving the exercises:

If any of these is hazy, re-read it before starting. The exercises below will not re-teach the leaf modules — they assume you can reproduce each one, or at least retrieve it from your own earlier work.

What you have at this point

Concretely, here is the inventory of finished, tested building blocks you carry into this capstone. Each is a module you built and ran in an earlier part; the capstone composes and extends them.

  • An Alu (Part 8) — sc_in<sc_uint<32>> a, b, sc_in<sc_uint<4>> op, sc_out<sc_uint<32>> result, sc_out<bool> zero. A single combinational SC_METHOD switching on a ten-value alu_op_t enum (ADD, SUB, AND, OR, XOR, SLT, SLTU, SLL, SRL, SRA).
  • An Rv32iDecoder (Part 9) — takes a 32-bit instr, emits rs1_addr, rs2_addr, rd_addr, funct3, a sign-extended imm, the alu_op, and the control bits alu_src, reg_write, mem_read, mem_write, branch, wb_sel, and a halt/illegal flag.
  • A Pc (Part 10) — clocked, active-high synchronous reset, a next_pc input and a pc_out output, built on the clocked state-register + combinational-output two-process split.
  • The two-process FSM idiom (Part 11) — a clocked state_register and a combinational next_state_and_outputs, sharing one sc_signal<state_t> that is the flip-flop.
  • A RegFile (Part 12) — two combinational read ports (rs1, rs2), one synchronous write port (rd), x0 always reading zero and never written.
  • A DataMemory (Part 13) — combinational load shaped by funct3 (LB/LH/LW/LBU/LHU), synchronous store with byte-enable lane masking.
  • The composition discipline (Part 14) — leaves as members, ~30 parent-owned interconnect sc_signals, glue SC_METHOD muxes each the single writer of one wire.

That is a complete parts bin. The capstone asks you to assemble it without a step-by-step guide.

What you're going to build

The deliverable is a configurable single-cycle RV32I CPU that fetches, decodes, executes, and writes back one instruction per clock, and that runs a small toy program to completion under a self-checking testbench. "Single-cycle" means there is no pipeline: each instruction flows all the way through the datapath — PC → instruction fetch → decode → register read → ALU → (optional memory) → writeback — and commits its register-file write on the one rising clock edge that ends its cycle. "Configurable" means the design is parameterized enough that the three exercises can extend it: Exercise 2 adds an instruction by touching only the decoder and the ALU op-table, and Exercise 3 swaps the PC-advance policy for an FSM-driven one without disturbing the datapath.

The architecture is the datapath-plus-control split you composed in Part 14. The datapath is the left-to-right flow of an instruction's data through the leaf modules. The control is the decoder's output bundle steering a handful of combinational muxes that sit in the seams between leaves: the ALU-B mux (select the second ALU operand — register value or immediate), the writeback mux (select what gets written back — ALU result, load data, or PC+4), and the next-PC logic (sequential PC+4, a taken branch target, or a frozen PC on halt). Those muxes are the glue. Each is a single SC_METHOD that is the only writer of its output signal — the single-driver rule from Part 14 that keeps every wire well-defined.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b'}}}%%
flowchart LR
    PC["Pc"] -->|pc| IMEM["Imem"]
    IMEM -->|instr| DEC["Rv32iDecoder"]
    DEC -->|rs1,rs2,rd| RF["RegFile"]
    RF -->|rs1_data| ALU["Alu"]
    RF -->|rs2_data| BMUX["ALU-B mux"]
    DEC -->|imm,alu_src| BMUX
    BMUX -->|alu_b| ALU
    DEC -->|alu_op| ALU
    ALU -->|result| DMEM["DataMemory"]
    ALU -->|result| WB["WB mux"]
    DMEM -->|rd_data| WB
    PC -->|pc+4| WB
    WB -->|wr_data| RF
    ALU -->|zero| BR["next-PC logic"]
    DEC -->|branch,imm| BR
    PC -->|pc| BR
    BR -->|next_pc| PC

The toy program the harness runs is four instructions long:

0x00500093   addi x1, x0, 5     # x1 = 5
0x00300113   addi x2, x0, 3     # x2 = 3
0x002081b3   add  x3, x1, x2    # x3 = x1 + x2 = 8
0x00100073   ebreak             # halt

After it runs, the testbench checks the register file: x1 must hold 5, x2 must hold 3, x3 must hold 8. If the datapath is wired correctly, you get PASS: all registers correct. If a mux is miswired — say the ALU-B mux always selects the register operand and ignores the immediate — add x3, x1, x2 still works but the two addi instructions compute the wrong value, the register check fails, and the harness tells you exactly which register is wrong. That self-check is the whole point of the harness: it converts "it elaborated" into "it actually computes the right thing," which, as Part 14 stressed, are two very different claims.

You are not building the harness. That is provided below, complete and runnable. You are building the Cpu module that sits inside it — and then, in the later exercises, growing it.

Scaffold provided

Below is the complete scaffold. Everything here compiles and runs as given. The leaf modules are reproduced in skeleton form so the file is self-contained; the testbench harness — clock, reset sequencing, program loader, cycle trace, and the self-checking register comparison — is complete and you should not need to touch it. The one deliberate gap is the body of the three glue processes inside Cpu, each marked with an // EXERCISE 1 comment. As shipped, those stubs make the design elaborate and run (every port is bound, so there is no (E109) binding failure), but the writeback path drives zero, so the smoke test reports FAIL until you implement the muxes. That is intentional: a runnable scaffold that fails its own check is the starting line, not the finish.

Note Build with a C++17 compiler against your SystemC install. On the reference machine for this post: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-<arch> capstone.cpp -o capstone -lsystemc, then run with the SystemC library directory on your loader path. Substitute your own $SYSTEMC_HOME and lib- architecture suffix.
// file: capstone.cpp
// Single-cycle RV32I capstone scaffold.
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-<arch> \
//            capstone.cpp -o capstone -lsystemc
//
// As shipped this compiles and runs, but the three glue processes inside Cpu
// are stubbed (search for "EXERCISE 1"). The smoke test will report FAIL until
// you implement them. That is the assignment.

#include <systemc.h>
#include <iostream>
#include <iomanip>
#include <cstdint>

// ─── ALU operation encoding (from Part 8) ────────────────────────────────────
enum alu_op_t {
  ALU_ADD = 0, ALU_SUB = 1, ALU_AND = 2, ALU_OR  = 3, ALU_XOR = 4,
  ALU_SLT = 5, ALU_SLTU = 6, ALU_SLL = 7, ALU_SRL = 8, ALU_SRA = 9
};

// ─── ALU (Part 8) ────────────────────────────────────────────────────────────
SC_MODULE(Alu) {
  sc_in<sc_uint<32>>  a, b;
  sc_in<sc_uint<4>>   op;
  sc_out<sc_uint<32>> result;
  sc_out<bool>        zero;

  void compute() {
    sc_uint<32> res = 0;
    switch ((int)op.read()) {
      case ALU_ADD:  res = a.read() + b.read(); break;
      case ALU_SUB:  res = a.read() - b.read(); break;
      case ALU_AND:  res = a.read() & b.read(); break;
      case ALU_OR:   res = a.read() | b.read(); break;
      case ALU_XOR:  res = a.read() ^ b.read(); break;
      case ALU_SLT:  res = ((sc_int<32>)a.read() < (sc_int<32>)b.read()) ? 1 : 0; break;
      case ALU_SLTU: res = (a.read() < b.read()) ? 1 : 0; break;
      case ALU_SLL:  res = a.read() << b.read().range(4, 0); break;
      case ALU_SRL:  res = a.read() >> b.read().range(4, 0); break;
      case ALU_SRA:  res = (sc_uint<32>)((sc_int<32>)a.read() >> b.read().range(4, 0)); break;
    }
    result.write(res);
    zero.write(res == 0);
  }

  SC_CTOR(Alu) { SC_METHOD(compute); sensitive << a << b << op; }
};

// ─── Program counter (Part 10) ───────────────────────────────────────────────
SC_MODULE(Pc) {
  sc_in<bool>         clk;
  sc_in<bool>         rst;            // active-high synchronous reset
  sc_in<sc_uint<32>>  next_pc;
  sc_out<sc_uint<32>> pc_out;

  sc_signal<sc_uint<32>> pc_reg;

  void reg_proc() {                    // clocked: the state register
    if (rst.read()) pc_reg.write(0);
    else            pc_reg.write(next_pc.read());
  }
  void out_proc() { pc_out.write(pc_reg.read()); }   // combinational mirror

  SC_CTOR(Pc) {
    SC_METHOD(reg_proc); sensitive << clk.pos(); dont_initialize();
    SC_METHOD(out_proc); sensitive << pc_reg;
  }
};

// ─── Instruction memory (a small ROM for the toy program) ────────────────────
SC_MODULE(Imem) {
  sc_in<sc_uint<32>>  addr;            // byte address (from pc_out)
  sc_out<sc_uint<32>> instr;

  uint32_t mem[256] = {};              // 256 words, value-initialized to 0

  void load_word(uint32_t byte_addr, uint32_t word) {
    mem[(byte_addr >> 2) & 0xFF] = word;
  }
  void read_proc() { instr.write(mem[(addr.read() >> 2) & 0xFF]); }

  SC_CTOR(Imem) { SC_METHOD(read_proc); sensitive << addr; }
};

// ─── Instruction decoder (Part 9) ────────────────────────────────────────────
// Decodes the subset the toy program needs: R-type ADD/SUB, I-type ADDI, and
// the SYSTEM EBREAK (halt). Everything else decodes as a NOP that writes no
// register. Extending this decoder is the work of Exercise 2.
SC_MODULE(Rv32iDecoder) {
  sc_in<sc_uint<32>>  instr;

  sc_out<sc_uint<5>>  rs1_addr, rs2_addr, rd_addr;
  sc_out<sc_uint<3>>  funct3;
  sc_out<sc_uint<32>> imm;
  sc_out<sc_uint<4>>  alu_op;
  sc_out<bool>        alu_src;          // false = rs2, true = immediate
  sc_out<bool>        reg_write;
  sc_out<bool>        mem_read, mem_write;
  sc_out<bool>        branch;
  sc_out<sc_uint<2>>  wb_sel;           // 0 = ALU, 1 = MEM, 2 = PC+4
  sc_out<bool>        halt;

  void decode() {
    sc_uint<32> i      = instr.read();
    sc_uint<7>  opcode = i.range(6, 0);
    sc_uint<3>  f3     = i.range(14, 12);
    sc_uint<7>  f7     = i.range(31, 25);

    // Field outputs (always valid).
    rs1_addr.write(i.range(19, 15));
    rs2_addr.write(i.range(24, 20));
    rd_addr.write(i.range(11, 7));
    funct3.write(f3);

    // Control defaults: a NOP that touches nothing.
    alu_op.write(ALU_ADD); alu_src.write(false); reg_write.write(false);
    mem_read.write(false); mem_write.write(false); branch.write(false);
    wb_sel.write(0); halt.write(false); imm.write(0);

    if (opcode == 0x33) {                       // R-type (register-register)
      reg_write.write(true); alu_src.write(false); wb_sel.write(0);
      if (f3 == 0) alu_op.write(f7.range(5, 5) ? ALU_SUB : ALU_ADD);
    } else if (opcode == 0x13) {                // I-type ADDI
      reg_write.write(true); alu_src.write(true); wb_sel.write(0);
      // Sign-extend the 12-bit I-immediate.
      sc_int<32> se = (sc_int<32>)(((int32_t)(i.range(31, 20).to_uint() << 20)) >> 20);
      imm.write((sc_uint<32>)se);
      if (f3 == 0) alu_op.write(ALU_ADD);
    } else if (opcode == 0x73) {                // SYSTEM
      if (i.range(31, 20) == 1) halt.write(true);   // EBREAK
    }
  }

  SC_CTOR(Rv32iDecoder) { SC_METHOD(decode); sensitive << instr; }
};

// ─── Register file (Part 12) ─────────────────────────────────────────────────
SC_MODULE(RegFile) {
  sc_in<bool>          clk, rst;
  sc_in<sc_uint<5>>    rs1_addr, rs2_addr, rd_addr;
  sc_in<sc_uint<32>>   rd_data;
  sc_in<bool>          reg_write;
  sc_out<sc_uint<32>>  rs1_data, rs2_data;

  uint32_t regs[32] = {};

  uint32_t read_reg(int r) const { return r == 0 ? 0u : regs[r]; }  // for the TB

  void write_proc() {                  // synchronous write; x0 never written
    if (rst.read()) { for (int i = 0; i < 32; ++i) regs[i] = 0; }
    else if (reg_write.read() && rd_addr.read() != 0)
      regs[rd_addr.read()] = rd_data.read();
  }
  void read1_proc() { sc_uint<5> a = rs1_addr.read(); rs1_data.write(a == 0 ? 0u : regs[a]); }
  void read2_proc() { sc_uint<5> a = rs2_addr.read(); rs2_data.write(a == 0 ? 0u : regs[a]); }

  SC_CTOR(RegFile) {
    SC_METHOD(write_proc); sensitive << clk.pos(); dont_initialize();
    SC_METHOD(read1_proc); sensitive << rs1_addr;
    SC_METHOD(read2_proc); sensitive << rs2_addr;
  }
};

// ─── Data memory (Part 13) ───────────────────────────────────────────────────
// The toy program performs no loads or stores; this leaf is present so the
// datapath is complete and the writeback mux has a memory source to select.
SC_MODULE(DataMemory) {
  sc_in<bool>          clk;
  sc_in<sc_uint<32>>   addr, wdata;
  sc_out<sc_uint<32>>  rdata;
  sc_in<bool>          mem_read, mem_write;
  sc_in<sc_uint<3>>    funct3;

  uint8_t mem[4096] = {};

  void read_proc() {
    if (!mem_read.read()) { rdata.write(0); return; }
    uint32_t a = (uint32_t)addr.read();
    rdata.write((uint32_t)mem[a]       | ((uint32_t)mem[a + 1] << 8)
              | ((uint32_t)mem[a + 2] << 16) | ((uint32_t)mem[a + 3] << 24));
  }
  void write_proc() {
    if (!mem_write.read()) return;
    uint32_t a = (uint32_t)addr.read();
    uint32_t d = (uint32_t)wdata.read();
    for (int k = 0; k < 4; ++k) mem[a + k] = (uint8_t)((d >> (8 * k)) & 0xFF);
  }

  SC_CTOR(DataMemory) {
    SC_METHOD(read_proc);  sensitive << addr << mem_read << funct3;
    SC_METHOD(write_proc); sensitive << clk.pos(); dont_initialize();
  }
};

// ─── Cpu — the composed top module (Exercise 1 fills the glue) ───────────────
SC_MODULE(Cpu) {
  sc_in<bool>          clk, rst;
  sc_out<sc_uint<32>>  dbg_pc, dbg_instr;
  sc_out<bool>         dbg_halt;

  // Leaf instances (members).
  Pc            i_pc;
  Imem          i_imem;
  Rv32iDecoder  i_dec;
  RegFile       i_rf;
  Alu           i_alu;
  DataMemory    i_dmem;

  // Interconnect signals — parent owns every wire.
  sc_signal<sc_uint<32>> sig_pc, sig_pc_plus4, sig_instr, sig_next_pc;
  sc_signal<sc_uint<5>>  sig_rs1_addr, sig_rs2_addr, sig_rd_addr;
  sc_signal<sc_uint<32>> sig_imm;
  sc_signal<sc_uint<4>>  sig_alu_op;
  sc_signal<bool>        sig_alu_src, sig_reg_write, sig_mem_read, sig_mem_write, sig_branch, sig_halt;
  sc_signal<sc_uint<3>>  sig_funct3;
  sc_signal<sc_uint<2>>  sig_wb_sel;
  sc_signal<sc_uint<32>> sig_rs1_data, sig_rs2_data;
  sc_signal<sc_uint<32>> sig_alu_b, sig_alu_result;
  sc_signal<bool>        sig_alu_zero;
  sc_signal<sc_uint<32>> sig_mem_rd_data, sig_wr_data;

  // ── Glue process 1: ALU-B mux. Single writer of sig_alu_b. ──
  void alu_b_mux() {
    // EXERCISE 1: drive sig_alu_b. When sig_alu_src is true the ALU's second
    // operand is the immediate (sig_imm); when false it is the register value
    // (sig_rs2_data). Replace the stub below with that selection.
    sig_alu_b.write(0);          // EXERCISE 1 stub — always zero (wrong on purpose)
  }

  // ── Glue process 2: writeback mux. Single writer of sig_wr_data. ──
  void writeback_mux() {
    // EXERCISE 1: drive sig_wr_data by selecting on sig_wb_sel:
    //   0 -> sig_alu_result, 1 -> sig_mem_rd_data, 2 -> sig_pc_plus4.
    // A single sc_signal may have only ONE writer, so this whole selection is
    // one process. Replace the stub below.
    sig_wr_data.write(0);        // EXERCISE 1 stub — always zero (wrong on purpose)
  }

  // ── Glue process 3: next-PC logic. Single writer of sig_next_pc. ──
  void branch_logic() {
    // EXERCISE 1: drive sig_next_pc.
    //   if sig_halt    -> hold the current PC (sig_pc)
    //   else if a branch is taken (sig_branch && BEQ condition via sig_alu_zero
    //        on funct3==0) -> sig_pc + sig_imm
    //   else -> sig_pc + 4
    // The stub below always advances sequentially, which is enough to fetch the
    // toy program but never halts and never branches. Replace it.
    sig_next_pc.write(sig_pc.read() + 4);   // EXERCISE 1 stub
  }

  // PC+4 helper (complete — not part of the exercise).
  void pc_plus4()  { sig_pc_plus4.write(sig_pc.read() + 4); }

  // Observation outputs (complete).
  void drive_dbg() {
    dbg_pc.write(sig_pc.read());
    dbg_instr.write(sig_instr.read());
    dbg_halt.write(sig_halt.read());
  }

  SC_CTOR(Cpu)
    : i_pc("i_pc"), i_imem("i_imem"), i_dec("i_dec"),
      i_rf("i_rf"), i_alu("i_alu"), i_dmem("i_dmem")
  {
    // PC
    i_pc.clk(clk); i_pc.rst(rst); i_pc.next_pc(sig_next_pc); i_pc.pc_out(sig_pc);
    // IMEM
    i_imem.addr(sig_pc); i_imem.instr(sig_instr);
    // Decoder
    i_dec.instr(sig_instr);
    i_dec.rs1_addr(sig_rs1_addr); i_dec.rs2_addr(sig_rs2_addr); i_dec.rd_addr(sig_rd_addr);
    i_dec.funct3(sig_funct3); i_dec.imm(sig_imm); i_dec.alu_op(sig_alu_op);
    i_dec.alu_src(sig_alu_src); i_dec.reg_write(sig_reg_write);
    i_dec.mem_read(sig_mem_read); i_dec.mem_write(sig_mem_write);
    i_dec.branch(sig_branch); i_dec.wb_sel(sig_wb_sel); i_dec.halt(sig_halt);
    // Register file
    i_rf.clk(clk); i_rf.rst(rst);
    i_rf.rs1_addr(sig_rs1_addr); i_rf.rs2_addr(sig_rs2_addr); i_rf.rd_addr(sig_rd_addr);
    i_rf.rd_data(sig_wr_data); i_rf.reg_write(sig_reg_write);
    i_rf.rs1_data(sig_rs1_data); i_rf.rs2_data(sig_rs2_data);
    // ALU
    i_alu.a(sig_rs1_data); i_alu.b(sig_alu_b); i_alu.op(sig_alu_op);
    i_alu.result(sig_alu_result); i_alu.zero(sig_alu_zero);
    // Data memory
    i_dmem.clk(clk); i_dmem.addr(sig_alu_result); i_dmem.wdata(sig_rs2_data);
    i_dmem.mem_read(sig_mem_read); i_dmem.mem_write(sig_mem_write);
    i_dmem.funct3(sig_funct3); i_dmem.rdata(sig_mem_rd_data);

    // Glue processes — single writer of each output signal.
    SC_METHOD(alu_b_mux);     sensitive << sig_rs2_data << sig_imm << sig_alu_src;
    SC_METHOD(writeback_mux); sensitive << sig_alu_result << sig_mem_rd_data
                                        << sig_pc_plus4 << sig_wb_sel;
    SC_METHOD(branch_logic);  sensitive << sig_branch << sig_funct3 << sig_alu_zero
                                        << sig_pc << sig_imm << sig_halt;
    SC_METHOD(pc_plus4);      sensitive << sig_pc;
    SC_METHOD(drive_dbg);     sensitive << sig_pc << sig_instr << sig_halt;
  }
};

// ─── Testbench harness (COMPLETE — do not modify) ────────────────────────────
SC_MODULE(Tb) {
  sc_clock               clk;
  sc_signal<bool>        rst, dbg_halt;
  sc_signal<sc_uint<32>> dbg_pc, dbg_instr;

  Cpu u_cpu;

  static const uint32_t prog[4];

  void run() {
    rst.write(true);
    wait(clk.posedge_event());        // first edge: PC loads its reset value 0
    rst.write(false);
    for (int d = 0; d < 6; ++d) wait(SC_ZERO_TIME);  // settle fetch/decode/exec

    std::cout << "=== single-cycle RV32I smoke test ===\n";
    int cyc = 0;
    while (cyc < 20) {
      std::cout << "[cyc " << std::dec << std::setw(2) << cyc << "] PC=0x"
                << std::hex << std::setw(2) << std::setfill('0')
                << (uint32_t)dbg_pc.read()
                << " instr=0x" << std::setw(8) << (uint32_t)dbg_instr.read()
                << std::setfill(' ') << "\n";
      if (dbg_halt.read()) break;     // EBREAK seen — stop fetching
      wait(clk.posedge_event());      // commit this instruction, advance PC
      for (int d = 0; d < 6; ++d) wait(SC_ZERO_TIME);
      ++cyc;
    }

    bool pass = true;
    auto chk = [&](int r, uint32_t want) {
      uint32_t got = u_cpu.i_rf.read_reg(r);
      if (got != want) {
        std::cout << "FAIL x" << std::dec << r << " = 0x" << std::hex << got
                  << " (want 0x" << want << ")\n";
        pass = false;
      }
    };
    chk(1, 5); chk(2, 3); chk(3, 8);
    std::cout << (pass ? "PASS: all registers correct\n"
                       : "FAIL: register mismatch\n");
    sc_stop();
  }

  SC_CTOR(Tb) : clk("clk", 10, SC_NS), u_cpu("u_cpu") {
    u_cpu.clk(clk); u_cpu.rst(rst);
    u_cpu.dbg_pc(dbg_pc); u_cpu.dbg_instr(dbg_instr); u_cpu.dbg_halt(dbg_halt);
    for (int i = 0; i < 4; ++i) u_cpu.i_imem.load_word(i * 4, prog[i]);
    SC_THREAD(run);
  }
};

const uint32_t Tb::prog[4] = {
  0x00500093u,   // addi x1, x0, 5
  0x00300113u,   // addi x2, x0, 3
  0x002081b3u,   // add  x3, x1, x2
  0x00100073u,   // ebreak
};

int sc_main(int, char*[]) {
  Tb tb("tb");
  sc_start();
  return 0;
}

Build it and run it exactly as shipped. Because the three glue processes are stubbed, the writeback bus is always zero and the smoke test fails — but it fails cleanly, fetching all four instructions and halting on ebreak:

Scaffold output (as shipped, before you do anything):

=== single-cycle RV32I smoke test ===
[cyc  0] PC=0x00 instr=0x00500093
[cyc  1] PC=0x04 instr=0x00300113
[cyc  2] PC=0x08 instr=0x002081b3
[cyc  3] PC=0x0c instr=0x00100073
FAIL x1 = 0x0 (want 0x5)
FAIL x2 = 0x0 (want 0x3)
FAIL x3 = 0x0 (want 0x8)
FAIL: register mismatch

Read that trace carefully, because it tells you the harness already works: the PC advances 0x00 → 0x04 → 0x08 → 0x0c, the right instruction word is fetched at each address, and the simulation halts on the ebreak. The only thing wrong is the architectural result — every register is zero, because the writeback mux stub drives zero. Your job in Exercise 1 is to turn those three FAIL lines into PASS: all registers correct by filling in the three glue processes. Nothing else in the file needs to change.

Exercise 1: Wire up the datapath

Goal. Make the scaffold's smoke test print PASS: all registers correct by implementing the three glue processes inside Cpu — and nothing else.

What to do. The three SC_METHOD glue processes (alu_b_mux, writeback_mux, branch_logic) are each marked // EXERCISE 1 and currently drive a placeholder value. Replace each stub body with the correct selection logic. The leaf modules, the interconnect signals, the bindings, and the sensitivity lists are already in place; you are filling in three short function bodies, no more.

Specification, process by process.

  • alu_b_mux is the single writer of sig_alu_b, the ALU's second operand. The decoder raises sig_alu_src when the instruction's second operand is an immediate rather than a register. So this mux selects between sig_imm (when alu_src is asserted) and sig_rs2_data (when it is not). For the toy program, the two addi instructions assert alu_src and want the immediate; the add instruction deasserts it and wants the register value rs2. Get this wrong and addi x1, x0, 5 adds x0 + rs2 instead of x0 + 5.
  • writeback_mux is the single writer of sig_wr_data, the value written back to the register file's rd port. It selects on the 2-bit sig_wb_sel the decoder produces: 0 selects the ALU result (sig_alu_result), 1 selects a load's data from memory (sig_mem_rd_data), and 2 selects the link value PC+4 (sig_pc_plus4) used by jump-and-link. The toy program only exercises selector 0, but write the full three-way selection now — Exercises 2 and 3 will lean on it. Remember the single-writer rule: this entire selection lives in one process. You may not fix a miswire by adding a second process that also writes sig_wr_data; that trips the (W116) multiple-driver warning and produces an implementation-defined value.
  • branch_logic is the single writer of sig_next_pc, the value the PC latches on the next clock edge. It implements three cases in priority order: if sig_halt is asserted (the decoder saw ebreak), hold the PC at its current value so the machine stops advancing; otherwise, if this is a taken branch — sig_branch asserted and the branch condition holds (for BEQ, funct3 == 0, the condition is sig_alu_zero) — redirect to sig_pc + sig_imm; otherwise advance sequentially to sig_pc + 4. The toy program never branches, but it does halt: without the halt case, the PC keeps advancing past ebreak, the trace never stops on the dbg_halt check, and the run only ends when the 20-cycle safety cap in the harness trips.

Acceptance. The smoke test prints the same PC/instruction trace as the scaffold (the fetch path was already correct) followed by PASS: all registers correct. If you see a register mismatch, the harness names the offending register — use that to localize which mux is wrong. A wrong x1/x2 but correct x3 points at the ALU-B mux (immediates broken, register path fine); all three wrong points at the writeback mux (nothing is being written back); a run that never halts points at the next-PC logic.

Do not change the leaf modules, the bindings, the sensitivity lists, or the harness. If you find yourself editing anything outside the three stubbed bodies, step back — the exercise is solvable by touching only those three functions.

Exercise 2: Add a multiply instruction

Goal. Extend the working CPU from Exercise 1 to execute MUL — the low-32-bits integer multiply from the RISC-V "M" standard extension — and prove it with a new toy program.

Background. MUL rd, rs1, rs2 computes rd = (rs1 * rs2)[31:0]: the full product is 64 bits wide, and MUL writes back only the lower 32. It is an R-type instruction sharing the 0x33 opcode with ADD/SUB, distinguished by funct7 == 0x01 (the M-extension marker) together with funct3 == 0x0. The other funct3 values under funct7 == 0x01 are the high-half multiplies (MULH, MULHSU, MULHU) and the divides/remainders — you are adding only MUL here.

What to do — three coordinated edits.

  1. ALU op-table. Add a new alu_op_t enumerator for multiply (the next free value after ALU_SRA) and a corresponding case in the ALU's compute() switch that writes the low 32 bits of the product. The operands are sc_uint<32>; produce the 32-bit truncated product and write it to result. Set the zero flag consistently with the other operations.
  1. Decoder. In the R-type branch of the decoder (opcode == 0x33), add the discrimination on funct7. When funct7 == 0x01 and funct3 == 0x0, emit your new multiply alu_op instead of ADD/SUB. Keep reg_write asserted, alu_src deasserted (both operands are registers), and wb_sel == 0 (writeback from the ALU). Do not disturb the existing ADD/SUB decode for funct7 == 0x00.
  1. Test. Extend the toy program (or write a second one) that loads two registers with addi, multiplies them into a third with your new MUL, and ends with ebreak. Add a register check in the harness's run() for the expected product. For a worked target you can hand-assemble: MUL x4, x1, x2 with x1 = 5, x2 = 3 should leave x4 = 15. The R-type encoding packs funct7 in bits [31:25], rs2 in [24:20], rs1 in [19:15], funct3 in [14:12], rd in [11:7], and opcode in [6:0] — assemble the bits yourself as practice; the encoding for MUL x4, x1, x2 works out to 0x02208233. (Notice how close that is to add x3, x1, x2 = 0x002081b3 — only the funct7 field and rd differ. That single funct7 bit is exactly what your decoder edit must key on.)

Acceptance. Your extended program runs to ebreak, and the register check confirms the product. Critically, verify that ADD and SUB still pass — a common mistake is to let the multiply case shadow the add case so that ordinary add instructions start multiplying. Keep the original four-instruction smoke test in place (or re-run it) to confirm no regression.

Stretch. Add MULH (high 32 bits, signed×signed) as a second op. It needs a 64-bit signed intermediate product and writes back the upper word. This exposes the sign-handling subtlety the low multiply hides — but it is genuinely optional.

Exercise 3: Convert PC update to a Moore FSM

Goal. Replace the always-advance, purely combinational PC-update policy with a small two-state Moore finite state machine that gates instruction commit — re-using the exact FSM discipline from Part 11.

Why. A real single-cycle machine commits every instruction in one cycle, so a controller is not strictly required. But the moment you want to stretch one operation across multiple cycles — a slow multiply, a memory access that takes two cycles, a halt-and-resume — you need a controller FSM that decides, each cycle, whether the datapath advances. This exercise has you insert that controller in its simplest form and wire it into the existing datapath, so the structure is in place for the multi-cycle work that Section 5 will build on.

What to build. A ControlFsm leaf module following the two-process pattern from Part 11: one clocked state_register (SC_METHOD, sensitive to clk.pos(), with dont_initialize(), the only writer of the state signal) and one combinational next_state_and_outputs (SC_METHOD, sensitive to state and every input it reads). Two states:

  • S_FETCH — the machine is fetching and decoding the current instruction. Output: advance = 0 (do not yet commit the PC update / register write).
  • S_EXEC — the instruction's datapath result is valid. Output: advance = 1 (commit: let the PC take next_pc and let the register file write).

The FSM alternates S_FETCH → S_EXEC → S_FETCH → … each clock, except that when halt is asserted it parks in a terminal state (or holds S_FETCH with advance = 0) so the machine stops. This is a Moore machine: the advance output is a function of state only, so it is registered and glitch-free — exactly the property you want for an enable that gates a register write.

How it threads into the datapath. The FSM's advance output becomes a gating term. Two wires need it:

  • The PC's commit: the next_pc the PC latches should be the computed next-PC when advance is high, and the current PC (hold) when it is low. You can express this by routing advance into the existing branch_logic glue, or by adding a hold term.
  • The register-file write: gate reg_write with advance so a register commit happens only in S_EXEC. Introduce one combinational glue process (single writer) that ANDs the decoder's reg_write with the FSM's advance, and bind that to the register file's reg_write input instead of the decoder's raw output.

Because each instruction now takes two cycles (one S_FETCH, one S_EXEC), your trace will show each PC value held for two clock edges before advancing. Update the harness's expected cycle accounting accordingly — the architectural result (x1=5, x2=3, x3=8) must be unchanged; only the timing changes.

Specification recap (no code). New leaf ControlFsm(clk, rst, halt → advance), two Moore states, two-process pattern, output advance gates both the PC commit and the register write. The datapath leaves you built are untouched; you are inserting a controller and two gating terms.

Acceptance. The smoke test still ends in PASS: all registers correct, but the cycle trace now shows each instruction occupying two cycles. Confirm the FSM parks on halt and the machine stops. If the registers come out wrong, the usual suspect is that the write-enable gate is missing — reg_write is still wired straight from the decoder, so the write happens in both S_FETCH and S_EXEC, committing twice (harmless for these instructions but wrong in principle, and a real bug the moment an instruction's operands are not yet stable in S_FETCH).

Hints

These are hints, not solutions. They point at the shape of the answer and the trap to avoid; they stop short of the code.

Hints for Exercise 1

  • Each glue process reads exactly the signals already named in its sensitive << ... list. The list is your checklist of inputs: alu_b_mux is sensitive to sig_alu_src, sig_imm, sig_rs2_data, so those three are the only signals it should read. If you reach for a signal that is not in the sensitivity list, either you are over-reaching or the list needs that entry — but for Exercise 1, the lists are already correct, so reaching outside them is the warning sign.
  • A two-way mux is one if/else. A three-way mux is a switch (or cascading if/else) on the selector. Both are a handful of lines. If a glue body grows past a few lines, you are probably solving more than the exercise asks.
  • The writeback selector values are fixed by the decoder: 0 = ALU, 1 = memory, 2 = PC+4. Match those exactly; an off-by-one in the selector silently routes the wrong source.
  • For the next-PC logic, think in priority order: halt beats branch, branch beats sequential. Writing the three cases in the wrong order — sequential first — means the halt and branch cases never get a chance to override.
  • When the harness reports which register is wrong, map the failure back to a mux per the localization guide in the exercise. Do not guess — the harness is telling you where to look.

Hints for Exercise 2

  • The multiply is an R-type instruction. Everything about how an R-type flows through your datapath — register read, ALU, writeback — is already correct from Exercise 1. The only new behavior is one ALU operation and one decoder discrimination. Resist the urge to add new datapath wires; you do not need any.
  • The discriminator is funct7. ADD and SUB already live under opcode 0x33 with funct3 == 0, separated by a bit of funct7. MUL lives in the same neighborhood with funct7 == 0x01. Your decoder edit is a nested condition inside the branch you already wrote.
  • The product of two 32-bit values can overflow 32 bits — that is expected. MUL defines its result as the low 32 bits, so a 32-bit truncating multiply is correct, not a bug. Do the arithmetic in a width that holds the operands and let the assignment to the 32-bit result truncate.
  • Test the boring case first: 5 * 3 = 15. Only once that passes should you try a case that overflows 32 bits, to confirm truncation behaves.
  • Guard against regression by keeping ADD/SUB in your test. The single most common Exercise-2 bug is a decoder edit that accidentally routes funct7 == 0x00 adds into the multiply op.

Hints for Exercise 3

  • Rebuild the two-process FSM from muscle memory, not by copying the datapath leaves: one clocked state_register that is the sole writer of state, one combinational next_state_and_outputs that reads state plus its inputs and writes next_state and the outputs. If you find yourself writing state.write(...) anywhere but the clocked process, stop — that is the "combinational write to a register" bug from Part 11.
  • dont_initialize() goes on the clocked process, never on the combinational one. The combinational process must run once at startup to establish a defined advance output before the first edge.
  • The Moore output advance is a function of state alone — one line. If you find yourself reading an input to compute advance, you have drifted into Mealy territory; for this exercise keep it Moore so the enable is registered and glitch-free.
  • The write-enable gate is a separate, tiny combinational process: it ANDs the decoder's reg_write with the FSM's advance and is the single writer of the gated signal you feed to the register file. Do not let two processes drive the register file's reg_write input.
  • Keep the datapath leaves untouched. The whole point is that a correctly composed datapath accepts a controller in front of its commit points without internal surgery — you are adding a leaf and two gating terms, not rewriting the machine.

Common mistakes

These are drawn from the worked-example bugs across Parts 8–14 — the same traps resurface the moment you assemble the pieces yourself.

  • Missing sensitivity entry on a glue mux. SystemC has no always @(*) auto-sensitivity (Part 11): every signal a combinational SC_METHOD reads must appear in its sensitive << ... list. If you add a third input to a mux in Exercise 2 or 3 but forget to extend the sensitivity list, the mux stops re-evaluating when that input changes — it updates only when one of the listed signals moves. The output looks "stuck" or one event behind. Fix: every .read() in the body corresponds to one entry in the list.
  • A combinational write to a register. In Exercise 3, writing state from the combinational next_state_and_outputs process (instead of next_state) bypasses the clock and turns your FSM into an asynchronous latch network with no memory (Part 11). The state changes whenever an input wiggles and the machine has no notion of "next cycle." Fix: the only state.write(...) lives in the clocked state_register process; the combinational process writes next_state.
  • Writing x0 and expecting it to stick. The register file hardwires x0 to zero — it discards writes to address 0 and always reads it as zero (Part 12). If a test writes a value to x0 and then checks for it, the check fails by design, not by bug. This bites when you hand-assemble a test program and accidentally target rd = x0. Fix: never use x0 as a destination when you expect a result; it is the architectural zero register.
  • Byte-enable / width errors on memory access. The toy program does no loads or stores, but the moment you extend toward LW/SW you re-enter the byte-enable territory of Part 13: the store lane mask is (width_mask << (addr & 0x3)), and a load must sign- or zero-extend per funct3. Getting the lane shift wrong writes the right bytes to the wrong offset; getting the extension wrong corrupts the upper bits of a loaded byte or halfword. Fix: re-derive the byte-enable from the address's low two bits and the access width, exactly as Part 13 did — do not hard-code a full-word mask.
  • An unbound port at elaboration. If you add a new leaf in Exercise 3 (the ControlFsm) and forget to bind one of its ports, the kernel raises (E109) complete binding failed at sc_start() — not at the constructor where you wrote the binding (Part 14). Engineers misread the timing and hunt in the wrong place. Fix: read the hierarchical name in the (E109) message; it names the exact unbound port (tb.u_cpu.i_fsm.advance). Bind it.
  • Two drivers on one interconnect signal. Trying to "fix" a writeback or next-PC bug by adding a second process that also writes the signal produces a (W116) multiple-driver warning and a value that depends on implementation-defined update ordering (Part 14). It may even look correct on your machine and break on a colleague's. Fix: every interconnect signal has exactly one writer; funnel competing sources through a single mux process.

Where to go next

  • Section 3 — TLM & Virtual Platforms is where the single-cycle CPU you just built stops being the end of the story and becomes one component in a system. Register-transfer modeling is precise but slow: simulating every signal every delta cycle does not scale to booting an operating system. This is where virtual platforms come in — they trade cycle-accurate signal wiggling for transaction-level messages between models, so a whole SoC can run real software fast enough to be useful before silicon exists.
  • The CPU here is a perfect on-ramp to that world: you will wrap a processor model in a TLM-2.0 initiator socket, connect it to memory and peripheral models through a bus, and watch the same kind of addi/add/load/store traffic you simulated signal-by-signal here flow instead as transactions — orders of magnitude faster.
  • If you want to go deeper on the RTL side first, the natural extensions are: fill out the rest of the RV32I instruction set (all branches, loads, stores, LUI/AUIPC, JAL/JALR), then verify your machine against a software reference model running in lockstep — the exact sign-off methodology real CPU vendors use.

Further reading

Section 2 (this section), Parts 1–7

Standards and references

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC Language Reference Manual — §4.3.4 (elaboration and the complete-binding check), §5.2.16 (SC_METHOD), §6.4 (sc_signal and the single-writer rule).
  • The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA — the RV32I base and the "M" standard extension (the MUL encoding used in Exercise 2: opcode 0x33, funct3 = 0x0, funct7 = 0x01).
  • CV32E40P top-level RTL (OpenHW Group, public source) — a production RISC-V core whose datapath-plus-control composition mirrors, at industrial scale, the structure of the CPU you assembled here; useful as an idiom reference for how the controller FSM and datapath meet.
  • Berkeley CS152 lecture notes, single-cycle datapath module — the canonical pedagogical treatment of the datapath/control partition this capstone follows.

Next in the series

→ Section 3 — TLM & Virtual Platforms opens the next stretch of this series with the next global post number. It picks up exactly where this capstone leaves off: you have a working RTL CPU; Section 3 lifts it into a transaction-level model and shows how virtual platforms run real software on it long before silicon. No link yet — coming next in this series.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment