21. SystemC Tutorial - Bus Fabric: Routing & Decoding
Why this matters
Written 2026-06-05 for the concept-first series.
Every virtual platform with more than one memory-mapped component has a piece in the middle that nobody draws on the marketing diagram but that every transaction passes through: the bus fabric. The processor model issues a read to 0x1000_0040 and does not know — and must not need to know — whether that address lands in RAM, in a UART register block, or in nothing at all. Something between the initiator and the targets reads that address, decides which target owns it, hands the transaction to exactly that target, and routes the answer back. That something is the bus. In real silicon it is an AXI crossbar or a network-on-chip with arbiters, QoS, and out-of-order reordering. In a transaction-level model it collapses to a strikingly small idea: a bus is an address-indexed function-call router. It looks at one field of the payload — the address — uses it as a key into a table, and forwards the same transaction object to the matching target. That is the load-bearing 90% of every interconnect model you will ever write, and this post builds it from nothing three times over, each compile-verified against an installed SystemC.
This is the post where the system stops being a single initiator wired to a single target and becomes a platform — many initiators, many targets, one fabric mediating between them. The mechanics are deceptively small and deceptively easy to get subtly wrong in ways that pass casual testing and corrupt results later. Forget to restore the address after localizing it and every scoreboard above the bus sees garbage. Forget that the payload is forwarded, not copied and you will reach for a deep-copy that breaks the response path. Believe the bus needs an arbiter to be correct and you will over-engineer a loosely-timed model that never had contention to arbitrate. Believe address translation breaks DMI and you will disable an optimization that the standard specifically tells you how to keep. This post teaches the routing primitive — memory map as a table of {base, size, target} tuples, decode, address localization, error response on a miss — then layers the three things that make a fabric real: multiple initiators through tagged sockets, latency modeled as a hop delay on the timing annotation, and the forwarding of debug and DMI traffic through the same decode logic. By the end you will write a router from scratch, route between RAM at 0x0000_0000 and a peripheral at 0x1000_0000, extend it to several masters, nest one fabric behind another, and name exactly which payload fields an interconnect is and is not allowed to touch. That is the bar, and the capstone later in this section assumes you have cleared it.
Prerequisites
- Part 2 — The Generic Payload & Blocking Transport. The bus moves
tlm_generic_payloadobjects throughb_transport, so you must already be fluent in the payload's fields, who owns each one, theb_transport(trans, delay)signature, the response-status discipline (TLM_INCOMPLETE_RESPONSEbefore,TLM_OK_RESPONSEor an error after), and timing annotation through thesc_time& delayargument. This post adds one new payload rule on top of that foundation — that an interconnect may mutate the address — and reuses everything else verbatim. - Part 3 — Initiator & Target Sockets. A bus is the canonical reason convenience sockets exist. You need the forward/backward interface picture from Part 3 to understand why a bus has both a target socket (facing the initiator) and initiator sockets (facing the targets), and why a single component can be both at once. This post introduces two new convenience sockets —
simple_target_socket_taggedandmulti_passthrough_initiator_socket— that are direct extensions of the simple sockets Part 3 covered. - The RTL Memory-Interface post (data memory). The decode logic in this post is the transaction-level twin of the combinational address decoder you build in RTL. Having seen the RTL version — an
always_combpriority chain selecting a target by address bits — makes the TLM version land as "the same decision, expressed as a table lookup instead of a synthesizable mux." We call back to it explicitly when we build the decoder.
Three ideas from Part 2 get heavy use here: the payload is one object passed by reference (the bus forwards it, never copies it); the response status is the only success/failure channel (the bus forwards it untouched, except on an unmapped miss); and timing is annotated onto delay, never realized with wait() inside a callback (the bus adds its hop delay and returns). If any of those is hazy, re-read Part 2 before continuing.
Mental model (first principles)
Start with the most reductive correct statement of what a bus is and build up.
A bus is a function-call router keyed on the address. It is neither an initiator (it originates no transactions of its own) nor a target (it is the ultimate destination of none); it sits between them and redirects. In TLM-2.0 vocabulary a component with this in-the-middle role is called an interconnect. It has a target socket on its initiator-facing side — that is the socket the upstream master calls b_transport on — and one or more initiator sockets on its target-facing side, which it calls b_transport on in turn. The whole behavior of the simplest possible bus is: receive a transaction on the target socket, pick an initiator socket by address, forward the same transaction, return.
The routing decision is a table lookup, and the table is the system's memory map. Write the map down as exactly what it is — an ordered list of {base, size, target} tuples:
| Region | Base | Size | Target |
|---|---|---|---|
| RAM | 0x0000_0000 |
64 KB | target 0 |
| Scratch peripheral | 0x1000_0000 |
4 KB | target 1 |
| (anything else) | — | — | error |
That table is the bus's configuration. To route a transaction the bus reads the payload's address, walks the table, and finds the first tuple whose half-open window [base, base + size) contains the address. The matching tuple names the outbound socket. Nothing else about the transaction enters the decision — not the command, not the length, not the data. The bus is a pure address-to-target function with the memory map as its lookup structure. This is the same decision an RTL address decoder makes with an always_comb priority chain; the only difference is that the RTL version is frozen at synthesis while the TLM table is ordinary runtime data you can build, inspect, and extend in sc_main.
Two operations turn a raw table lookup into a correct forward, and both are about the address field, the one payload attribute an interconnect is permitted to modify:
- Localization (forward path). A target models its own private address space starting at zero. RAM does not know it lives at
0x0000_0000and scratch does not know it lives at0x1000_0000; each expects to be addressed from offset 0. So before forwarding, the bus subtracts the matched region's base: a global0x1000_0010becomes a scratch-local0x0000_0010. The bus writes that local address back into the payload withset_address, then forwards.
- Restoration (return path). After the target returns, the bus writes the original global address back into the payload with
set_address, before it returns control upstream. This is mandatory and is the single most-forgotten step in interconnect modeling. The initiator — and every monitor, scoreboard, and coverage collector layered above the bus — inspectsget_address()after the call to know which address the completed transaction touched. If the bus leaves the local offset in place, everything upstream sees0x0000_0010instead of0x1000_0010and reaches wrong conclusions while the data values look perfectly fine. The bug hides because functionality is unaffected; only the address field is corrupted.
Now state the ownership rule precisely, because it is the heart of correct interconnect modeling. An interconnect forwards the same payload object — it never copies it — and may modify exactly two fields on the way through: the address (which it must restore on return) and the DMI-allowed hint. It must not touch the command, data pointer, data length, byte enables, streaming width, or response status. The response status belongs to the ultimate target; the bus forwards it back untouched. The one exception is the unmapped case: when no tuple covers the address, the bus has no target to forward to, so it becomes the responder of last resort and stamps TLM_ADDRESS_ERROR_RESPONSE itself, exactly as a target would for an out-of-range access. That is the bus's only license to write the response status.
The lifecycle of one routed transaction is therefore completely deterministic and clock-free:
- The initiator fills a payload (command, global address, data pointer, length, status =
TLM_INCOMPLETE_RESPONSE) and callsbus_socket->b_transport(trans, delay). - The bus reads
trans.get_address(), walks the memory map, and finds the covering region — or finds none. - Hit: the bus saves the global address, writes the local address (
global − base), optionally adds a hop delay todelay, and calls the matching target'sb_transport(trans, delay). Control flows into the target, which services the access and sets the response status exactly as in Part 2. - The target returns. The bus restores the global address into the payload and returns control upstream.
- Miss: the bus sets
TLM_ADDRESS_ERROR_RESPONSEand returns without touching any target. - The initiator inspects the response status and (for a read) the data — both visible because it was the same object all along.
That loop is the entire substance of bus modeling. Multiple initiators, latency, nested fabrics, and debug/DMI forwarding are all refinements of step 3.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
sequenceDiagram
participant I as Initiator
participant B as Bus (interconnect)
participant T as Target (RAM / scratch)
I->>B: b_transport(trans @ global addr, delay)
B->>B: lookup region in memory map
B->>B: save global, set_address(global - base)
B->>B: delay += hop_delay
B->>T: b_transport(trans @ local addr, delay)
T->>T: service access, set response status
T-->>B: return (same payload)
B->>B: set_address(global) restore
B-->>I: return (same payload)
I->>I: check status, use data
Read the diagram as a redirection, not a relay: one payload, addressed globally on the way in, addressed locally for the single hop into the target, and addressed globally again on the way out. No clock anywhere — time advances only when someone annotates delay and the initiator pays it.
Beginner: First Principles
The simplest interesting fabric is one initiator routed to two targets. We will build a Bus with a fixed two-entry memory map — RAM at 0x0000_0000 and a scratch peripheral at 0x1000_0000, the memory-map convention this whole section uses — a tiny Memory target reused for both, and a traffic-generator initiator that writes and reads each region and then deliberately misses the map to provoke an address error. Every routing idea from the mental model appears exactly once, so this doubles as the field-by-field tour of an interconnect.
A note on sockets before the code. The bus needs two kinds of socket. Facing the initiator it has a simple_target_socket — the same convenience target socket a memory uses, because from the initiator's point of view the bus is the thing being called. Facing the targets it has a multi_passthrough_initiator_socket: an array of initiator sockets exposed as one named object, bound by repeated .bind() (the first bind becomes index 0, the second index 1) and called by index, out[i]->b_transport(...). That one socket replaces the N separate initiator-socket declarations a naive bus would need and scales to any number of targets. What these convenience sockets wrap — the raw forward/backward interfaces — is the subject of Part 3; here we use them as the binding points.
// file: bus_decode.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
// -I$SYSTEMC_HOME/include bus_decode.cpp \
// -o bus_decode -L$SYSTEMC_HOME/lib -lsystemc
// Run: DYLD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./bus_decode
#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <tlm_utils/multi_passthrough_initiator_socket.h>
#include <iostream>
#include <iomanip>
// A simple memory target: SIZE bytes, addressed from local offset 0.
SC_MODULE(Memory) {
tlm_utils::simple_target_socket<Memory> socket;
unsigned int size;
unsigned char* mem;
std::string label;
Memory(sc_module_name nm, unsigned int sz, const char* lbl)
: sc_module(nm), socket("socket"), size(sz), label(lbl) {
mem = new unsigned char[size];
for (unsigned i = 0; i < size; i++) mem[i] = 0;
socket.register_b_transport(this, &Memory::b_transport);
}
~Memory() { delete[] mem; }
void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
sc_dt::uint64 addr = trans.get_address();
unsigned int len = trans.get_data_length();
unsigned char* ptr = trans.get_data_ptr();
// The address we receive is already target-LOCAL (the bus subtracted base).
if (addr + len > size) {
trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE);
return;
}
if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
for (unsigned i = 0; i < len; i++) mem[addr + i] = ptr[i];
else
for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];
std::cout << " [" << label << "] "
<< (trans.get_command() == tlm::TLM_WRITE_COMMAND ? "WR" : "RD")
<< " local=0x" << std::hex << addr
<< " len=" << std::dec << len << "\n";
trans.set_response_status(tlm::TLM_OK_RESPONSE);
}
};
// One-initiator router: decode by fixed memory map, localize, forward, restore.
SC_MODULE(Bus) {
tlm_utils::simple_target_socket<Bus> in; // from initiator
tlm_utils::multi_passthrough_initiator_socket<Bus> out; // to targets
// Memory map: { base, size, target index, name }.
struct Region { sc_dt::uint64 base, size; int idx; const char* name; };
Region map_[2] = {
{ 0x00000000ull, 0x10000, 0, "RAM" }, // 64 KB at 0x0000_0000
{ 0x10000000ull, 0x1000, 1, "SCRATCH" }, // 4 KB at 0x1000_0000
};
SC_CTOR(Bus) : in("in"), out("out") {
in.register_b_transport(this, &Bus::b_transport);
}
void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
sc_dt::uint64 global = trans.get_address();
for (auto& r : map_) {
if (global >= r.base && global < r.base + r.size) {
sc_dt::uint64 local = global - r.base; // localize
std::cout << "[BUS] @0x" << std::hex << std::setw(8) << std::setfill('0')
<< global << " -> " << r.name
<< " local=0x" << local << std::dec << "\n";
trans.set_address(local); // forward LOCAL address
out[r.idx]->b_transport(trans, delay);
trans.set_address(global); // restore GLOBAL address
return;
}
}
std::cout << "[BUS] @0x" << std::hex << std::setw(8) << std::setfill('0')
<< global << " -> UNMAPPED" << std::dec << "\n";
trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE);
}
};
SC_MODULE(TrafficGen) {
tlm_utils::simple_initiator_socket<TrafficGen> socket;
SC_CTOR(TrafficGen) : socket("socket") { SC_THREAD(run); }
void access(tlm::tlm_command cmd, sc_dt::uint64 addr, unsigned int& word) {
tlm::tlm_generic_payload trans;
sc_time delay = SC_ZERO_TIME;
trans.set_command(cmd);
trans.set_address(addr);
trans.set_data_ptr(reinterpret_cast<unsigned char*>(&word));
trans.set_data_length(4);
trans.set_streaming_width(4);
trans.set_byte_enable_ptr(nullptr);
trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
socket->b_transport(trans, delay);
std::cout << " result: addr=0x" << std::hex << trans.get_address()
<< " resp=" << trans.get_response_string() << std::dec << "\n";
}
void run() {
unsigned int w;
// 1) write+read RAM at global 0x0000_0010
w = 0xCAFEBABE; access(tlm::TLM_WRITE_COMMAND, 0x00000010, w);
w = 0; access(tlm::TLM_READ_COMMAND, 0x00000010, w);
std::cout << " RAM read-back = 0x" << std::hex << w << std::dec << "\n";
// 2) write+read SCRATCH at global 0x1000_0020
w = 0x12345678; access(tlm::TLM_WRITE_COMMAND, 0x10000020, w);
w = 0; access(tlm::TLM_READ_COMMAND, 0x10000020, w);
std::cout << " SCRATCH read-back = 0x" << std::hex << w << std::dec << "\n";
// 3) miss: nothing maps 0x2000_0000
w = 0xDEAD; access(tlm::TLM_WRITE_COMMAND, 0x20000000, w);
}
};
int sc_main(int, char*[]) {
TrafficGen gen("gen");
Bus bus("bus");
Memory ram("ram", 0x10000, "RAM");
Memory scratch("scratch", 0x1000, "SCRATCH");
gen.socket.bind(bus.in);
bus.out.bind(ram.socket); // becomes out index 0
bus.out.bind(scratch.socket); // becomes out index 1
sc_start();
return 0;
}
Compile and run it before reading the walkthrough. Try to predict each line, paying attention to the difference between the [BUS] global address and the [RAM]/[SCRATCH] local address.
Expected output:
[BUS] @0x00000010 -> RAM local=0x10
[RAM] WR local=0x10 len=4
result: addr=0x10 resp=TLM_OK_RESPONSE
[BUS] @0x00000010 -> RAM local=0x10
[RAM] RD local=0x10 len=4
result: addr=0x10 resp=TLM_OK_RESPONSE
RAM read-back = 0xcafebabe
[BUS] @0x10000020 -> SCRATCH local=0x20
[SCRATCH] WR local=0x20 len=4
result: addr=0x10000020 resp=TLM_OK_RESPONSE
[BUS] @0x10000020 -> SCRATCH local=0x20
[SCRATCH] RD local=0x20 len=4
result: addr=0x10000020 resp=TLM_OK_RESPONSE
SCRATCH read-back = 0x12345678
[BUS] @0x20000000 -> UNMAPPED
result: addr=0x20000000 resp=TLM_ADDRESS_ERROR_RESPONSE
(The SystemC banner prints to stderr; the transcript above is the program's stdout.)
Walk through it. The initiator's first access is a write to global 0x0000_0010. The bus reads that address, walks its two-entry map, and matches the RAM region [0x0000_0000, 0x0001_0000). It prints the route, subtracts the base (0x0000_0010 − 0x0000_0000 = 0x10 — RAM happens to live at base 0, so the local address equals the global one here), writes the local address into the payload, and forwards to out[0] — RAM. RAM sees a local address 0x10, well inside its 64 KB, writes four bytes, and stamps TLM_OK_RESPONSE. Control returns to the bus, which restores the global address 0x10 into the payload, and returns to the initiator. The initiator prints addr=0x10 resp=TLM_OK_RESPONSE — and crucially, the address it sees is the global one it issued, because the bus restored it.
The scratch access is where localization earns its keep. The initiator writes to global 0x1000_0020. The bus matches the scratch region [0x1000_0000, 0x1000_1000) and subtracts the base: 0x1000_0020 − 0x1000_0000 = 0x20. It forwards local 0x20 to out[1] — scratch — which sees a tiny offset 0x20 into its 4 KB, not the enormous 0x1000_0020 that would have overrun a 4 KB target and triggered a spurious address error. The target services it, the bus restores the global 0x1000_0020, and the initiator's result line shows addr=0x10000020 — the global address again. That is the localize-then-restore contract working end to end: the target lives in a clean zero-based space, and the initiator never learns the local offset existed.
The miss closes the loop. The initiator writes to global 0x2000_0000, which no region covers. The bus walks the whole map, matches nothing, prints UNMAPPED, and — having no target to forward to — sets TLM_ADDRESS_ERROR_RESPONSE itself and returns. The initiator's get_response_string() reads TLM_ADDRESS_ERROR_RESPONSE. This is the bus acting as responder of last resort, the one and only situation in which an interconnect writes the response status.
Three operational details are worth fixing in your mind now.
First, the bus forwarded the same payload object — it never copied it. When RAM wrote 0xCAFEBABE into the data buffer and stamped the status, it was writing the exact object the initiator constructed, because the bus passed trans straight through by reference. That is why the read-back values are visible to the initiator after the call: there was always one payload, threaded through initiator → bus → target and back, with the bus only redirecting the call. If you ever feel tempted to deep-copy a payload inside a bus, stop — the response and read data would land in the copy, not in the object the initiator holds.
Second, the address restore is not optional and its omission is silent. Comment out the trans.set_address(global) line on the return path and the data values stay perfectly correct — every read still returns the right bytes — but the initiator's result line for the scratch access would print addr=0x20 instead of addr=0x10000020. Any monitor that keys on the post-call address (for logging, coverage, or scoreboard lookup) would mis-attribute the transaction. The restore is a two-line contract — set local before the forward, set global after — and it is the first thing to check when a working model produces wrong addresses in its logs.
Third, the decode is pure address arithmetic, exactly like the RTL decoder, just reconfigurable. This for loop over map_ is the transaction-level twin of the RTL data-memory address decode — the always_comb priority chain that selects a memory bank by comparing address bits. Same decision, same first-match-wins semantics; the difference is that the RTL decoder is frozen at synthesis and this table is ordinary C++ data you can build and extend at runtime. The Intermediate section makes that table a real std::vector you populate in sc_main.
0x0000_0000, its local and global addresses are numerically equal, which can make localization look like a no-op. The scratch region at 0x1000_0000 is where the subtraction visibly matters — global 0x1000_0020 becomes local 0x20. Always test a router with at least one non-zero-based region, or you will not exercise the localize/restore path at all.Intermediate: How It Really Works
The Beginner bus had one initiator and a fixed two-entry map hardcoded in the module. Real fabrics have several initiators contending for several targets, and a memory map you configure from outside the module rather than baking into it. This section adds both: multiple masters through tagged target sockets, a routing table built as a std::vector in sc_main, and a first encounter with arbitration — including the crucial point about when arbitration actually matters.
Multiple initiators: tagged target sockets
A target socket accepts transactions from whatever is bound to it. If you bind two initiators to one plain simple_target_socket, the bus's single b_transport callback cannot tell which master a given transaction came from — it sees only the payload. For logging, per-master arbitration, or security checks the bus needs to know the origin. The convenience socket that supplies it is simple_target_socket_tagged: it adds an integer tag, bound at registration time, to the callback signature, so the callback becomes b_transport(int tag, payload&, sc_time&). You give each initiator its own tagged target socket, register each with a distinct tag, and the shared callback receives that tag with every call.
The target side uses the same multi_passthrough_initiator_socket as before. The memory map becomes a std::vector<Region> populated by an add_route helper that also guards against overlapping regions — the configuration error that silently produces first-match routing. The lookup is factored into its own function so debug and DMI forwarding can reuse it later.
// file: bus_multi.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
// -I$SYSTEMC_HOME/include bus_multi.cpp \
// -o bus_multi -L$SYSTEMC_HOME/lib -lsystemc
// Run: DYLD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./bus_multi
#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <tlm_utils/multi_passthrough_initiator_socket.h>
#include <iostream>
#include <iomanip>
#include <vector>
#include <string>
SC_MODULE(Memory) {
tlm_utils::simple_target_socket<Memory> socket;
unsigned int size;
unsigned char* mem;
std::string label;
Memory(sc_module_name nm, unsigned int sz, const char* lbl)
: sc_module(nm), socket("socket"), size(sz), label(lbl) {
mem = new unsigned char[size];
for (unsigned i = 0; i < size; i++) mem[i] = 0;
socket.register_b_transport(this, &Memory::b_transport);
}
~Memory() { delete[] mem; }
void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
sc_dt::uint64 addr = trans.get_address();
unsigned int len = trans.get_data_length();
unsigned char* ptr = trans.get_data_ptr();
if (addr + len > size) {
trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE);
return;
}
if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
for (unsigned i = 0; i < len; i++) mem[addr + i] = ptr[i];
else
for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];
trans.set_response_status(tlm::TLM_OK_RESPONSE);
}
};
// M initiators (one tagged target socket each) x N targets.
// Routing table is data-driven: a std::vector of entries + a lookup helper.
SC_MODULE(Bus) {
static const int M = 2;
// One tagged target socket per initiator; the tag (= initiator id) is bound
// at registration so a single shared callback knows who issued the request.
tlm_utils::simple_target_socket_tagged<Bus>* in[M];
tlm_utils::multi_passthrough_initiator_socket<Bus> out;
struct Region { sc_dt::uint64 base, size; int idx; std::string name; };
std::vector<Region> routes;
int rr_last = -1; // round-robin bookkeeping (LT illustration)
unsigned int grants[M] = {0, 0}; // per-initiator grant counters
SC_CTOR(Bus) : out("out") {
for (int i = 0; i < M; i++) {
in[i] = new tlm_utils::simple_target_socket_tagged<Bus>(
("in" + std::to_string(i)).c_str());
in[i]->register_b_transport(this, &Bus::b_transport, /*tag=*/i);
}
}
~Bus() { for (int i = 0; i < M; i++) delete in[i]; }
void add_route(sc_dt::uint64 base, sc_dt::uint64 size,
int idx, const std::string& name) {
for (auto& r : routes) { // overlap guard
bool overlap = (base < r.base + r.size) && (base + size > r.base);
if (overlap) { SC_REPORT_ERROR("Bus", ("overlap: " + name).c_str()); return; }
}
routes.push_back({base, size, idx, name});
}
const Region* lookup(sc_dt::uint64 global) {
for (auto& r : routes)
if (global >= r.base && global < r.base + r.size) return &r;
return nullptr;
}
// tag == initiator id, bound at registration.
void b_transport(int tag, tlm::tlm_generic_payload& trans, sc_time& delay) {
// Round-robin "fairness" bookkeeping: in LT this only records the grant
// order; it does not gate concurrency, because b_transport is atomic here.
rr_last = tag;
grants[tag]++;
sc_dt::uint64 global = trans.get_address();
const Region* r = lookup(global);
if (!r) { trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE); return; }
std::cout << "[BUS] grant M" << tag << " "
<< (trans.get_command() == tlm::TLM_WRITE_COMMAND ? "WR" : "RD")
<< " @0x" << std::hex << std::setw(8) << std::setfill('0') << global
<< " -> " << r->name << std::dec << "\n";
sc_dt::uint64 local = global - r->base;
trans.set_address(local);
out[r->idx]->b_transport(trans, delay);
trans.set_address(global);
}
};
SC_MODULE(Master) {
tlm_utils::simple_initiator_socket<Master> socket;
int id;
sc_dt::uint64 target_addr;
unsigned int pattern;
sc_time offset;
Master(sc_module_name nm, int id_, sc_dt::uint64 addr,
unsigned int pat, sc_time off)
: sc_module(nm), socket("socket"), id(id_), target_addr(addr),
pattern(pat), offset(off) {
SC_THREAD(run);
}
void run() {
wait(offset); // stagger starts so masters interleave
for (int n = 0; n < 2; n++) {
tlm::tlm_generic_payload trans;
sc_time delay = SC_ZERO_TIME;
unsigned int w = pattern + n;
trans.set_command(tlm::TLM_WRITE_COMMAND);
trans.set_address(target_addr + n * 4);
trans.set_data_ptr(reinterpret_cast<unsigned char*>(&w));
trans.set_data_length(4);
trans.set_streaming_width(4);
trans.set_byte_enable_ptr(nullptr);
trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
socket->b_transport(trans, delay);
wait(20, SC_NS);
}
}
};
int sc_main(int, char*[]) {
Master m0("m0", 0, 0x00000000, 0xA000, sc_time(0, SC_NS)); // -> RAM
Master m1("m1", 1, 0x10000000, 0xB000, sc_time(10, SC_NS)); // -> SCRATCH
Bus bus("bus");
Memory ram("ram", 0x10000, "RAM");
Memory scratch("scratch", 0x1000, "SCRATCH");
m0.socket.bind(*bus.in[0]); // tag 0
m1.socket.bind(*bus.in[1]); // tag 1
bus.out.bind(ram.socket); // idx 0
bus.out.bind(scratch.socket); // idx 1
bus.add_route(0x00000000, 0x10000, 0, "RAM");
bus.add_route(0x10000000, 0x1000, 1, "SCRATCH");
sc_start();
std::cout << "grants: M0=" << bus.grants[0] << " M1=" << bus.grants[1]
<< " (last=M" << bus.rr_last << ")\n";
return 0;
}
Expected output:
[BUS] grant M0 WR @0x00000000 -> RAM
[BUS] grant M1 WR @0x10000000 -> SCRATCH
[BUS] grant M0 WR @0x00000004 -> RAM
[BUS] grant M1 WR @0x10000004 -> SCRATCH
grants: M0=2 M1=2 (last=M1)
Read the interleaving. m0 starts at time 0 and writes to RAM; m1 starts 10 ns later and writes to scratch; each waits 20 ns between its two writes. The result is a time-ordered interleave — M0, M1, M0, M1 — and the bus's single tagged callback correctly attributes each transaction to its origin via the tag argument. The grant counters confirm both masters were served twice. The tag is what makes a shared callback work for distinct masters: without it, the four transactions would be indistinguishable inside b_transport.
The structural pattern here is the one every SoC-level virtual platform uses: an array of tagged target sockets on the master side (one per initiator), a multi_passthrough_initiator_socket on the target side (one indexed object for all targets), and a routing table built in sc_main. The bus module knows nothing about specific addresses or specific masters at compile time; the topology is configured at elaboration. Move the scratch peripheral from 0x1000_0000 to 0x4000_0000 and exactly one line changes — the add_route call — with no edit to the bus module and no recompilation of its logic.
A word on arbitration — and when it actually matters
The rr_last and grants[] bookkeeping above records which master the bus served, but notice what it does not do: it does not make any master wait for another. It cannot, and it does not need to. Under loosely-timed b_transport, each transaction runs to completion atomically on the calling master's stack before control returns and the next master's thread gets a turn. There is no moment when two transactions are simultaneously in flight competing for the target, so there is nothing to arbitrate for functional correctness. Routing alone is the whole job; the grant counters are pure observability.
True arbitration — round-robin that actually stalls a loser, priority, QoS — changes observable behavior only when requests genuinely overlap in time and contend for a shared resource, and that overlap is a property of the approximately-timed (AT) coding style, where a transaction is split into phases that return control before completion. There, two requests really can be outstanding at once, the order in which the fabric grants them changes latencies and throughput, and an arbiter earns its place. If you need real arbitration, you need AT — covered in Part 4. For an LT fabric, model the grant policy for realism if you like, but know that it is annotation and logging, not a correctness mechanism.
Advanced: Edge Cases & LRM Corners
The Beginner and Intermediate buses cover what most LT fabrics do: decode, localize, forward, restore, with multiple masters. This section is the part tutorials skip — the precise rules for what an interconnect may legally modify, fabrics nested behind fabrics, forwarding the other two transport methods (debug and DMI) through the same decode, and modeling the fabric's own latency. One program demonstrates all four; we build it up, then dissect.
Corner 1: what an interconnect may legally modify
State the rule exactly, because it is the contract that keeps a fabric conformant. On the forward path an interconnect may modify exactly two payload attributes: the address (for decoding/localization) and the DMI-allowed hint. It must not modify the command, data pointer, data length, byte enables, or streaming width. If it modifies the address it must restore the original before the transaction returns to the initiator. It must not modify the response status — that belongs to the ultimate target — with the single exception of stamping an error itself when no target maps the address. Everything the Beginner bus did is within these bounds: it rewrote the address twice (localize, restore) and otherwise forwarded the payload untouched. A bus that "helpfully" clamped a length, swapped a command, or rewrote a target's TLM_OK_RESPONSE into something else would be violating the base protocol, and any checker or downstream interconnect that trusts the protocol would be misled. The discipline is conservative on purpose: redirect the call, rewrite only the address, leave the rest alone.
Corner 2: nested fabrics — a bus behind a bus
Because a bus is both a target (to its upstream) and an initiator (to its downstream), nothing stops a bus's downstream target from being another bus. This is how real hierarchical interconnects are built: a top-level fabric routes a coarse address range to a sub-fabric, which routes finely among its own targets. The localize/restore contract composes cleanly: each level subtracts its own region's base on the way in and restores on the way out, so the address is correct at every boundary and globally correct when it finally returns to the initiator. The example below puts an outer bus in front of an inner bus; the outer forwards the entire low address range to the inner as a single target, and the inner does the fine RAM/scratch decode.
Corner 3: forwarding debug and DMI through the router
b_transport is not the only call that arrives at a bus. Two others must be forwarded through the same decode logic, each with its own rules:
transport_dbg— debug transport. A debugger or loader uses it to peek/poke memory without perturbing the model: it consumes no simulation time, carries nosc_time& delay, and must not alter state beyond the access itself. The bus forwards it exactly likeb_transport— decode, localize, forward, restore the global address — but returns the target's byte count and never touchesdelay(there is none). Forwardingtransport_dbgis what lets a debugger reach through the fabric to inspect any target by its global address. The debug-transport semantics (why it must be side-effect-free, how loaders use it) are the subject of Part 7; here we just wire its forwarding into the router.
get_direct_mem_ptr(DMI) — the direct-memory-interface request. An initiator asks "can I have a raw pointer to the backing store for this address, so I can bypassb_transportfor a hot region?" The bus decodes and localizes the requested address, forwards the DMI request to the target, and — this is the rule people miss — rebases the returned region back into global address space before returning it. The target reports its DMI window in local addresses (it lives at base 0); the initiator caches the pointer keyed by global addresses. So the bus must add its region base back to the returnedstart_address/end_address. Get this wrong and the initiator caches a window labeled with local addresses, then uses it for global accesses, and reads the wrong bytes. The full DMI lifecycle (granting, using, and invalidating pointers) is Part 5 material; the forwarding-and-rebase rule is what belongs in the fabric.
Corner 4: modeling the fabric's own latency
A real interconnect adds traversal latency — arbitration, pipeline stages, wire delay. The bus models its hop cost the same way any component does in LT: it adds to the sc_time& delay annotation and returns immediately. It must never call wait() — it runs on the initiator's call stack, and a wait() there is illegal (the master is already suspended inside its own b_transport). So the bus does delay += hop_delay before forwarding, and the hop cost accumulates alongside the target's own latency in the same annotation the initiator eventually pays. In a nested fabric each level adds its own hop, so the total annotated delay is the sum of every hop plus the final target's access latency — which is exactly what a multi-level interconnect should cost.
Here is the program that demonstrates all four corners: an outer bus, an inner bus, RAM and scratch with distinct access latencies, a timed write that crosses both fabrics, a zero-time debug read-back, and a DMI request whose window comes back in global space.
// file: bus_advanced.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
// -I$SYSTEMC_HOME/include bus_advanced.cpp \
// -o bus_advanced -L$SYSTEMC_HOME/lib -lsystemc
// Run: DYLD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./bus_advanced
#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <tlm_utils/multi_passthrough_initiator_socket.h>
#include <iostream>
#include <iomanip>
#include <vector>
#include <string>
// Memory target supporting b_transport, transport_dbg, and DMI.
SC_MODULE(Memory) {
tlm_utils::simple_target_socket<Memory> socket;
unsigned int size;
unsigned char* mem;
std::string label;
sc_time latency;
Memory(sc_module_name nm, unsigned int sz, const char* lbl, sc_time lat)
: sc_module(nm), socket("socket"), size(sz), label(lbl), latency(lat) {
mem = new unsigned char[size];
for (unsigned i = 0; i < size; i++) mem[i] = 0;
socket.register_b_transport(this, &Memory::b_transport);
socket.register_transport_dbg(this, &Memory::transport_dbg);
socket.register_get_direct_mem_ptr(this, &Memory::get_dmi);
}
~Memory() { delete[] mem; }
void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
sc_dt::uint64 addr = trans.get_address();
unsigned int len = trans.get_data_length();
unsigned char* ptr = trans.get_data_ptr();
if (addr + len > size) {
trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE);
return;
}
if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
for (unsigned i = 0; i < len; i++) mem[addr + i] = ptr[i];
else
for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];
delay += latency; // target's own access latency
trans.set_response_status(tlm::TLM_OK_RESPONSE);
}
// Debug transport: side-effect-free, consumes NO simulation time.
unsigned int transport_dbg(tlm::tlm_generic_payload& trans) {
sc_dt::uint64 addr = trans.get_address();
unsigned int len = trans.get_data_length();
unsigned char* ptr = trans.get_data_ptr();
if (addr >= size) return 0;
unsigned int n = std::min<unsigned int>(len, size - addr);
if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
for (unsigned i = 0; i < n; i++) mem[addr + i] = ptr[i];
else
for (unsigned i = 0; i < n; i++) ptr[i] = mem[addr + i];
return n; // bytes actually transferred
}
// DMI: hand back a raw pointer to the backing store for this LOCAL range.
bool get_dmi(tlm::tlm_generic_payload& trans, tlm::tlm_dmi& dmi) {
dmi.allow_read_write();
dmi.set_dmi_ptr(mem);
dmi.set_start_address(0); // local address space
dmi.set_end_address(size - 1);
dmi.set_read_latency(latency);
dmi.set_write_latency(latency);
return true;
}
};
// Router: decode, localize, forward b_transport / transport_dbg / DMI, add hop.
SC_MODULE(Bus) {
tlm_utils::simple_target_socket<Bus> in;
tlm_utils::multi_passthrough_initiator_socket<Bus> out;
struct Region { sc_dt::uint64 base, size; int idx; std::string name; };
std::vector<Region> routes;
sc_time hop_delay;
SC_CTOR(Bus) : in("in"), out("out"), hop_delay(2, SC_NS) {
in.register_b_transport(this, &Bus::b_transport);
in.register_transport_dbg(this, &Bus::transport_dbg);
in.register_get_direct_mem_ptr(this, &Bus::get_dmi);
}
void add_route(sc_dt::uint64 base, sc_dt::uint64 size,
int idx, const std::string& name) {
routes.push_back({base, size, idx, name});
}
const Region* lookup(sc_dt::uint64 g) {
for (auto& r : routes) if (g >= r.base && g < r.base + r.size) return &r;
return nullptr;
}
void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
sc_dt::uint64 g = trans.get_address();
const Region* r = lookup(g);
if (!r) { trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE); return; }
delay += hop_delay; // model fabric traversal cost
trans.set_address(g - r->base); // localize (legal: address mutable)
out[r->idx]->b_transport(trans, delay);
trans.set_address(g); // restore global address
}
// Debug forwarding: same decode/localize, but no time, no delay arg.
unsigned int transport_dbg(tlm::tlm_generic_payload& trans) {
sc_dt::uint64 g = trans.get_address();
const Region* r = lookup(g);
if (!r) return 0;
trans.set_address(g - r->base);
unsigned int n = out[r->idx]->transport_dbg(trans);
trans.set_address(g);
return n;
}
// DMI forwarding: localize the requested address, ask the target, then
// rebase the returned window back into GLOBAL space for the initiator.
bool get_dmi(tlm::tlm_generic_payload& trans, tlm::tlm_dmi& dmi) {
sc_dt::uint64 g = trans.get_address();
const Region* r = lookup(g);
if (!r) return false;
trans.set_address(g - r->base);
bool ok = out[r->idx]->get_direct_mem_ptr(trans, dmi);
trans.set_address(g);
if (ok) { // shift local window to global
dmi.set_start_address(dmi.get_start_address() + r->base);
dmi.set_end_address(dmi.get_end_address() + r->base);
}
return ok;
}
};
SC_MODULE(TrafficGen) {
tlm_utils::simple_initiator_socket<TrafficGen> socket;
SC_CTOR(TrafficGen) : socket("socket") { SC_THREAD(run); }
void run() {
unsigned int w;
sc_time delay = SC_ZERO_TIME;
// 1) timed write through TWO fabrics (nested) to SCRATCH at 0x1000_0010.
tlm::tlm_generic_payload t;
w = 0xFEEDFACE;
t.set_command(tlm::TLM_WRITE_COMMAND);
t.set_address(0x10000010);
t.set_data_ptr(reinterpret_cast<unsigned char*>(&w));
t.set_data_length(4); t.set_streaming_width(4);
t.set_byte_enable_ptr(nullptr);
t.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
std::cout << "[" << sc_time_stamp() << "] issue write\n";
socket->b_transport(t, delay);
std::cout << " resp=" << t.get_response_string()
<< " annotated delay=" << delay << "\n";
wait(delay); delay = SC_ZERO_TIME;
std::cout << "[" << sc_time_stamp() << "] write paid\n";
// 2) DEBUG read-back through both fabrics (zero time).
unsigned int dbg = 0;
tlm::tlm_generic_payload d;
d.set_command(tlm::TLM_READ_COMMAND);
d.set_address(0x10000010);
d.set_data_ptr(reinterpret_cast<unsigned char*>(&dbg));
d.set_data_length(4);
unsigned int n = socket->transport_dbg(d);
std::cout << "[" << sc_time_stamp() << "] dbg read " << n
<< " bytes = 0x" << std::hex << dbg << std::dec << "\n";
// 3) DMI request: get a global-space pointer window for the same address.
tlm::tlm_generic_payload dm;
dm.set_command(tlm::TLM_READ_COMMAND);
dm.set_address(0x10000010);
tlm::tlm_dmi dmi;
bool ok = socket->get_direct_mem_ptr(dm, dmi);
std::cout << " DMI ok=" << ok
<< " global window [0x" << std::hex << dmi.get_start_address()
<< ", 0x" << dmi.get_end_address() << "]" << std::dec << "\n";
}
};
int sc_main(int, char*[]) {
TrafficGen gen("gen");
// Nested fabric: gen -> outer bus -> inner bus -> {RAM, SCRATCH}
Bus outer("outer");
Bus inner("inner");
Memory ram("ram", 0x10000, "RAM", sc_time(5, SC_NS));
Memory scratch("scratch", 0x1000, "SCRATCH", sc_time(3, SC_NS));
gen.socket.bind(outer.in);
// outer forwards the whole low 512 MB to the inner bus as target 0.
outer.out.bind(inner.in);
outer.add_route(0x00000000, 0x20000000, 0, "INNER");
inner.out.bind(ram.socket); // idx 0
inner.out.bind(scratch.socket); // idx 1
inner.add_route(0x00000000, 0x10000, 0, "RAM");
inner.add_route(0x10000000, 0x1000, 1, "SCRATCH");
sc_start();
return 0;
}
Expected output:
[0 s] issue write
resp=TLM_OK_RESPONSE annotated delay=7 ns
[7 ns] write paid
[7 ns] dbg read 4 bytes = 0xfeedface
DMI ok=1 global window [0x10000000, 0x10000fff]
Read the four results against the four corners. The timed write crosses the outer bus (+2 ns hop), then the inner bus (+2 ns hop), then lands in scratch (+3 ns access latency) — total annotated delay 7 ns, which the initiator pays with wait(delay), advancing simulation time from 0 s to 7 ns. That is latency modeling composing across a nested fabric: each hop adds to the same annotation, and the sum is exactly what a two-level interconnect to a 3 ns target should cost. Note the outer bus's route base is 0x0000_0000, so its localization is a numeric no-op, but the inner bus genuinely subtracts 0x1000_0000 to turn the global address into a scratch-local 0x10 — and both restore on the way out, so the write completes correctly.
The debug read-back runs through both fabrics' transport_dbg forwarders in zero simulation time — sc_time_stamp() still reads 7 ns after it, unchanged — and returns 4 bytes equal to 0xfeedface, proving the write landed and that debug transport reaches the right target through the same nested decode without perturbing the clock.
The DMI request is the subtle one. Scratch reports its window in local addresses [0x0, 0xfff] (it lives at base 0 and is 4 KB). The inner bus rebases it by adding scratch's base 0x1000_0000, producing [0x1000_0000, 0x1000_0fff]; the outer bus rebases by adding its route base 0x0, leaving it unchanged. The initiator receives a window labeled in global addresses [0x1000_0000, 0x1000_0fff] — exactly the global range it would use for direct accesses. That global-space window is the whole point of the rebase: a pointer the initiator can index by the same global addresses it uses everywhere else. Forget the rebase at either level and the initiator would cache [0x0, 0xfff] and read RAM-base bytes for scratch addresses.
wait() inside any of its forwarding callbacks. It runs on the initiator's call stack — the initiator is already suspended inside its own b_transport — so a wait() in the bus attempts to re-suspend an already-blocked thread, a SystemC runtime error. Model every fabric latency by adding to the delay annotation and returning, exactly as the hop_delay line does. This is the same rule as "targets annotate, initiators pay" from Part 2, extended to interconnects.Version differences
The interconnect model, the address-mutability rule, the tagged and multi-passthrough convenience sockets, and the debug/DMI forwarding semantics are identical across SystemC 2.3.0, 2.3.1, 2.3.3, 2.3.4, and 3.0.x. The tlm and tlm_utils headers ship bundled with the kernel from 2.3.0 onward. Every example here was compiled and run against an installed SystemC 3.0.1 build with a C++17 compiler; on 3.0.x you compile with -DSC_ALLOW_DEPRECATED_IEEE_API to silence deprecation notes for the classic SC_MODULE/SC_CTOR style used here, with no change to behavior. The tagged-socket registration signature — register_b_transport(module, callback, tag) — and the multi_passthrough_initiator_socket bind-then-index pattern are stable across all these releases.
Hands-on exercise
Extend the Beginner router into a three-target fabric with overlapping-region detection.
Start from the bus_decode.cpp Beginner program. Replace its fixed two-entry map_ array with a runtime routing table — a std::vector<Region> populated by an add_route(base, size, idx, name) helper called in sc_main, in the style of the Intermediate bus. Then:
- Add a third target. Introduce a second peripheral — call it
TIMER— at base0x2000_0000, size 256 bytes, bound asoutindex 2. Route it with a thirdadd_routecall. Drive a write+read pair to a global address inside it (say0x2000_0010) and confirm the round-trip, with the bus localizing0x2000_0010to0x10before forwarding.
- Add overlapping-region detection. Make
add_routereject any new region whose[base, base+size)window intersects an already-registered region, the way the Intermediateadd_routedid. The intersection test is(base < e.base + e.size) && (base + size > e.base)for each existing entrye. On an overlap, callSC_REPORT_ERROR(or print a diagnostic and skip the insertion) naming both colliding regions.
- Prove the detection fires. After successfully registering RAM, scratch, and the timer, attempt to register a fourth region that overlaps scratch — for example base
0x1000_0800, size0x1000(which straddles the end of the 4 KB scratch window at0x1000_0000–0x1000_0fff). Confirm your overlap check catches it before the bad entry corrupts the map.
Then exercise the finished fabric: write distinct sentinel values to RAM, scratch, and the timer, read all three back, and confirm none contaminates another — the same non-cross-contamination check the old RTL-era bus testbench made, now driven through your runtime table. Predict the localized address the bus forwards for each global access before you run, then check the trace.
When that works, extend it once more: make the timer a read-only region by having its target return TLM_COMMAND_ERROR_RESPONSE for any TLM_WRITE_COMMAND, and have the initiator confirm a write to the timer is rejected while a read succeeds. This exercises the response-status path through the bus — the bus forwards the target's error untouched, exactly as it forwards TLM_OK_RESPONSE.
Hints
- Reuse the Beginner
Memorytarget verbatim for RAM, scratch, and the timer — only the sizes and the routing-table entries differ. The target never knows its base; it always sees local addresses. - The overlap test belongs only in
add_route, beforeroutes.push_back(...). It runs once at elaboration, never on the hotb_transportpath. Do not put it in the routing loop. - Bind in the same order you
add_route: the firstout.bind(...)is index 0, so theidxyou pass toadd_routemust match the bind order. A mismatch routes correctly-decoded transactions to the wrong target with no error — verify the order in one place insc_main. - For the read-only timer, put the command check at the very top of its
b_transport, setTLM_COMMAND_ERROR_RESPONSE, andreturnbefore touching storage — the same early-return discipline as the address-range check. - Keep the localize/restore pair around every forward, including the timer's. The third region at
0x2000_0000is where forgetting the restore would show up as a wrong post-call address for timer transactions. - To see the overlap detection without aborting the whole run, either print-and-skip in
add_route, or wrap the deliberate bad registration withsc_report_handler::set_actionsto downgrade the error to a log for that one call.
No solution is provided. The understanding lives in getting the runtime table, the localize/restore discipline, and the overlap guard right yourself.
Common mistakes
- Forgetting to restore the global address after localizing. The bus subtracts the region base before forwarding but neglects to write the original address back on the return path. Data values stay correct — every read returns the right bytes — so the bug passes functional tests, but the initiator and every monitor above the bus see the local offset when they inspect
get_address()after the call, corrupting logs, coverage, and scoreboards. Fix: localize before the forward, restore immediately after, as a two-line pair around everyb_transport/transport_dbg/DMI forward.
- Deep-copying the payload inside the bus. Treating the payload like a value to be copied to the target breaks the response path: the target writes its status and read data into the copy, not the object the initiator holds, so the initiator sees a stale
TLM_INCOMPLETE_RESPONSEand garbage data. Fix: forward the same payload object by reference, exactly asb_transportpasses it. An interconnect redirects a call; it never duplicates the transaction.
- Believing an LT bus needs an arbiter to be correct. Adding a stalling arbiter to a loosely-timed
b_transportfabric models contention that does not exist — each transaction already completes atomically before the next begins, so there is nothing in flight to arbitrate. The result is wasted complexity and sometimes deadlock. Fix: route in LT; reach for real arbitration only at AT, where transactions overlap in time. Grant counters for observability are fine; stalling logic is not.
- Disabling address translation because "it breaks DMI." Translation does not break DMI — forgetting to rebase the DMI window does. The target reports its window in local addresses; the bus must add the region base back so the initiator caches a global-space pointer. Fix: localize the DMI request's address on the way in, and shift the returned
start_address/end_addressby the region base on the way out. The same applies toinvalidate_direct_mem_ptron the backward path.
- Letting the
add_routeindex drift out of sync with the bind order.add_route(2, ...)declares that index 2 handles a region, but ifout.bind(...)for that target happened at a different position, transactions decode correctly and then route to the wrong target with no error or warning. Fix: bind andadd_routein the same order, in the same block ofsc_main, and use named index constants so the correspondence is visible.
- Calling
wait()inside a bus callback. The bus runs on the initiator's call stack while the initiator is suspended inside its ownb_transport; await()there tries to re-suspend an already-blocked thread, a SystemC runtime error. Fix: model fabric latency by adding to thesc_time& delayannotation (delay += hop_delay) and returning immediately. Targets annotate, interconnects annotate, initiators pay.
- Ignoring overlapping regions in the memory map. Nothing in TLM-2.0 checks the map for you; two overlapping entries silently produce first-match routing, so accesses to the overlap region reach whichever entry was registered first while the other is shadowed. Fix: add an overlap check to
add_routethat rejects any region intersecting an existing one, and treat overlap as the configuration error it is.
Recap
After working through this post you can now:
- State the core abstraction — a bus is an address-indexed function-call router whose memory map is a table of
{base, size, target}tuples — and explain why routing alone is the whole functional job of an LT fabric. - Write a router from scratch with a
simple_target_socketfacing the initiator and amulti_passthrough_initiator_socketfacing the targets, decode an address against the map, and forward the same payload by reference. - Localize an address on the forward path (subtract the region base so the target sees a zero-based offset) and restore the global address on the return path, and explain why the restore is mandatory and its omission silent.
- Return
TLM_ADDRESS_ERROR_RESPONSEfrom the bus on an unmapped access, the one and only case in which an interconnect writes the response status. - Connect multiple initiators through tagged target sockets, use the registration-bound tag to attribute each transaction to its origin, and build a runtime routing table with an overlap-guarded
add_routehelper. - Explain why arbitration changes observable behavior only under approximately-timed modeling, and why an LT bus needs routing but not a stalling arbiter for correctness.
- Name exactly which payload attributes an interconnect may modify (the address, which it must restore, and the DMI hint) and which it must not (command, data pointer, length, byte enables, streaming width, and — except on a miss — response status).
- Nest one fabric behind another, with the localize/restore contract composing cleanly across levels, and model per-hop latency by adding to the timing annotation.
- Forward
transport_dbgthrough the same decode in zero simulation time, and forward DMI requests with the returned window rebased from target-local to global address space.
Further reading
Standards
- IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §10.3.4 (attribute-modification rules: which payload fields an interconnect may change and the address-restore obligation), §10.3.5 (response status), §10.4.2 (
transport_dbgsemantics), §11 (DMI:get_direct_mem_ptr, address rebasing through interconnects, invalidation), and §14 (interconnect components and the forward/backward socket interfaces).
Vendor and consortium documents
- Aynsley (Doulos), OSCI TLM-2.0 Language Reference Manual (JA32) — the canonical narrative treatment of interconnect components, address localization, and the base-protocol forwarding rules, with a worked router example.
- Doulos TLM-2.0 Tutorial — interconnect/router examples and the base-protocol checklist, including the tagged-socket and multi-passthrough-socket patterns.
- Accellera Systems Initiative, SystemC 3.0.0 distribution,
include/tlm_utils/simple_target_socket.h(thesimple_target_socket_taggedvariant and itsregister_b_transport(mod, cb, tag)signature) andinclude/tlm_utils/multi_passthrough_initiator_socket.h— authoritative source for the convenience sockets used here.
Textbooks
- Grötker, Liao, Martin, and Swan, System Design with SystemC — bus and interconnect modeling chapters; the memory-map-as-routing-table framing and the hierarchy of fabrics.
Next in this section
→ Part 7: Memory-Mapped Peripherals & Debug Transport — the semantics behind one of the two forwarding methods this post wired through the router, and the targets that make it matter: register-block peripherals whose reads and writes have side effects (read-to-clear, write-1-to-clear, FIFO-backed data), and why transport_dbg must be side-effect-free and consume no simulation time so a debugger or testbench can peek and poke those registers through the fabric without perturbing them. Read it here: 22. SystemC Tutorial — Memory-Mapped Peripherals & Debug Transport.
Comments (0)
Leave a Comment