17. SystemC Tutorial - The Generic Payload & Blocking Transport

Why this matters

Written 2026-06-05 for the concept-first series.

Every virtual platform you will ever build — the kind that boots an operating system in seconds instead of the hours a pin-accurate RTL model would take — moves data between components as transactions, not as wires wiggling on a clock edge. A transaction is the whole "read 64 bytes from address 0x1000_0040" idea expressed as a single object handed from one model to another, in one function call, with no notion of clocks, no individual address/data/strobe pins, and no per-cycle handshake. The component that issues the request is the initiator; the component that services it is the target; and the object they pass between them is, in TLM-2.0, almost always a tlm_generic_payload. If you understand that one struct and the one function that carries it, you understand the load-bearing 80% of every transaction-level model in the SystemC ecosystem — every memory model, every bus fabric, every memory-mapped peripheral, and every processor front-end in every commercial and open-source virtual platform.

This is the post where the abstraction either clicks or it does not. The generic payload is deceptively simple to look at — it is "just a struct with some fields" — and deceptively easy to get wrong in ways that produce models that appear to work and silently corrupt data, drop error reports, or race. Forget to set the response status and your initiator cannot tell success from failure. Read the return data from the wrong process and you have a race the simulator will not warn you about. Reuse a payload without resetting a stale field and a byte-enable mask from a previous transaction silently masks the next one. Annotate timing the wrong way and your "loosely-timed" model violates causality. This post teaches the payload field by field, who is allowed to set each field, the exact semantics of the b_transport call that carries it, how the sc_time delay argument models latency without a clock, the response-status discipline that separates correct models from broken ones, and — in the Advanced section — the reference-counted memory management you need when a payload outlives a single call. By the end you will write an initiator/target pair from scratch, name every field you set and who owns it, and recognize the three most common payload bugs on sight. That is the bar, and the six posts after this one in the section assume you have cleared it.

Prerequisites

  • Part 1 — Why TLM Exists. You need the motivating picture from Part 1: why transaction-level modeling trades pin-and-cycle accuracy for simulation speed, the distinction between loosely-timed (LT) and approximately-timed (AT) coding styles, and the role of initiators, targets, and sockets at a high level. This post makes that picture concrete.
  • Part 4 — Processes & Sensitivity (SC_METHOD vs SC_THREAD vs SC_CTHREAD). The initiators in this post are SC_THREAD processes, because they need to call wait() to consume simulation time between transactions. You need to be comfortable with how an SC_THREAD is registered, why it can block where an SC_METHOD cannot, and what wait(sc_time) does to simulation time.
  • Part 6 — Your First SystemC Testbench. You need to be fluent in sc_main, instantiating modules, binding them together, and reading sc_time_stamp() output. Every example here is a complete sc_main program you compile and run.
  • SystemC 2.3.x or newer with TLM-2.0 headers (bundled since 2.3.0). The tlm and tlm_utils headers ship inside the SystemC distribution from version 2.3.0 onward — you do not download TLM-2.0 separately. All examples in this post compile with a C++17 compiler against any 2.3.x or 3.0.x build. If your installation predates 2.3.0, upgrade before continuing.

Three ideas from Part 1 get heavy use here: a transaction is the unit of communication (not a pin event); an initiator drives transactions into a target; and sockets are the binding points that connect them. If any of those is hazy, re-read Part 1 before continuing — the rest of this post builds directly on them.

Mental model (first principles)

Start with the most reductive correct statement of what a TLM-2.0 transaction is, and build up from there.

A transaction is a function call that carries a struct. That is the whole thing. When an initiator wants a target to do something — read some bytes, write some bytes — it does not toggle pins over several clock cycles. It fills in a struct describing the request and calls a function on the target, passing the struct by reference. The target's function reads the struct, does the work, writes any results back into the same struct, and returns. Control flow goes into the target and comes back to the initiator, exactly like any C++ method call. The return of the call is the completion of the transaction.

Think of it as a courier delivering a work order. The initiator is a company that needs something done. It fills out a single work-order form — what operation, at what address, pointing at which buffer of bytes, how many bytes, which bytes count — and hands the form to a courier. The courier carries the same physical form to the target (a warehouse, say). The warehouse reads the form, does the work (pulls the requested bytes off a shelf into the buffer the form points at, or files the bytes from the buffer onto a shelf), stamps the form with a result ("done OK" or "no such shelf"), and the courier carries the same form back. The initiator reads the stamp to learn whether it worked. Crucially, the form is never photocopied — there is exactly one form, passed by reference, and both parties write on it. That single-form-passed-by-reference detail is the source of every subtle payload bug, so hold onto it.

In TLM-2.0 the "form" is tlm::tlm_generic_payload and the "hand it to the courier" call is b_transport. Here is the payload's anatomy, field by field, in the order an initiator typically fills them:

  • command (tlm_command): TLM_READ_COMMAND or TLM_WRITE_COMMAND. Initiator sets it. This is "what operation".
  • address (64-bit): where in the target's space the access lands. Initiator sets it. An interconnect (a bus) is allowed to modify it for decoding and must restore it on the way back, but for a directly-bound initiator/target pair it is just what the initiator wrote.
  • data pointer (unsigned char*): points at the initiator's own buffer. For a write, the target reads bytes from here; for a read, the target writes bytes into here. Initiator sets it. The buffer is owned by the initiator; the target only borrows it for the duration of the call.
  • data length (bytes): how many bytes the transfer covers. Initiator sets it. Note: bytes, not words — a 32-bit access has length 4.
  • byte-enable pointer + byte-enable length: an optional per-byte mask for sub-word access. A non-null pointer of 0xff/0x00 bytes says "this byte participates / this byte is masked". Initiator sets it. nullptr means "all bytes participate".
  • streaming width (bytes): for fixed-address bursts, how far into the data buffer the transfer wraps. For an ordinary contiguous transfer it equals the data length. Initiator sets it.
  • DMI hint (bool): "may this transaction use direct memory interface acceleration?" A pure performance hint, covered in a later part. Initiator sets it (commonly false).
  • response status (tlm_response_status): the result stamp. Target sets it. The initiator must set it to TLM_INCOMPLETE_RESPONSE before the call as a "not done yet" sentinel, and the target must overwrite it with TLM_OK_RESPONSE on success (or an error code on failure). This field is the entire contract of "did it work".

The division of labor is the part people get wrong, so state it as a rule. The initiator owns and sets everything about the request. The target sets exactly two things: the response status (always) and, for a read, the bytes in the data buffer. A target that modifies the address, the command, or the length on the return path is violating the base protocol. An initiator that fails to reset the response status before reusing a payload is shipping a stale stamp.

Now the call itself. b_transport — "blocking transport" — has this signature:

void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay);

Two arguments. The first is the payload, by reference (the work-order form). The second is a timing-annotation delay, also by reference, and it is the one genuinely new idea in this post. There is no clock in this picture, so how does a model express "this access takes 8 ns"? The target annotates the delay: it adds its latency to the delay argument. The caller decides when to actually pay that time — either immediately by calling wait(delay), or by accumulating it and paying later (temporal decoupling, the subject of a later part). "Blocking" means the call is allowed to consume simulation time before returning — the target may wait() internally — and the initiator's process is suspended until the call returns. That is why initiators are SC_THREADs: only a thread can sit inside a blocking call across simulated time.

The lifecycle of one transaction, then, is completely deterministic and clock-free:

  1. The initiator constructs (or reuses) a payload, fills in command, address, data pointer, length, byte enables, streaming width, DMI hint, and sets response status to TLM_INCOMPLETE_RESPONSE.
  2. The initiator calls socket->b_transport(trans, delay). Control transfers into the target's registered b_transport method, with trans referring to the same object the initiator filled in.
  3. The target reads the fields, does the work (copies bytes into or out of the buffer the data pointer names), optionally annotates delay with its latency, and sets the response status to TLM_OK_RESPONSE (or an error).
  4. The target returns. Control comes back to the initiator at the line after the call.
  5. The initiator inspects the response status. If is_response_ok(), the data buffer (for a read) now holds valid bytes; otherwise the transaction failed and the buffer is not to be trusted.

That five-step loop is the entire substance of LT transaction modeling. Everything else in this post is detail layered on it: timing annotation is step 3's delay; response discipline is steps 1 and 5's status field; byte enables are a refinement of step 3's copy; memory management is what changes when the payload must outlive the call. Internalize the loop and the rest is filling in fields.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
sequenceDiagram
    participant I as Initiator (SC_THREAD)
    participant T as Target (b_transport)
    I->>I: fill payload, status = INCOMPLETE
    I->>T: b_transport(trans, delay)
    T->>T: read fields, copy bytes
    T->>T: delay += latency
    T->>T: status = TLM_OK_RESPONSE
    T-->>I: return (same payload)
    I->>I: check is_response_ok(), use data

Read the diagram as the courier round-trip: one form, handed over, stamped, handed back. No clock anywhere — time advances only if the target annotates delay and the initiator pays it.

Beginner: First Principles

The simplest interesting TLM-2.0 system is one initiator talking to one memory. We will build a 256-byte memory target and a traffic-generator initiator that writes one word, reads it back, and checks the result. This is the "hello world" of transaction-level modeling, and every field of the generic payload appears in it exactly once, so it doubles as the field-by-field tour.

A note on sockets before the code. We use tlm_utils::simple_initiator_socket and tlm_utils::simple_target_socket — the convenience sockets from the tlm_utils namespace. They let the target register a plain method as its b_transport handler with one line (socket.register_b_transport(this, &Memory::b_transport)) and let the initiator call through the socket with socket->b_transport(...). What these convenience sockets actually wrap — the raw tlm_initiator_socket / tlm_target_socket and the interface classes underneath — is the subject of Part 3 — Initiator & Target Sockets. For now, treat them as the binding points: bind one initiator socket to one target socket and you have a channel.

// file: gp_first_transaction.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include gp_first_transaction.cpp \
//            -o gp_first_transaction -L$SYSTEMC_HOME/lib -lsystemc
// Run:   LD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./gp_first_transaction

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>
#include <iomanip>

// A 256-byte memory target. It answers b_transport calls synchronously.
SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;

  static const unsigned int SIZE = 256;
  unsigned char mem[SIZE];

  SC_CTOR(Memory) : socket("socket") {
    for (unsigned i = 0; i < SIZE; i++) mem[i] = 0;
    socket.register_b_transport(this, &Memory::b_transport);
  }

  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    tlm::tlm_command  cmd  = trans.get_command();
    sc_dt::uint64     addr = trans.get_address();
    unsigned char*    ptr  = trans.get_data_ptr();
    unsigned int      len  = trans.get_data_length();

    if (cmd == tlm::TLM_WRITE_COMMAND) {
      for (unsigned i = 0; i < len; i++) mem[addr + i] = ptr[i];
    } else {
      for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];
    }
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

// A traffic-generator initiator: writes a word, reads it back, checks the result.
SC_MODULE(TrafficGen) {
  tlm_utils::simple_initiator_socket<TrafficGen> socket;

  SC_CTOR(TrafficGen) : socket("socket") {
    SC_THREAD(run);
  }

  void run() {
    tlm::tlm_generic_payload trans;     // one payload, reused, lives on the stack
    sc_time delay = SC_ZERO_TIME;

    // ----- WRITE 0xDEADBEEF to address 0x10 -----
    unsigned int wdata = 0xDEADBEEF;
    trans.set_command(tlm::TLM_WRITE_COMMAND);
    trans.set_address(0x10);
    trans.set_data_ptr(reinterpret_cast<unsigned char*>(&wdata));
    trans.set_data_length(4);
    trans.set_streaming_width(4);            // contiguous transfer: == data_length
    trans.set_byte_enable_ptr(nullptr);      // no masking: every byte participates
    trans.set_dmi_allowed(false);            // no DMI hint
    trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE); // sentinel: "not done"

    socket->b_transport(trans, delay);

    std::cout << "WRITE 0x" << std::hex << wdata
              << " -> addr 0x" << trans.get_address()
              << " resp=" << trans.get_response_string() << "\n";

    // ----- READ it back -----
    unsigned int rdata = 0;
    trans.set_command(tlm::TLM_READ_COMMAND);
    trans.set_address(0x10);
    trans.set_data_ptr(reinterpret_cast<unsigned char*>(&rdata));
    trans.set_data_length(4);
    trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE); // reset before reuse!

    socket->b_transport(trans, delay);

    std::cout << "READ  addr 0x" << trans.get_address()
              << " -> 0x" << std::hex << rdata
              << " resp=" << trans.get_response_string() << "\n";

    if (trans.is_response_ok() && rdata == 0xDEADBEEF)
      std::cout << "MATCH: round-trip succeeded\n";
    else
      std::cout << "MISMATCH\n";
  }
};

int sc_main(int, char*[]) {
  TrafficGen gen("gen");
  Memory     mem("mem");
  gen.socket.bind(mem.socket);     // initiator socket bound to target socket
  sc_start();
  return 0;
}

Compile and run it before reading the walkthrough. Try to predict each output line.

Expected output:

WRITE 0xdeadbeef -> addr 0x10 resp=TLM_OK_RESPONSE
READ  addr 0x10 -> 0xdeadbeef resp=TLM_OK_RESPONSE
MATCH: round-trip succeeded

Walk through it. The initiator's run() thread starts at simulation time 0. It constructs one tlm_generic_payload on the stack — a single object it will reuse for both the write and the read. For the write, it sets the command to TLM_WRITE_COMMAND, the address to 0x10, points the data pointer at its own local wdata variable, sets the length to 4 (four bytes — this is a 32-bit access), sets streaming width to 4 (a contiguous transfer), clears byte enables (every byte counts), clears the DMI hint, and — critically — sets the response status to TLM_INCOMPLETE_RESPONSE. That last write is not optional. It is the initiator declaring "this transaction has not been serviced yet"; the target will overwrite it.

The call socket->b_transport(trans, delay) transfers control into Memory::b_transport, with trans referring to the exact same object the initiator just filled. The memory reads the command (TLM_WRITE_COMMAND), the address (0x10), the data pointer (pointing at the initiator's wdata), and the length (4). It copies four bytes from the initiator's buffer into mem[0x10..0x13]. Then it sets the response status to TLM_OK_RESPONSE and returns. Control comes back to the initiator, which prints the write line. Note the printed address is 0x10 — the memory did not modify it, exactly as the base protocol requires for a target.

For the read, the initiator reuses the same payload. It changes the command to TLM_READ_COMMAND, repoints the data pointer at a fresh local rdata, and — again critically — resets the response status to TLM_INCOMPLETE_RESPONSE. Forgetting this reset is one of the most common payload bugs: the status would still read TLM_OK_RESPONSE from the previous call, and if this read somehow failed without the target setting an error, the initiator would wrongly believe it succeeded. The memory services the read by copying four bytes from mem[0x10..0x13] into the initiator's rdata, sets TLM_OK_RESPONSE, and returns. The initiator finds rdata == 0xDEADBEEF and is_response_ok() true, so it prints MATCH.

Three operational details are worth fixing in your mind now.

First, the data pointer points at the initiator's memory, and the target borrows it only during the call. The memory never allocates a buffer for the returned data — it writes directly into rdata, a variable owned by the initiator's stack frame. This is the "single form passed by reference" rule in action: there is no copy of the data buffer, only one buffer that both sides touch. It also means the initiator must keep that buffer alive for the duration of the call, which a stack variable in a blocking call trivially does.

Second, b_transport is blocking, so the read of rdata after the call is safe. The initiator reads rdata only on the line after b_transport returns. Because the call is synchronous from the initiator's point of view — control does not come back until the target is done — the data is guaranteed valid at that point. If you tried to read rdata from a different process while the b_transport call was still in flight, you would have a race. Read return data only after the call returns, only in the calling process.

Third, the response status is the only channel for "did it work". There is no exception thrown, no error code returned from the function (it returns void), no out-parameter besides the payload itself. The target communicates success or failure exclusively by writing the response status field, and the initiator learns it exclusively by reading that field. This is why the discipline around it — initiator sets TLM_INCOMPLETE_RESPONSE before, target sets TLM_OK_RESPONSE on success, initiator checks after — is not a style preference but the actual error-handling contract of the protocol. We will see it enforced in the Intermediate section with a target that returns an address error.

Note The delay argument was passed as SC_ZERO_TIME and ignored by the memory in this first example — the transaction completed in zero simulation time. That is a legal but un-annotated LT model. The next section makes the target charge realistic latency.

Intermediate: How It Really Works

The Beginner example completed every transaction in zero simulation time and never failed. Real models do neither. A memory access takes time, an address can be out of range, and a sub-word write must touch only some bytes. This section adds all three — timing annotation, response discipline, and byte enables — each with a complete program you compile and run.

Timing annotation: how a clock-free model expresses latency

A transaction-level model has no clock, yet a memory access plainly takes time. TLM-2.0 expresses that time through the sc_time& delay argument of b_transport. The rule is simple and the consequences are deep: the target annotates its latency by adding to delay; the caller decides when to actually advance simulation time. There are two conformant idioms for the caller, and the difference between them is the difference between a naively-timed model and a temporally-decoupled one.

In idiom A, the initiator pays the delay immediately: after b_transport returns, it calls wait(delay) to advance simulation time by exactly the annotated amount, then resets delay to zero. The model's wall-clock-to-simulation mapping is honest — every access advances time right when it happens. In idiom B, the initiator accumulates the delay across several calls without waiting, then pays the whole sum at once. This is the seed of temporal decoupling: a process runs ahead of simulation time, banking up delay, and only synchronizes occasionally. Idiom B is what makes LT models fast, and it is the subject of the quantum-keeper post later in this section. Here we just show both so the mechanism is concrete.

// file: gp_timing_annotation.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include gp_timing_annotation.cpp \
//            -o gp_timing_annotation -L$SYSTEMC_HOME/lib -lsystemc

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>

// Target that charges 2 ns per byte of access latency, annotated onto `delay`.
SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  static const unsigned int SIZE = 256;
  unsigned char mem[SIZE];

  SC_CTOR(Memory) : socket("socket") {
    for (unsigned i = 0; i < SIZE; i++) mem[i] = 0;
    socket.register_b_transport(this, &Memory::b_transport);
  }

  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    unsigned int   len  = trans.get_data_length();
    unsigned char* ptr  = trans.get_data_ptr();
    sc_dt::uint64  addr = trans.get_address();

    if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
      for (unsigned i = 0; i < len; i++) mem[addr + i] = ptr[i];
    else
      for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];

    // Annotate access latency: 2 ns per byte.
    delay += sc_time(2 * len, SC_NS);
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

SC_MODULE(TrafficGen) {
  tlm_utils::simple_initiator_socket<TrafficGen> socket;
  SC_CTOR(TrafficGen) : socket("socket") { SC_THREAD(run); }

  void run() {
    tlm::tlm_generic_payload trans;
    unsigned int data = 0;

    // ---- Idiom A: realize the annotated delay immediately with wait() ----
    sc_time delay = SC_ZERO_TIME;
    trans.set_command(tlm::TLM_WRITE_COMMAND);
    trans.set_address(0x00);
    trans.set_data_ptr(reinterpret_cast<unsigned char*>(&data));
    trans.set_data_length(4);
    trans.set_streaming_width(4);
    trans.set_byte_enable_ptr(nullptr);
    trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);

    std::cout << "[" << sc_time_stamp() << "] before write\n";
    socket->b_transport(trans, delay);
    wait(delay);                 // consume the 8 ns the target annotated
    delay = SC_ZERO_TIME;        // delay is "spent"
    std::cout << "[" << sc_time_stamp() << "] after write+wait\n";

    // ---- Idiom B: accumulate the delay across two calls, pay once ----
    sc_time acc = SC_ZERO_TIME;
    trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
    socket->b_transport(trans, acc);     // acc -> 8 ns (no wait yet)
    trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
    socket->b_transport(trans, acc);     // acc -> 16 ns (still no wait)
    std::cout << "[" << sc_time_stamp()
              << "] two more writes done, accumulated delay = " << acc << "\n";
    wait(acc);                           // pay both at once
    std::cout << "[" << sc_time_stamp() << "] after paying accumulated delay\n";
  }
};

int sc_main(int, char*[]) {
  TrafficGen gen("gen");
  Memory mem("mem");
  gen.socket.bind(mem.socket);
  sc_start();
  return 0;
}

Expected output:

[0 s] before write
[8 ns] after write+wait
[8 ns] two more writes done, accumulated delay = 16 ns
[24 ns] after paying accumulated delay

Read the timestamps carefully. The first write is a 4-byte access; the target annotates 2 ns × 4 = 8 ns. In idiom A the initiator immediately calls wait(delay), so simulation time jumps from 0 s to 8 ns — the second print line confirms it. Then the initiator does two more 4-byte writes without waiting between them, accumulating delay into acc. After both calls sc_time_stamp() still reads 8 ns — no time has passed, because the initiator never waited — but acc has grown to 16 ns (8 ns per write, two writes). Only when the initiator finally calls wait(acc) does time advance, from 8 ns to 24 ns. The two idioms produce the same total elapsed time (16 ns of write latency in the second phase) but reach it differently: idiom A pays as it goes, idiom B banks and pays in bulk.

The load-bearing insight: the target never calls wait() itself in these LT examples; it only annotates delay. That is what keeps the target reusable across both idioms. A target that called wait() internally would force idiom A on every caller and break temporal decoupling. The convention "targets annotate, initiators pay" is what lets the same memory model plug into a slow honest model and a fast decoupled one without modification.

Response status discipline: how errors are reported

In the Beginner example every access succeeded, so the response status was always TLM_OK_RESPONSE. Real targets reject out-of-range addresses, unsupported commands, and malformed bursts. The mechanism is the response status field, and the discipline has three parts, all mandatory:

  1. The initiator sets TLM_INCOMPLETE_RESPONSE before the call. This is the default-constructed value of a fresh payload anyway, but a reused payload carries the previous call's status, so the initiator must reset it. TLM_INCOMPLETE_RESPONSE means "no target has serviced this yet".
  2. The target sets a definite status before returning. On success, TLM_OK_RESPONSE. On failure, a specific error: TLM_ADDRESS_ERROR_RESPONSE for out-of-range, TLM_COMMAND_ERROR_RESPONSE for an unsupported command, and so on. A target that returns without setting the status leaves it at TLM_INCOMPLETE_RESPONSE, which a conformant initiator must treat as a failure.
  3. The initiator checks the status after the call. is_response_ok() is true only for TLM_OK_RESPONSE; is_response_error() is true for any of the error codes. The initiator must not trust read data unless the response is OK.

The next program builds a memory that range-checks every access and returns TLM_ADDRESS_ERROR_RESPONSE for anything past its 256 bytes. The initiator drives one in-range write and one out-of-range write and reports each.

// file: gp_response_discipline.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include gp_response_discipline.cpp \
//            -o gp_response_discipline -L$SYSTEMC_HOME/lib -lsystemc

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>

SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  static const unsigned int SIZE = 256;
  unsigned char mem[SIZE];

  SC_CTOR(Memory) : socket("socket") {
    for (unsigned i = 0; i < SIZE; i++) mem[i] = 0;
    socket.register_b_transport(this, &Memory::b_transport);
  }

  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    sc_dt::uint64 addr = trans.get_address();
    unsigned int  len  = trans.get_data_length();

    // Range check: this target owns 256 bytes only.
    if (addr + len > SIZE) {
      trans.set_response_status(tlm::TLM_ADDRESS_ERROR_RESPONSE);
      return;                                  // do NOT touch memory
    }
    unsigned char* ptr = trans.get_data_ptr();
    if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
      for (unsigned i = 0; i < len; i++) mem[addr + i] = ptr[i];
    else
      for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];

    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

SC_MODULE(TrafficGen) {
  tlm_utils::simple_initiator_socket<TrafficGen> socket;
  SC_CTOR(TrafficGen) : socket("socket") { SC_THREAD(run); }

  void do_write(sc_dt::uint64 addr) {
    tlm::tlm_generic_payload trans;
    sc_time delay = SC_ZERO_TIME;
    unsigned int data = 0xCAFE;
    trans.set_command(tlm::TLM_WRITE_COMMAND);
    trans.set_address(addr);
    trans.set_data_ptr(reinterpret_cast<unsigned char*>(&data));
    trans.set_data_length(4);
    trans.set_streaming_width(4);
    trans.set_byte_enable_ptr(nullptr);
    trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);

    socket->b_transport(trans, delay);

    if (trans.is_response_error())
      std::cout << "addr 0x" << std::hex << addr
                << " FAILED: " << trans.get_response_string() << "\n";
    else
      std::cout << "addr 0x" << std::hex << addr
                << " OK: " << trans.get_response_string() << "\n";
  }

  void run() {
    do_write(0x10);    // in range
    do_write(0xFE);    // 0xFE + 4 = 0x102 > 256: out of range
  }
};

int sc_main(int, char*[]) {
  TrafficGen gen("gen");
  Memory mem("mem");
  gen.socket.bind(mem.socket);
  sc_start();
  return 0;
}

Expected output:

addr 0x10 OK: TLM_OK_RESPONSE
addr 0xfe FAILED: TLM_ADDRESS_ERROR_RESPONSE

The in-range write at 0x10 succeeds and the target stamps TLM_OK_RESPONSE. The write at 0xFE would touch bytes 0xFE through 0x101, past the 256-byte (0x00–0xFF) memory, so the target returns TLM_ADDRESS_ERROR_RESPONSE without touching memory and the initiator's is_response_error() catches it. Two things to internalize. First, the target sets the error status and returns before any memory access — never partially service a failed transaction. Second, the initiator's behavior is entirely driven by the status field; it has no other way to know the second write failed. If the target had forgotten to set any status, the field would have stayed at the TLM_INCOMPLETE_RESPONSE the initiator set, is_response_error() would still report a failure (because TLM_INCOMPLETE_RESPONSE is not OK), and the model would correctly flag the problem — which is exactly why the "set incomplete before, check after" discipline is robust even against a buggy target.

Byte enables: sub-word access

The final Intermediate refinement is the byte-enable mask, which lets a transaction touch only some bytes of its address range — the transaction-level equivalent of write strobes. The initiator supplies a byte-enable buffer parallel to the data buffer: TLM_BYTE_ENABLED (0xff) for a byte that participates, TLM_BYTE_DISABLED (0x00) for a byte the target must leave untouched. The mask repeats if its length is shorter than the data length. The following program pre-fills a word with a known pattern, then does a sub-word write that updates only byte 1, and reads the word back to prove the other three bytes survived.

// file: gp_byte_enables.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include gp_byte_enables.cpp \
//            -o gp_byte_enables -L$SYSTEMC_HOME/lib -lsystemc

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>
#include <iomanip>

SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  static const unsigned int SIZE = 256;
  unsigned char mem[SIZE];

  SC_CTOR(Memory) : socket("socket") {
    for (unsigned i = 0; i < SIZE; i++) mem[i] = 0;
    // Pre-fill 0x20..0x23 with 0x11 0x22 0x33 0x44.
    mem[0x20] = 0x11; mem[0x21] = 0x22; mem[0x22] = 0x33; mem[0x23] = 0x44;
    socket.register_b_transport(this, &Memory::b_transport);
  }

  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    sc_dt::uint64  addr  = trans.get_address();
    unsigned int   len   = trans.get_data_length();
    unsigned char* ptr   = trans.get_data_ptr();
    unsigned char* be    = trans.get_byte_enable_ptr();
    unsigned int   belen = trans.get_byte_enable_length();

    if (trans.get_command() == tlm::TLM_WRITE_COMMAND) {
      for (unsigned i = 0; i < len; i++) {
        // Honour byte enables: if a byte is disabled, leave memory untouched.
        if (be == nullptr || be[i % belen] == TLM_BYTE_ENABLED)
          mem[addr + i] = ptr[i];
      }
    } else {
      for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];
    }
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

SC_MODULE(TrafficGen) {
  tlm_utils::simple_initiator_socket<TrafficGen> socket;
  SC_CTOR(TrafficGen) : socket("socket") { SC_THREAD(run); }

  void run() {
    tlm::tlm_generic_payload trans;
    sc_time delay = SC_ZERO_TIME;

    // Sub-word WRITE: update only byte 1 of the 4-byte word at 0x20.
    unsigned char wbuf[4] = { 0xAA, 0xBB, 0xCC, 0xDD };
    unsigned char be[4]   = { TLM_BYTE_DISABLED, TLM_BYTE_ENABLED,
                              TLM_BYTE_DISABLED, TLM_BYTE_DISABLED };
    trans.set_command(tlm::TLM_WRITE_COMMAND);
    trans.set_address(0x20);
    trans.set_data_ptr(wbuf);
    trans.set_data_length(4);
    trans.set_streaming_width(4);
    trans.set_byte_enable_ptr(be);
    trans.set_byte_enable_length(4);
    trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
    socket->b_transport(trans, delay);

    // READ the whole word back (no byte enables).
    unsigned char rbuf[4] = {0,0,0,0};
    trans.set_command(tlm::TLM_READ_COMMAND);
    trans.set_address(0x20);
    trans.set_data_ptr(rbuf);
    trans.set_data_length(4);
    trans.set_byte_enable_ptr(nullptr);          // clear byte enables for the read
    trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
    socket->b_transport(trans, delay);

    std::cout << "word @0x20 after masked write: ";
    for (int i = 0; i < 4; i++)
      std::cout << "0x" << std::hex << std::setw(2) << std::setfill('0')
                << (int)rbuf[i] << " ";
    std::cout << "\n";
  }
};

int sc_main(int, char*[]) {
  TrafficGen gen("gen");
  Memory mem("mem");
  gen.socket.bind(mem.socket);
  sc_start();
  return 0;
}

Expected output:

word @0x20 after masked write: 0x11 0xbb 0x33 0x44

The memory started with 0x11 0x22 0x33 0x44 at 0x20. The write carried data 0xAA 0xBB 0xCC 0xDD but a byte-enable mask of {disabled, enabled, disabled, disabled} — only byte 1 may change. After the write, byte 1 went from 0x22 to 0xBB (the enabled byte from the write buffer), and bytes 0, 2, and 3 kept their original 0x11, 0x33, 0x44. That is exactly what the read-back shows. The target honored the mask by skipping memory writes for disabled bytes.

One subtlety worth flagging now, because it is a classic reuse bug: the read transaction explicitly calls set_byte_enable_ptr(nullptr). Because the payload is reused from the write, it still carries the byte-enable pointer from the masked write. If the read did not clear it, the byte-enable buffer (now possibly out of scope or stale) would be consulted on the read too, masking part of the read. Always clear byte enables when reusing a payload for an access that does not want them. This is the same family of bug as forgetting to reset the response status — a stale field from a previous transaction corrupting the next one.

This concludes the Intermediate section. You can now annotate latency two ways, enforce the response-status contract, and do sub-word access with byte enables — the working vocabulary for the LT memory and peripheral models that fill the rest of this section. The Advanced section drills into the one piece of the payload we have deliberately deferred: what happens when a payload must outlive the b_transport call.

Advanced: Edge Cases & LRM Corners

The Beginner and Intermediate sections cover what the great majority of LT models do: stack-allocated payloads, reused within one process, serviced synchronously by b_transport. This section is the part the tutorials skip — payload memory management, why the stack payload stops working when a transaction outlives its call, and a handful of LRM corners that bite senior engineers.

Corner 1: when a stack payload is no longer enough

Every example so far put the payload on the stack of the initiator's run() thread and reused it. That is correct and recommended for loosely-timed models where each transaction completes before the next begins — because b_transport is blocking, the call returns before the initiator reuses the payload, so a single stack object suffices. The stack payload's lifetime trivially covers the call.

That breaks the moment a transaction must outlive the call that created it. Two situations force this. First, the approximately-timed (AT) coding style with non-blocking transport (the next major topic in this section) splits a transaction into phases — request, response — that return control to the initiator before the transaction is complete. The payload must stay alive across those returns, so it cannot live on a stack frame that unwinds. Second, even within LT, a target that defers completion (queues the transaction, services it later from another process) needs the payload to persist after b_transport returns. In both cases the payload must be heap-allocated and its lifetime managed explicitly.

TLM-2.0's answer is reference-counted memory management. A payload can be associated with a memory manager (an object implementing tlm::tlm_mm_interface). The payload then carries a reference count. acquire() increments it; release() decrements it; when the count reaches zero, the payload calls its memory manager's free(), which typically resets the payload and returns it to a pool for reuse. Whoever holds a reference to a payload acquire()s it; whoever is done release()s it; the last releaser triggers cleanup. This is the standard reference-counting pattern, scoped to transactions.

The following program builds a minimal pooling memory manager and an initiator that allocates a payload from the pool for each transaction, acquire()s it, uses it, and release()s it. The pool recycles freed payloads, so even though the initiator issues three transactions, only one payload object is ever allocated.

// file: gp_memory_management.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include gp_memory_management.cpp \
//            -o gp_memory_management -L$SYSTEMC_HOME/lib -lsystemc

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>
#include <vector>

// A trivial pooling memory manager. The payload calls free() (via release())
// when its reference count drops to zero.
class SimplePool : public tlm::tlm_mm_interface {
public:
  tlm::tlm_generic_payload* allocate() {
    tlm::tlm_generic_payload* p;
    if (free_list.empty()) {
      p = new tlm::tlm_generic_payload(this);   // register THIS as the manager
      created++;
    } else {
      p = free_list.back();
      free_list.pop_back();
    }
    return p;
  }
  // Called by release() when the reference count hits zero.
  void free(tlm::tlm_generic_payload* p) override {
    p->reset();                 // clear extensions / byte-enable ptr
    free_list.push_back(p);
  }
  int created = 0;
  ~SimplePool() { for (auto* p : free_list) delete p; }
private:
  std::vector<tlm::tlm_generic_payload*> free_list;
};

SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  unsigned char mem[256];
  SC_CTOR(Memory) : socket("socket") {
    for (int i = 0; i < 256; i++) mem[i] = 0;
    socket.register_b_transport(this, &Memory::b_transport);
  }
  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    sc_dt::uint64  addr = trans.get_address();
    unsigned int   len  = trans.get_data_length();
    unsigned char* ptr  = trans.get_data_ptr();
    if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
      for (unsigned i = 0; i < len; i++) mem[addr + i] = ptr[i];
    else
      for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

SC_MODULE(TrafficGen) {
  tlm_utils::simple_initiator_socket<TrafficGen> socket;
  SimplePool pool;
  SC_CTOR(TrafficGen) : socket("socket") { SC_THREAD(run); }

  void run() {
    unsigned int data[3] = { 0xA1, 0xB2, 0xC3 };
    for (int n = 0; n < 3; n++) {
      tlm::tlm_generic_payload* trans = pool.allocate();
      trans->acquire();                 // ref count: 1 (we hold a reference)
      sc_time delay = SC_ZERO_TIME;
      trans->set_command(tlm::TLM_WRITE_COMMAND);
      trans->set_address(n * 4);
      trans->set_data_ptr(reinterpret_cast<unsigned char*>(&data[n]));
      trans->set_data_length(4);
      trans->set_streaming_width(4);
      trans->set_byte_enable_ptr(nullptr);
      trans->set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
      socket->b_transport(*trans, delay);
      std::cout << "txn " << n << " resp=" << trans->get_response_string() << "\n";
      trans->release();                 // ref count: 0 -> pool.free() recycles it
    }
    std::cout << "payloads actually allocated: " << pool.created << "\n";
  }
};

int sc_main(int, char*[]) {
  TrafficGen gen("gen");
  Memory mem("mem");
  gen.socket.bind(mem.socket);
  sc_start();
  return 0;
}

Expected output:

txn 0 resp=TLM_OK_RESPONSE
txn 1 resp=TLM_OK_RESPONSE
txn 2 resp=TLM_OK_RESPONSE
payloads actually allocated: 1

Three transactions, one payload allocated. The first allocate() finds an empty free list, news a payload (registering the pool as its memory manager) and bumps created to 1. The initiator acquire()s it (count 1), uses it, and release()s it. Because the payload was constructed with a memory manager, release() drops the count to zero and calls pool.free(), which resets the payload and pushes it back on the free list. The next allocate() finds it there and reuses it — created never grows past 1. That is the whole value of pooling: in an AT model issuing millions of transactions, you allocate a handful of payloads and recycle them, instead of new/delete per transaction.

Two rules make reference counting safe. First, only payloads constructed with a memory manager may be acquire()/release()d — calling them on a stack payload with no manager is undefined. The constructor tlm_generic_payload(this) (passing the manager) is what arms reference counting; the default constructor does not. Second, balance every acquire() with exactly one release(). An extra acquire() leaks the payload (it never returns to the pool); an extra release() frees it while someone still holds a reference, producing a use-after-free. The discipline is the same as any reference-counted resource.

Note For loosely-timed b_transport models, you almost never need this. Reach for the stack payload first. Memory management earns its complexity only when payloads outlive their call — AT pipelines, deferred completion, and the analysis ports of a later section. The next post on sockets, and the AT post after it, build on exactly this mechanism.

Corner 2: the payload is one object — aliasing and reuse hazards

Because b_transport passes the payload by reference, the target writes the same object the initiator holds. This is efficient and is the source of a family of bugs. Reusing a payload without resetting a field carries that field forward: a stale response status (looks like success), a stale byte-enable pointer (silently masks the next access, as the Intermediate byte-enable example warned), a stale data length (transfers the wrong number of bytes), or a stale data pointer (writes into a buffer that has since gone out of scope). The reset() method clears the extension array and byte-enable pointer; it does not clear command, address, data pointer, or length, which you must overwrite per transaction. The safe habit is to set every field you care about on every transaction and never assume a reused payload is clean. The pooling manager above calls reset() on free() precisely to drop the byte-enable pointer and extensions before the payload is handed out again.

Corner 3: b_transport may consume time, and that has scheduling consequences

b_transport is permitted to call wait() internally — that is what "blocking" licenses. A target that models a long, time-consuming operation can wait() inside b_transport, and the initiator's thread is suspended for that simulated duration. This is legal and sometimes the cleanest way to model a slow peripheral. But it has a consequence: while the initiator is blocked inside such a call, it cannot issue another transaction, and the target — if it can only service one call at a time — serializes all initiators bound to it. For a single-initiator LT model this is fine. For a shared target with multiple initiators, an internally-waiting b_transport can model contention but can also accidentally serialize traffic you intended to overlap. The convention used throughout this post — targets annotate delay and return immediately, initiators pay the delay — sidesteps the issue: the target never blocks, so it never serializes, and timing is still modeled through annotation. Use an internally-waiting b_transport only when you specifically want the blocking-contention semantics.

Corner 4: response status is not the only error channel, but it is the one you use

The base protocol defines seven response-status values. TLM_OK_RESPONSE and TLM_INCOMPLETE_RESPONSE you have seen. The five error codes — TLM_GENERIC_ERROR_RESPONSE, TLM_ADDRESS_ERROR_RESPONSE, TLM_COMMAND_ERROR_RESPONSE, TLM_BURST_ERROR_RESPONSE, and TLM_BYTE_ENABLE_ERROR_RESPONSE — let a target be specific about why it failed: address out of range, command not supported, burst parameters illegal, byte-enable pattern unsupported, or a catch-all generic error. A well-behaved target picks the most specific applicable code; a well-behaved initiator at least distinguishes OK from not-OK via is_response_ok(), and a thorough one switches on get_response_string() for diagnostics. Targets must not invent their own out-of-band error signaling (throwing exceptions across a b_transport boundary, setting global flags) — the response status is the protocol's error channel, and tools, checkers, and interconnects all rely on it being the single source of truth.

Version differences

TLM-2.0 — the generic-payload layout, the b_transport signature, timing annotation, the response-status enum, and the acquire()/release()/memory-manager protocol — is identical across SystemC 2.3.0, 2.3.1, 2.3.3, 2.3.4, and 3.0.x. The headers ship bundled with the kernel from 2.3.0 onward. SystemC 3.0.x tracks IEEE 1666-2023, which deprecates the SC_HAS_PROCESS(name) macro form (3.0 prefers declaring processes directly in the constructor) and emits a deprecation warning if you use it, but none of the TLM-2.0 semantics in this post changed. Every example here compiles and runs identically on any 2.3.x or 3.0.x build; on 3.0.x you may see the SC_HAS_PROCESS deprecation note if you use that macro, which the SC_CTOR form in these examples avoids.

Hands-on exercise

Build a loosely-timed register-file target with byte-enable support and an initiator that exercises it.

The target models a small block of sixteen 32-bit registers (64 bytes total, addresses 0x00–0x3F). It services b_transport with these rules:

  • A read returns the addressed bytes from its register storage.
  • A write updates only the bytes whose byte enable is TLM_BYTE_ENABLED, leaving disabled bytes untouched (the Intermediate byte-enable pattern).
  • Any access whose address-plus-length exceeds 64 bytes returns TLM_ADDRESS_ERROR_RESPONSE without touching storage.
  • Every successful access annotates a latency of 1 ns per byte onto the delay argument (do not call wait() inside the target — annotate only).
  • Every successful access sets TLM_OK_RESPONSE.

Drive it with an SC_THREAD initiator that, using a single reused stack payload:

  1. Writes 0x11223344 to register 0 (address 0x00) with all bytes enabled, paying the annotated delay with wait() after the call.
  2. Does a sub-word write to register 0 that changes only byte 2, then reads register 0 back and prints all four bytes to prove the masked write touched exactly one byte.
  3. Attempts a 4-byte write at address 0x3E (which would run off the end) and prints the response status to show the address error is caught.

Print sc_time_stamp() before and after the first write so you can see the annotated latency advance simulation time. Predict the final simulation time and the masked-write byte pattern before you run, then check.

When that works, extend it: add a second initiator bound to the same target through a second target socket (or revisit this once you have read the sockets post, which shows how one target serves multiple initiators). Have both initiators write to different registers and confirm neither corrupts the other's data. Then make the target reject TLM_READ_COMMAND to a write-only register (say register 15) with TLM_COMMAND_ERROR_RESPONSE, and have an initiator confirm it.

Hints

  • Store the registers as a plain unsigned char storage[64] and index it by byte address, exactly like the Memory targets in this post. You do not need an array of uint32_t.
  • The byte-enable loop is if (be == nullptr || be[i % belen] == TLM_BYTE_ENABLED) storage[addr + i] = ptr[i]; — the same line from the Intermediate example.
  • Reset the response status to TLM_INCOMPLETE_RESPONSE and clear the byte-enable pointer (set_byte_enable_ptr(nullptr)) on the reused payload before any access that does not want a mask. The masked-write-then-read sequence is the exact place this bites.
  • Remember the address check (addr + len > 64) must happen before any storage access and must return immediately after setting the error status.
  • Annotate latency with delay += sc_time(len, SC_NS); and let the initiator wait(delay); delay = SC_ZERO_TIME;.
  • For the read-back print, copy the four bytes into a local buffer via a read transaction, then print them with std::hex and std::setw(2) << std::setfill('0').

No solution is provided. The understanding lives in getting the field-ownership and reuse-reset discipline right yourself.

Common mistakes

  • Forgetting to set the response status (or resetting it on reuse). A fresh payload defaults to TLM_INCOMPLETE_RESPONSE, but a reused one still carries the previous call's status. If a target forgets to call set_response_status(TLM_OK_RESPONSE), the transaction is silently marked failed. If an initiator forgets to reset to TLM_INCOMPLETE_RESPONSE before reusing a payload, a stale TLM_OK_RESPONSE can mask a real failure. Fix: targets always set a definite status before returning; initiators always reset to incomplete before each call and check after.
  • Reading return data before b_transport returns, or from another process. b_transport is blocking; the read data buffer is valid only after the call returns, only in the calling process. Reading it from a different process while the call is in flight, or assuming it is ready before the return, is a race the simulator will not flag. Fix: consume read data only on the line after the call, in the same thread that made the call.
  • Reusing a payload without clearing stale fields. Because the payload is passed by reference and reused, a leftover byte-enable pointer silently masks the next access, a leftover data length transfers the wrong byte count, and a leftover data pointer can point at a buffer that has gone out of scope. Fix: set every field you depend on (command, address, data ptr, length) on every transaction, and explicitly set_byte_enable_ptr(nullptr) when the next access does not want a mask. Use reset() (or a pooling manager that calls it) to drop extensions and byte enables between uses.
  • Treating data length as words instead of bytes. Every length in the payload — data length, byte-enable length, streaming width — is counted in bytes. A 32-bit access has length 4, not 1. Setting length 1 for a 32-bit transfer moves one byte and leaves the other three untouched. Fix: count bytes everywhere.
  • Calling acquire()/release() on a stack payload with no memory manager. Reference counting only works on payloads constructed with a memory manager (tlm_generic_payload(mm)). Calling acquire()/release() on a default-constructed stack payload is undefined behavior. Fix: use a stack payload (no acquire/release) for simple LT models; only switch to a manager-backed heap payload when the transaction must outlive the call, and then balance every acquire() with exactly one release().
  • Making the target call wait() when annotation would do. A target that calls wait() internally forces every initiator into pay-as-you-go timing and serializes shared targets, breaking temporal decoupling. Fix: annotate latency onto delay and return immediately; let the initiator decide when to pay. Reserve internally-blocking b_transport for the rare case where you specifically want blocking-contention semantics.

Recap

After working through this post you can now:

  • State the core abstraction — a TLM-2.0 transaction is a function call (b_transport) carrying a tlm_generic_payload struct by reference — and explain why the call's return is the transaction's completion.
  • Name every generic-payload attribute (command, address, data pointer, data length, byte enables, streaming width, DMI hint, response status) and say who sets each: the initiator owns the request, the target sets only the response status and, for a read, the data bytes.
  • Write a loosely-timed initiator/target pair from scratch with tlm_utils convenience sockets, fill the payload field by field, and bind them.
  • Model latency without a clock by annotating the sc_time& delay argument, and contrast the immediate-wait idiom with the accumulate-and-pay idiom that seeds temporal decoupling.
  • Enforce the response-status discipline — initiator sets TLM_INCOMPLETE_RESPONSE before, target sets TLM_OK_RESPONSE or a specific error, initiator checks is_response_ok()/is_response_error() after — and explain why it is the protocol's only error channel.
  • Perform sub-word access with byte enables, and recognize a stale byte-enable pointer as a reuse hazard.
  • Decide when a stack payload suffices (simple LT) versus when reference-counted memory management with acquire()/release() and a pool is required (payloads that outlive their call, AT pipelines).
  • Diagnose the common payload bugs on sight: missing or stale response status, reading return data too early, reusing a payload without resetting fields, and byte-vs-word length confusion.

Further reading

Standards

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §10.3 (generic payload: attributes §10.3.1, who-sets-what base-protocol rules §10.3.4, byte enables §10.3.3, response status §10.3.5, memory management §10.3.2) and §10.4 (b_transport semantics §10.4, timing annotation §10.4.6).

Vendor and consortium documents

  • Aynsley (Doulos), OSCI TLM-2.0 Language Reference Manual (JA32) — the canonical narrative treatment of the generic payload, base protocol, and LT/AT coding styles.
  • Doulos TLM-2.0 Tutorial — worked initiator/target/interconnect examples and the base-protocol checklist.
  • Accellera Systems Initiative, SystemC 3.0.0 distribution, include/tlm_core/tlm_2/ and include/tlm_utils/ — authoritative source for the payload class, convenience sockets, and the memory-manager interface.

Textbooks

  • Grötker, Liao, Martin, and Swan, System Design with SystemC — transaction-level modeling chapters; the transaction-as-function-call framing.

Next in this section

→ Part 3: Initiator & Target Sockets — what the tlm_utils convenience sockets actually wrap: the raw tlm_initiator_socket/tlm_target_socket, the forward and backward interface classes, how b_transport is routed from socket to method, and binding one initiator to many targets. Read it here: 18. SystemC Tutorial — Initiator & Target Sockets.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment