19. SystemC Tutorial - Non-Blocking Transport & the AT Protocol

Why this matters

Written 2026-06-05 for the concept-first series.

In Part 2 you learned that a TLM-2.0 transaction is a function call carrying a tlm_generic_payload by reference: the initiator fills in a struct, calls b_transport, and the target services the whole thing and returns. One call, one round-trip, transaction done. That model — loosely-timed, LT — is the right tool for the overwhelming majority of virtual platforms, the kind that boot Linux in seconds. It is simple, it is fast, and it is where you should start every model. This post is about the other coding style, the one you reach for when a single round-trip call is no longer enough to express what the hardware does: the approximately-timed (AT) style, built on non-blocking transport.

Here is the gap LT cannot close. A real bus does not freeze the initiator from the instant it issues a request until the instant the data comes back. It accepts the request, frees the initiator to issue the next request, and delivers the response later — so that several transactions are in flight at once, overlapping in time the way a real pipelined interconnect overlaps them. A real target with a finite request buffer pushes back when it is full, stalling the initiator until a slot frees — backpressure. None of that — overlap, pipelining, backpressure — can be expressed by a function call that does not return until the transaction is complete, because by construction b_transport holds the initiator hostage for the whole round-trip. To model the overlap, you must be able to return control to the initiator after the request is accepted but before the response arrives. That is precisely what splitting one blocking call into separate request and response phases buys you, and that splitting is the entire subject of this post.

This is the hardest post in the section, and it is hard for an honest reason: the protocol that governs how those phases may be sequenced — when a target may accept a request, when it must push back, when the response may begin, who advances the transaction and how completion is signalled — is a genuine state machine with rules you must get exactly right or your model deadlocks, races, or silently corrupts a payload that has gone out of scope. By the end of this post you will know the four base-protocol phases (BEGIN_REQ, END_REQ, BEGIN_RESP, END_RESP) and their one legal ordering; the three tlm_sync_enum return codes (TLM_ACCEPTED, TLM_UPDATED, TLM_COMPLETED) and exactly what each one obligates the caller to do; the early-completion shortcut that lets a fast target collapse all four phases into one call; how to build a pipelined initiator and target that keep two transactions in flight and apply backpressure by withholding a phase; why the stack payload that served you so well in LT is now a use-after-free waiting to happen, and the reference-counted memory management that replaces it; the peq_with_cb_and_phase utility that every production AT model uses to schedule phase transitions; and the decision table for when AT is worth its cost versus when you should stay in LT. That is the bar. It is a high one, and it is the last conceptual wall in the transport half of this section.

Prerequisites

  • Part 2 — The Generic Payload & Blocking Transport. This post is the direct continuation of Part 2. You must be fluent in the generic payload — every field, who sets it, the response-status discipline, timing annotation through the sc_time& delay argument — and in the b_transport round-trip, because the whole point of this post is to break that single call into phases. Critically, Part 2's Advanced section introduced reference-counted payload memory management (acquire()/release() with a memory manager); under AT that is no longer optional, so if it was hazy, re-read it.
  • Part 3 — Initiator & Target Sockets. The non-blocking calls in this post travel both ways — the initiator's nb_transport_fw reaches the target, and the target's nb_transport_bw reaches back to the initiator. You need to understand the forward and backward interfaces a socket carries, and how register_nb_transport_fw / register_nb_transport_bw wire a method to each path, which Part 3 covers.
  • Part 5 (F5) — Events & Notifications. AT models defer phase transitions to future simulation times: a target receives a request now and emits END_REQ 2 ns later. That deferral is built on sc_event and timed notify(). The peq_with_cb_and_phase utility we use wraps exactly this machinery, and the pipelined initiator uses a bare sc_event to gate its request stream, so you need to be comfortable with event notification and wait(event).
  • Part 2 (F2) — Simulation Time. The traces in this post are read in simulation time — sc_time_stamp() advancing as phases fire at annotated delays. You need to be fluent in how sc_time works, how delay offsets the effective time of a phase, and how to read a timestamped trace.

If any of these is shaky, fix it before continuing. This post does not re-teach the payload, sockets, events, or time — it composes all four into a protocol.

Mental model (first principles)

Start from the one sentence that frames everything: non-blocking transport is b_transport cut into pieces so that control can return to the initiator in the middle of a transaction.

In LT, the lifecycle of a transaction is a single function call. Control goes into the target on b_transport and does not come back until the transaction is complete. That is fine until you need two transactions to overlap — and you cannot, because the initiator is stuck inside the call. So AT splits the one call into a sequence of shorter calls, each of which returns immediately, and uses a small piece of state — the phase — to track where in the transaction's life each call lands.

A transaction in the base protocol has exactly two halves and four boundaries. The request half: it begins (BEGIN_REQ) when the initiator hands the request to the target, and it ends (END_REQ) when the target has taken the request in and the initiator is free to send the next one. The response half: it begins (BEGIN_RESP) when the target has the result ready and hands it back, and it ends (END_RESP) when the initiator has accepted the result and the transaction is retired. Four phase markers, and exactly one legal order for a single transaction:

BEGIN_REQ  →  END_REQ  →  BEGIN_RESP  →  END_RESP

The request half closes before the response half opens — that is the request/response exclusion rule, and it is why the target must send END_REQ before it sends BEGIN_RESP. The two halves never overlap for one transaction. But — and this is the whole reason AT exists — they freely overlap across different transactions. While transaction A is sitting in its response half (waiting for its data), transaction B can be in its request half. That is pipelining, and it is exactly what b_transport could not express.

An analogy makes the split intuitive. LT b_transport is a phone call: you dial, the other party picks up, you state your business, they handle it while you hold the line, they give you the answer, you hang up — one continuous connection, and you cannot make a second call until this one ends. AT non-blocking transport is a ticket queue at a service counter: you walk up and hand over your request slip (BEGIN_REQ); the clerk takes it and waves you aside — "got it, next!" (END_REQ) — which frees you to immediately hand over a second slip while the first is still being worked; later the clerk calls your number and hands you your result (BEGIN_RESP); you take it and confirm "got it" (END_RESP) and your ticket is done. The slip (the payload) is passed back and forth; the four announcements are the phases; and crucially you were free to queue a second request the moment the clerk waved you aside on the first — that is the overlap a phone call cannot give you.

Now the two mechanics that carry the phases. There are two methods, one per direction, because phases travel both ways. The initiator sends BEGIN_REQ forward to the target by calling nb_transport_fw; the target sends END_REQ and BEGIN_RESP backward to the initiator by calling nb_transport_bw; the initiator sends END_RESP forward again. Both methods have the identical shape:

tlm::tlm_sync_enum nb_transport_fw(tlm::tlm_generic_payload& trans,
                                   tlm::tlm_phase& phase, sc_time& delay);
tlm::tlm_sync_enum nb_transport_bw(tlm::tlm_generic_payload& trans,
                                   tlm::tlm_phase& phase, sc_time& delay);

Three arguments, all by reference. trans is the same payload object throughout the transaction's life (so memory management matters — more below). phase is the marker being delivered. delay is the timing annotation, exactly as in b_transport: "this phase transition actually happens delay from now". And both methods are non-blocking in the literal sense — they must return immediately, must never call wait(), and must never consume simulation time. Elapsed time is expressed only through delay and through scheduling a future callback; never by blocking.

The return type is the second new idea, the tlm_sync_enum, and it has three values that each obligate the caller differently:

  • TLM_ACCEPTED — "I have noted your call, but I have advanced nothing. Ignore the phase and delay I am returning. Expect a future call on the opposite path to carry the transaction forward." This is the "got it, I'll get back to you" answer. A target that needs simulation time to produce END_REQ returns TLM_ACCEPTED to the initiator's BEGIN_REQ and later calls nb_transport_bw with END_REQ.
  • TLM_UPDATED — "I have advanced the transaction synchronously. The phase and delay I am returning are the new state — act on them now." A target that can answer BEGIN_REQ with END_REQ immediately, in the same call, returns TLM_UPDATED with phase = END_REQ.
  • TLM_COMPLETED — "The entire transaction is finished. There will be no further phase transitions on either path." This is the early-completion shortcut: a target can answer the very first BEGIN_REQ with TLM_COMPLETED and skip every intermediate phase, collapsing AT back down to a single call.

Two warnings that will save you hours. First, TLM_ACCEPTED and TLM_UPDATED are not interchangeable. TLM_ACCEPTED says "ignore my return arguments, wait for a callback"; TLM_UPDATED says "my return arguments are real, act now." Treat an TLM_ACCEPTED as if it carried a real phase and your handshake desynchronizes. Second, TLM_COMPLETED is a protocol statement, not a status statement. It means "no more phases", not "it succeeded." Whether the transaction worked still lives in trans.get_response_status(), exactly as in Part 2. A target may legally return TLM_COMPLETED carrying TLM_ADDRESS_ERROR_RESPONSE: finished and failed. Always check the response status independently of the sync enum.

One last piece of the mental model, the one that bites hardest. In LT, the payload could live on the initiator's stack because b_transport returned before the payload was reused — the stack frame outlived the call trivially. Under AT that is no longer true. The phases are separated in simulation time; the target schedules END_REQ and BEGIN_RESP to fire later, after the nb_transport_fw call that delivered BEGIN_REQ has already returned and the initiator's calling function may have moved on. The payload is touched by those future callbacks, so it must outlive the call — which means it cannot be a stack local that unwinds. The base protocol therefore requires a memory-managed heap payload: construct it with a memory manager, acquire() it before BEGIN_REQ, and release() it after END_RESP. This is the reference counting Part 2 introduced as an Advanced corner; under AT it is mandatory, not optional. We will see a payload-on-the-stack version segfault and fix it, because seeing it fail is the fastest way to internalize why AT needs memory management.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%%
sequenceDiagram
    participant I as Initiator
    participant T as Target
    I->>T: nb_transport_fw(trans, BEGIN_REQ, d)
    Note right of T: returns TLM_ACCEPTED
    T-->>I: nb_transport_bw(trans, END_REQ, d)
    Note left of I: request phase closed; initiator free to send next req
    T-->>I: nb_transport_bw(trans, BEGIN_RESP, d)
    Note left of I: target has the result
    I->>T: nb_transport_fw(trans, END_RESP, d)
    Note right of T: returns TLM_COMPLETED; transaction retired

Read the diagram as the ticket-queue round-trip: hand in the slip (BEGIN_REQ), get waved aside (END_REQ), get called back with the result (BEGIN_RESP), confirm receipt (END_RESP). Forward calls go down, backward calls go up, and the gap between END_REQ and BEGIN_RESP is exactly the window in which a second transaction's request can run. Hold the four-phase order and the two-path mechanic in your head; the rest of this post is making them concrete in compiled, running code.

Beginner: First Principles

The simplest interesting AT system is one initiator and one target walking a single transaction through all four phases, one transaction at a time. There is no pipelining yet and no backpressure — just the four-phase handshake made concrete, so you can watch BEGIN_REQ → END_REQ → BEGIN_RESP → END_RESP fire in order, on the right paths, at the right simulated times. Build this, run it, and read the trace before you read another word of explanation; the ordering is the whole lesson.

Two new ingredients appear here. First, the target registers two things on its socket compared to the LT target's one: it still has a forward handler (now nb_transport_fw instead of b_transport), and the initiator now also registers a backward handler (nb_transport_bw) so the target can call back to it. Second, the target needs to emit END_REQ 2 ns after it receives BEGIN_REQ, and BEGIN_RESP 8 ns after that — it must schedule those phase transitions for the future without blocking. The tool for that is tlm_utils::peq_with_cb_and_phase, a payload-event-queue: you call peq.notify(trans, phase, delay) and it invokes your callback peq_cb(trans, phase) at now + delay. It is the production-standard way to defer a phase, and we use it from the start so the idiom is familiar.

// file: nb_first_handshake.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
//            -I$SYSTEMC_HOME/include nb_first_handshake.cpp \
//            -o nb_first_handshake -L$SYSTEMC_HOME/lib -lsystemc
// Run:   LD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./nb_first_handshake

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <tlm_utils/peq_with_cb_and_phase.h>
#include <iostream>

using namespace tlm;
using namespace sc_core;

// A target that walks one transaction through all four base-protocol phases,
// one transaction at a time. Request occupancy 2 ns, response latency 8 ns.
SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  tlm_utils::peq_with_cb_and_phase<Memory> m_peq;   // schedules deferred phases
  unsigned char mem[256];

  SC_CTOR(Memory) : socket("socket"), m_peq(this, &Memory::peq_cb) {
    for (int i = 0; i < 256; i++) mem[i] = 0;
    socket.register_nb_transport_fw(this, &Memory::nb_transport_fw);
  }

  // FORWARD path: calls arriving from the initiator.
  tlm_sync_enum nb_transport_fw(tlm_generic_payload& trans,
                                tlm_phase& phase, sc_time& delay) {
    if (phase == BEGIN_REQ) {
      std::cout << "[" << sc_time_stamp() << "] TARGET  fw BEGIN_REQ\n";
      // Defer END_REQ to 2 ns from now; do NOT block, do NOT answer inline.
      m_peq.notify(trans, END_REQ, delay + sc_time(2, SC_NS));
      return TLM_ACCEPTED;          // "noted; a bw call will carry it forward"
    }
    if (phase == END_RESP) {
      std::cout << "[" << sc_time_stamp() << "] TARGET  fw END_RESP -> done\n";
      return TLM_COMPLETED;         // transaction retired
    }
    return TLM_ACCEPTED;
  }

  // The PEQ fires here at the scheduled simulation time.
  void peq_cb(tlm_generic_payload& trans, const tlm_phase& phase) {
    sc_time d = SC_ZERO_TIME;
    if (phase == END_REQ) {
      std::cout << "[" << sc_time_stamp() << "] TARGET  bw END_REQ\n";
      tlm_phase ph = END_REQ;
      socket->nb_transport_bw(trans, ph, d);   // close the request phase
      // Perform the access now, then schedule the response 8 ns out.
      sc_dt::uint64  a = trans.get_address();
      unsigned int   l = trans.get_data_length();
      unsigned char* p = trans.get_data_ptr();
      if (trans.get_command() == TLM_WRITE_COMMAND)
        for (unsigned i = 0; i < l; i++) mem[a + i] = p[i];
      else
        for (unsigned i = 0; i < l; i++) p[i] = mem[a + i];
      trans.set_response_status(TLM_OK_RESPONSE);
      m_peq.notify(trans, BEGIN_RESP, sc_time(8, SC_NS));
    } else if (phase == BEGIN_RESP) {
      std::cout << "[" << sc_time_stamp() << "] TARGET  bw BEGIN_RESP\n";
      tlm_phase ph = BEGIN_RESP;
      socket->nb_transport_bw(trans, ph, d);   // open the response phase
    }
  }
};

SC_MODULE(Cpu) {
  tlm_utils::simple_initiator_socket<Cpu> socket;
  unsigned w = 0x5555;
  tlm_generic_payload trans;   // member, not a stack local: it outlives run()

  SC_CTOR(Cpu) : socket("socket") {
    socket.register_nb_transport_bw(this, &Cpu::nb_transport_bw);
    SC_THREAD(run);
  }

  void run() {
    sc_time delay = SC_ZERO_TIME;
    trans.set_command(TLM_WRITE_COMMAND);
    trans.set_address(0x10);
    trans.set_data_ptr(reinterpret_cast<unsigned char*>(&w));
    trans.set_data_length(4);
    trans.set_streaming_width(4);
    trans.set_byte_enable_ptr(nullptr);
    trans.set_response_status(TLM_INCOMPLETE_RESPONSE);
    tlm_phase phase = BEGIN_REQ;
    std::cout << "[" << sc_time_stamp() << "] INIT    fw BEGIN_REQ\n";
    socket->nb_transport_fw(trans, phase, delay);   // send the request
  }

  // BACKWARD path: calls arriving from the target.
  tlm_sync_enum nb_transport_bw(tlm_generic_payload& trans,
                                tlm_phase& phase, sc_time& delay) {
    if (phase == END_REQ) {
      std::cout << "[" << sc_time_stamp() << "] INIT    bw END_REQ\n";
      return TLM_ACCEPTED;            // request done; (here we send no more reqs)
    }
    if (phase == BEGIN_RESP) {
      std::cout << "[" << sc_time_stamp() << "] INIT    bw BEGIN_RESP resp="
                << trans.get_response_string() << "\n";
      tlm_phase ph = END_RESP;
      sc_time d = SC_ZERO_TIME;
      socket->nb_transport_fw(trans, ph, d);   // accept response -> END_RESP
      std::cout << "[" << sc_time_stamp()
                << "] INIT    fw END_RESP -> transaction done\n";
      return TLM_ACCEPTED;
    }
    return TLM_ACCEPTED;
  }
};

int sc_main(int, char*[]) {
  Cpu cpu("cpu");
  Memory mem("mem");
  cpu.socket.bind(mem.socket);
  sc_start();
  return 0;
}

Expected output:

[0 s] INIT    fw BEGIN_REQ
[0 s] TARGET  fw BEGIN_REQ
[2 ns] TARGET  bw END_REQ
[2 ns] INIT    bw END_REQ
[10 ns] TARGET  bw BEGIN_RESP
[10 ns] INIT    bw BEGIN_RESP resp=TLM_OK_RESPONSE
[10 ns] TARGET  fw END_RESP -> done
[10 ns] INIT    fw END_RESP -> transaction done

Trace it phase by phase against the four-phase order. At time 0 the initiator's run() thread fills the payload — exactly the same fields as an LT transaction, including the TLM_INCOMPLETE_RESPONSE sentinel — sets phase = BEGIN_REQ, and calls nb_transport_fw. Control enters the target's forward handler with phase == BEGIN_REQ. The target does not service the transaction inline. It schedules END_REQ for 2 ns in the future via the PEQ and returns TLM_ACCEPTED — "noted, I'll call you back." Control returns to the initiator, which has nothing more to do, so run() ends. Note what just happened: the transaction is not complete, yet the forward call has returned. That return-before-completion is the thing b_transport could never do.

At 2 ns the PEQ fires peq_cb with END_REQ. The target calls back into the initiator via nb_transport_bw with phase = END_REQ — the request phase is now closed. Then, still inside the same callback, the target performs the actual memory write and schedules BEGIN_RESP for 8 ns later. The initiator's backward handler prints END_REQ and returns TLM_ACCEPTED; in a single-outstanding model like this one it has no next request to send.

At 10 ns (2 + 8) the PEQ fires again with BEGIN_RESP. The target calls nb_transport_bw with phase = BEGIN_RESP, handing the result back. The initiator's backward handler reads the response status (TLM_OK_RESPONSE — checked independently of any sync enum, exactly as in LT), then accepts the response by calling nb_transport_fw forward with phase = END_RESP. The target's forward handler sees END_RESP and returns TLM_COMPLETED: the transaction is retired. All four phases fired, in the one legal order, on the correct paths, with simulation time advancing only at the annotated boundaries.

Three details to fix in your mind now.

First, the request phase fully closes before the response phase opens. END_REQ fires at 2 ns; BEGIN_RESP does not fire until 10 ns. The 8 ns gap between them is the request/response exclusion rule in action, and — although this single-transaction example does not exploit it — that gap is exactly the window where a pipelined initiator would inject its next request. Hold that thought; the next section fills the gap.

Second, non-blocking really means non-blocking. Look at every method: not one of them calls wait(). The target expresses its 2 ns and 8 ns latencies entirely by scheduling PEQ notifications at a future time, and returns immediately each time. The initiator's handlers return immediately too. Time advances because the kernel runs the PEQ's events forward, not because any process blocked inside a transport call. This is the structural difference from b_transport, which was allowed to wait() internally.

Third — and this is the trap the next paragraph is built to spring — the payload is a member of Cpu, not a stack local in run(). Read that line again: tlm_generic_payload trans; sits at module scope. That is deliberate and load-bearing. The target's PEQ touches trans at 2 ns and 10 ns, long after run() returned at time 0. If trans had been a local in run(), its stack frame would have unwound the instant run() ended, and the PEQ callback at 2 ns would dereference a destroyed object. Let us prove that is not hypothetical.

Warning If you move trans from a Cpu member back into run() as a stack local — void run() { tlm_generic_payload trans; ... } — this exact program segfaults. The PEQ fires END_REQ at 2 ns and dereferences a payload whose stack frame unwound at time 0. Under LT a stack payload was safe because b_transport returned only after the transaction completed; under AT the transaction outlives the call that launched it, so the payload must outlive the call too. A module member works for a single-outstanding model; the Advanced section replaces it with the proper reference-counted heap payload that AT requires once transactions overlap.

That warning is the conceptual hinge of the whole post. Everything about AT memory management flows from "the payload outlives the call." Keep it in view as we move to pipelining, where a single member payload no longer suffices because multiple transactions are alive at once.

Intermediate: How It Really Works

The Beginner example walked one transaction through the phases with nothing overlapping — which throws away the entire reason AT exists. This section builds the real thing: a pipelined initiator that issues a new request the instant the previous one is accepted, and a target that keeps up to two transactions in flight and applies backpressure by withholding END_REQ when its pipeline is full. The trace will show two transactions genuinely overlapping in time, and a third stalling against the backpressure — the behaviors b_transport could not produce. This is also where the single member payload breaks down: with multiple transactions alive simultaneously, we need a pool of payloads, which is the AT memory management the Beginner warning foreshadowed.

The design in one paragraph. The target tracks an outstanding counter and a MAX_OUTSTANDING cap of 2. On BEGIN_REQ, if there is a free slot it accepts the request (schedules END_REQ 2 ns out) and bumps the counter; if the pipeline is full it parks the request on a backlog queue and says nothing — withholding END_REQ is the backpressure. Each accepted transaction then schedules its BEGIN_RESP 10 ns after its END_REQ. When a response retires (END_RESP), a slot frees, the counter drops, and a parked request is admitted. The initiator's generator thread issues BEGIN_REQ, then blocks on an sc_event that the backward END_REQ handler notifies — so it sends the next request exactly when the target accepts the current one, running ahead of the responses.

// file: at_pipeline.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
//            -I$SYSTEMC_HOME/include at_pipeline.cpp \
//            -o at_pipeline -L$SYSTEMC_HOME/lib -lsystemc
// Run:   LD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./at_pipeline

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <tlm_utils/peq_with_cb_and_phase.h>
#include <iostream>
#include <queue>
#include <vector>

using namespace tlm;
using namespace sc_core;

// ============================================================================
//  A tiny reference-counted payload pool (memory manager).
//  Under AT, a payload outlives the nb_transport call that launched it, so it
//  must live on the heap with explicit acquire()/release() lifetime.
// ============================================================================
class Pool : public tlm_mm_interface {
public:
  tlm_generic_payload* alloc() {
    tlm_generic_payload* p;
    if (free_list.empty()) { p = new tlm_generic_payload(this); created++; }
    else { p = free_list.back(); free_list.pop_back(); }
    return p;
  }
  void free(tlm_generic_payload* p) override { p->reset(); free_list.push_back(p); }
  int created = 0;
  ~Pool() { for (auto* p : free_list) delete p; }
private:
  std::vector<tlm_generic_payload*> free_list;
};

// ============================================================================
//  Pipelined AT target.
//  - Accepts at most MAX_OUTSTANDING requests at once (request pipeline depth).
//  - When full, it withholds END_REQ (backpressure) until a response drains.
//  - Request occupancy = 2 ns; access+response latency = 10 ns.
// ============================================================================
SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  unsigned char mem[256];
  tlm_utils::peq_with_cb_and_phase<Memory> m_peq;

  static const int MAX_OUTSTANDING = 2;
  int outstanding;                            // requests accepted, not responded
  std::queue<tlm_generic_payload*> backlog;   // requests parked by backpressure

  SC_CTOR(Memory)
    : socket("socket"), m_peq(this, &Memory::peq_cb), outstanding(0) {
    for (int i = 0; i < 256; i++) mem[i] = 0;
    socket.register_nb_transport_fw(this, &Memory::nb_transport_fw);
  }

  tlm_sync_enum nb_transport_fw(tlm_generic_payload& trans,
                                tlm_phase& phase, sc_time& delay) {
    if (phase == BEGIN_REQ) {
      std::cout << "[" << sc_time_stamp() << "] TARGET   BEGIN_REQ  addr=0x"
                << std::hex << trans.get_address() << std::dec
                << "  (outstanding=" << outstanding << ")\n";
      if (outstanding < MAX_OUTSTANDING) {
        accept_request(trans, delay);
      } else {
        // Backpressure: pipeline full -> park the request, withhold END_REQ.
        std::cout << "[" << sc_time_stamp() << "] TARGET   (full) parking addr=0x"
                  << std::hex << trans.get_address() << std::dec << "\n";
        backlog.push(&trans);
      }
      return TLM_ACCEPTED;
    }
    if (phase == END_RESP) {
      // Initiator finished with the response; this frees a pipeline slot.
      std::cout << "[" << sc_time_stamp() << "] TARGET   END_RESP   addr=0x"
                << std::hex << trans.get_address() << std::dec << "\n";
      outstanding--;
      trans.release();
      if (!backlog.empty()) {                 // a slot freed: admit a parked req
        tlm_generic_payload* p = backlog.front(); backlog.pop();
        sc_time d = SC_ZERO_TIME;
        accept_request(*p, d);
      }
      return TLM_COMPLETED;
    }
    return TLM_ACCEPTED;
  }

  void accept_request(tlm_generic_payload& trans, sc_time& delay) {
    outstanding++;
    m_peq.notify(trans, END_REQ, delay + sc_time(2, SC_NS));  // 2 ns occupancy
  }

  void peq_cb(tlm_generic_payload& trans, const tlm_phase& phase) {
    sc_time delay = SC_ZERO_TIME;
    if (phase == END_REQ) {
      std::cout << "[" << sc_time_stamp() << "] TARGET   END_REQ    addr=0x"
                << std::hex << trans.get_address() << std::dec << "\n";
      tlm_phase ph = END_REQ;
      socket->nb_transport_bw(trans, ph, delay);   // free initiator to send next
      do_access(trans);
      m_peq.notify(trans, BEGIN_RESP, sc_time(10, SC_NS));
    } else if (phase == BEGIN_RESP) {
      std::cout << "[" << sc_time_stamp() << "] TARGET   BEGIN_RESP addr=0x"
                << std::hex << trans.get_address() << std::dec << "\n";
      tlm_phase ph = BEGIN_RESP;
      tlm_sync_enum s = socket->nb_transport_bw(trans, ph, delay);
      if (s == TLM_COMPLETED) {
        // Initiator completed on the return path (skipped a separate END_RESP).
        std::cout << "[" << sc_time_stamp() << "] TARGET   END_RESP   addr=0x"
                  << std::hex << trans.get_address() << std::dec << " (via return)\n";
        outstanding--;
        trans.release();
        if (!backlog.empty()) {
          tlm_generic_payload* p = backlog.front(); backlog.pop();
          sc_time d = SC_ZERO_TIME;
          accept_request(*p, d);
        }
      }
    }
  }

  void do_access(tlm_generic_payload& trans) {
    sc_dt::uint64  a = trans.get_address();
    unsigned int   l = trans.get_data_length();
    unsigned char* p = trans.get_data_ptr();
    if (trans.get_command() == TLM_WRITE_COMMAND)
      for (unsigned i = 0; i < l; i++) mem[a + i] = p[i];
    else
      for (unsigned i = 0; i < l; i++) p[i] = mem[a + i];
    trans.set_response_status(TLM_OK_RESPONSE);
  }
};

// ============================================================================
//  Pipelined AT initiator: issues a new request the instant the previous one
//  is accepted (END_REQ), so it runs ahead of the responses.
// ============================================================================
SC_MODULE(Cpu) {
  tlm_utils::simple_initiator_socket<Cpu> socket;
  Pool pool;
  sc_event can_send;          // notified on END_REQ: a request slot is free
  unsigned data[6];

  SC_CTOR(Cpu) : socket("socket") {
    socket.register_nb_transport_bw(this, &Cpu::nb_transport_bw);
    SC_THREAD(gen);
  }

  void gen() {
    for (int i = 0; i < 6; i++) {
      tlm_generic_payload* t = pool.alloc();
      t->acquire();                       // payload now lives until release()
      data[i] = 0x100 + i;
      t->set_command(TLM_WRITE_COMMAND);
      t->set_address(i * 4);
      t->set_data_ptr(reinterpret_cast<unsigned char*>(&data[i]));
      t->set_data_length(4);
      t->set_streaming_width(4);
      t->set_byte_enable_ptr(nullptr);
      t->set_response_status(TLM_INCOMPLETE_RESPONSE);

      tlm_phase phase = BEGIN_REQ;
      sc_time delay = SC_ZERO_TIME;
      std::cout << "[" << sc_time_stamp() << "] INIT     send BEGIN_REQ  addr=0x"
                << std::hex << t->get_address() << std::dec << "\n";
      tlm_sync_enum s = socket->nb_transport_fw(*t, phase, delay);
      if (s == TLM_COMPLETED) { t->release(); continue; }
      wait(can_send);    // block until END_REQ frees us to send the next request
    }
  }

  tlm_sync_enum nb_transport_bw(tlm_generic_payload& trans,
                                tlm_phase& phase, sc_time& delay) {
    if (phase == END_REQ) {
      std::cout << "[" << sc_time_stamp() << "] INIT     recv END_REQ    addr=0x"
                << std::hex << trans.get_address() << std::dec << "\n";
      can_send.notify(SC_ZERO_TIME);     // release the generator for the next req
      return TLM_ACCEPTED;
    }
    if (phase == BEGIN_RESP) {
      std::cout << "[" << sc_time_stamp() << "] INIT     recv BEGIN_RESP addr=0x"
                << std::hex << trans.get_address()
                << " resp=" << trans.get_response_string() << std::dec << "\n";
      tlm_phase ph = END_RESP;
      sc_time d = SC_ZERO_TIME;
      socket->nb_transport_fw(trans, ph, d);   // accept response -> END_RESP
      return TLM_ACCEPTED;
    }
    return TLM_ACCEPTED;
  }
};

int sc_main(int, char*[]) {
  Cpu cpu("cpu");
  Memory mem("mem");
  cpu.socket.bind(mem.socket);
  sc_start();
  std::cout << "payloads allocated: " << cpu.pool.created << "\n";
  return 0;
}

Expected output:

[0 s] INIT     send BEGIN_REQ  addr=0x0
[0 s] TARGET   BEGIN_REQ  addr=0x0  (outstanding=0)
[2 ns] TARGET   END_REQ    addr=0x0
[2 ns] INIT     recv END_REQ    addr=0x0
[2 ns] INIT     send BEGIN_REQ  addr=0x4
[2 ns] TARGET   BEGIN_REQ  addr=0x4  (outstanding=1)
[4 ns] TARGET   END_REQ    addr=0x4
[4 ns] INIT     recv END_REQ    addr=0x4
[4 ns] INIT     send BEGIN_REQ  addr=0x8
[4 ns] TARGET   BEGIN_REQ  addr=0x8  (outstanding=2)
[4 ns] TARGET   (full) parking addr=0x8
[12 ns] TARGET   BEGIN_RESP addr=0x0
[12 ns] INIT     recv BEGIN_RESP addr=0x0 resp=TLM_OK_RESPONSE
[12 ns] TARGET   END_RESP   addr=0x0
[14 ns] TARGET   BEGIN_RESP addr=0x4
[14 ns] INIT     recv BEGIN_RESP addr=0x4 resp=TLM_OK_RESPONSE
[14 ns] TARGET   END_RESP   addr=0x4
[14 ns] TARGET   END_REQ    addr=0x8
[14 ns] INIT     recv END_REQ    addr=0x8
[14 ns] INIT     send BEGIN_REQ  addr=0xc
[14 ns] TARGET   BEGIN_REQ  addr=0xc  (outstanding=1)
[16 ns] TARGET   END_REQ    addr=0xc
[16 ns] INIT     recv END_REQ    addr=0xc
[16 ns] INIT     send BEGIN_REQ  addr=0x10
[16 ns] TARGET   BEGIN_REQ  addr=0x10  (outstanding=2)
[16 ns] TARGET   (full) parking addr=0x10
[24 ns] TARGET   BEGIN_RESP addr=0x8
[24 ns] INIT     recv BEGIN_RESP addr=0x8 resp=TLM_OK_RESPONSE
[24 ns] TARGET   END_RESP   addr=0x8
[26 ns] TARGET   BEGIN_RESP addr=0xc
[26 ns] INIT     recv BEGIN_RESP addr=0xc resp=TLM_OK_RESPONSE
[26 ns] TARGET   END_RESP   addr=0xc
[26 ns] TARGET   END_REQ    addr=0x10
[26 ns] INIT     recv END_REQ    addr=0x10
[26 ns] INIT     send BEGIN_REQ  addr=0x14
[26 ns] TARGET   BEGIN_REQ  addr=0x14  (outstanding=1)
[28 ns] TARGET   END_REQ    addr=0x14
[28 ns] INIT     recv END_REQ    addr=0x14
[36 ns] TARGET   BEGIN_RESP addr=0x10
[36 ns] INIT     recv BEGIN_RESP addr=0x10 resp=TLM_OK_RESPONSE
[36 ns] TARGET   END_RESP   addr=0x10
[38 ns] TARGET   BEGIN_RESP addr=0x14
[38 ns] INIT     recv BEGIN_RESP addr=0x14 resp=TLM_OK_RESPONSE
[38 ns] TARGET   END_RESP   addr=0x14
payloads allocated: 3

This trace is the payoff for the whole post, so read it slowly. The story is in the overlap, which the single-call b_transport model is structurally incapable of producing.

Two transactions in flight at once. Follow transaction 0x0. Its BEGIN_REQ lands at 0 s; its END_REQ fires at 2 ns; but its BEGIN_RESP does not arrive until 12 ns (2 ns occupancy + 10 ns access latency). For the entire 10 ns that 0x0 spends waiting for its response, the initiator is not idle. At 2 ns — the instant 0x0's request is accepted — the initiator fires 0x4's BEGIN_REQ, and 0x4 is accepted (outstanding goes from 1 to 2). So between 2 ns and 12 ns, both 0x0 and 0x4 are live transactions — 0x0 in its response half, 0x4 in its request-then-response half. Two transactions overlapping in time is precisely what AT delivers and LT cannot. Under b_transport, the initiator would have been frozen inside the 0x0 call from 0 to 12 ns and could not have touched 0x4 at all.

Backpressure, visible. Watch 0x8. At 4 ns the initiator sends 0x8's BEGIN_REQ, but the target is already full — outstanding == 2 (both 0x0 and 0x4 are in flight). The target prints (full) parking addr=0x8 and withholds END_REQ. With no END_REQ coming back, the initiator's generator is blocked in wait(can_send) — it cannot send another request. The pipeline has stalled the producer. That stall persists until 14 ns, when 0x0's END_RESP (at 12 ns) and 0x4's END_RESP (at 14 ns) drain slots; the admitted 0x8 finally gets its END_REQ at 14 ns, which unblocks the initiator to send 0xc. Withholding a phase is the flow-control mechanism — no extra signal, no credit counter visible to the initiator, just the absence of END_REQ. That is how a real bus pushes back, and AT expresses it natively.

The pool recycles. The final line reads payloads allocated: 3, even though the initiator issued six transactions. Because at most two or three payloads are alive at any instant (the cap is 2 outstanding, plus brief overlap at slot hand-off), the pool never needs to allocate more than three tlm_generic_payload objects — it reset()s and reuses them as transactions retire. In an AT model issuing millions of transactions this is the difference between three heap objects and millions of new/delete pairs. The lifetime is explicit and balanced: the initiator acquire()s each payload before BEGIN_REQ, and the target release()s it when it processes END_RESP. Every acquire() is matched by exactly one release(); the last release() drives the reference count to zero, calls Pool::free(), and recycles the object. Get that balance wrong — an extra acquire() leaks; an extra release() is a use-after-free — and the model breaks in exactly the ways the Beginner warning described.

One subtlety in the code worth calling out: the target's peq_cb handles the case where the initiator returns TLM_COMPLETED from the BEGIN_RESP backward call (completing the transaction on the return of nb_transport_bw rather than with a separate forward END_RESP call). Our initiator does not take that shortcut — it sends an explicit END_RESP and returns TLM_ACCEPTED — so the (via return) branch never fires here, but a conformant target must handle both, because the base protocol permits a peer to complete the response either way. Defensive handling of "the other legal path" is a recurring theme in AT code, and it is why AT models are several times the size of their LT equivalents.

Advanced: Edge Cases & LRM Corners

The Beginner and Intermediate sections cover the load-bearing AT mechanics: the four-phase handshake, pipelining, backpressure, and pooled payloads. This section drills into the corners the tutorials skip — the early-completion shortcut that bridges AT back to LT, the exact memory-management contract that Part 2 introduced and AT makes mandatory, the peq_with_cb_and_phase tool itself, and a handful of LRM subtleties that bite senior engineers.

Corner 1: the early-completion shortcut — collapsing four phases into one

Nothing forces a target to use all four phases. The tlm_sync_enum return value TLM_COMPLETED lets a target answer the very first BEGIN_REQ by declaring the entire transaction finished, in the same call, skipping END_REQ, BEGIN_RESP, and END_RESP as separate messages. This is the early-completion shortcut, and it is how a fast, single-outstanding target avoids the cost of the full handshake while still presenting an nb interface. It is also the conceptual bridge between AT and LT: a target that always early-completes is functionally an LT target wearing an nb socket.

// file: at_early_completion.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
//            -I$SYSTEMC_HOME/include at_early_completion.cpp \
//            -o at_early_completion -L$SYSTEMC_HOME/lib -lsystemc

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>

using namespace tlm;
using namespace sc_core;

// Target that completes on the BEGIN_REQ call: it services the access and
// returns TLM_COMPLETED with phase advanced to BEGIN_RESP, collapsing all
// four phases into one call.
SC_MODULE(FastMem) {
  tlm_utils::simple_target_socket<FastMem> socket;
  unsigned char mem[256];
  SC_CTOR(FastMem) : socket("socket") {
    for (int i = 0; i < 256; i++) mem[i] = 0;
    socket.register_nb_transport_fw(this, &FastMem::nb_transport_fw);
  }
  tlm_sync_enum nb_transport_fw(tlm_generic_payload& trans,
                                tlm_phase& phase, sc_time& delay) {
    if (phase == BEGIN_REQ) {
      sc_dt::uint64  a = trans.get_address();
      unsigned int   l = trans.get_data_length();
      unsigned char* p = trans.get_data_ptr();
      if (trans.get_command() == TLM_WRITE_COMMAND)
        for (unsigned i = 0; i < l; i++) mem[a + i] = p[i];
      else
        for (unsigned i = 0; i < l; i++) p[i] = mem[a + i];
      trans.set_response_status(TLM_OK_RESPONSE);
      delay += sc_time(5, SC_NS);     // annotate the access latency
      phase = BEGIN_RESP;             // tell the caller the response is here...
      return TLM_COMPLETED;           // ...and the whole transaction is done
    }
    return TLM_ACCEPTED;
  }
};

SC_MODULE(Cpu) {
  tlm_utils::simple_initiator_socket<Cpu> socket;
  SC_CTOR(Cpu) : socket("socket") { SC_THREAD(run); }

  void run() {
    tlm_generic_payload trans;   // a STACK payload is safe here: early completion
                                 // means the transaction never outlives the call.
    unsigned w = 0xABCD;
    sc_time delay = SC_ZERO_TIME;
    trans.set_command(TLM_WRITE_COMMAND);
    trans.set_address(0x40);
    trans.set_data_ptr(reinterpret_cast<unsigned char*>(&w));
    trans.set_data_length(4);
    trans.set_streaming_width(4);
    trans.set_byte_enable_ptr(nullptr);
    trans.set_response_status(TLM_INCOMPLETE_RESPONSE);

    tlm_phase phase = BEGIN_REQ;
    tlm_sync_enum s = socket->nb_transport_fw(trans, phase, delay);

    std::cout << "[" << sc_time_stamp() << "] return=";
    if      (s == TLM_COMPLETED) std::cout << "TLM_COMPLETED";
    else if (s == TLM_UPDATED)   std::cout << "TLM_UPDATED";
    else                         std::cout << "TLM_ACCEPTED";
    std::cout << " phase=" << (phase == BEGIN_RESP ? "BEGIN_RESP" : "?")
              << " resp=" << trans.get_response_string()
              << " annotated_delay=" << delay << "\n";

    if (s == TLM_COMPLETED) {
      wait(delay);                   // pay the annotated latency
      std::cout << "[" << sc_time_stamp() << "] transaction complete in one call\n";
    }
  }
};

int sc_main(int, char*[]) {
  Cpu cpu("cpu");
  FastMem mem("mem");
  cpu.socket.bind(mem.socket);
  sc_start();
  return 0;
}

Expected output:

[0 s] return=TLM_COMPLETED phase=BEGIN_RESP resp=TLM_OK_RESPONSE annotated_delay=5 ns
[5 ns] transaction complete in one call

Three things this trace teaches. First, TLM_COMPLETED collapses the handshake: the target serviced the access, set the response, advanced phase to BEGIN_RESP, and returned TLM_COMPLETED — all four phases' worth of meaning in one call, no backward calls at all. The initiator sees completion in the return value and does not send END_RESP (there is no one to send it to; the transaction is already retired). Second, TLM_COMPLETED is orthogonal to success: the initiator still reads get_response_status() to learn it was TLM_OK_RESPONSE. Had the address been out of range, the target could have returned TLM_COMPLETED with TLM_ADDRESS_ERROR_RESPONSE — completed and failed. The sync enum is about the protocol; the response status is about the outcome. Third, a stack payload is safe here: because the transaction completes synchronously inside the one call, the payload never outlives the call, so the use-after-free hazard of the multi-phase model does not apply. Early completion is the one AT path where the LT stack-payload habit still works — which is exactly why it is the natural bridge between the two styles.

Corner 2: payload memory management is mandatory under multi-phase AT

Part 2 introduced reference-counted memory management as an Advanced corner you could usually skip for LT. Under multi-phase AT it is not skippable, and the Beginner section's segfault warning showed why: the payload is dereferenced by callbacks that fire after the launching call returned. The contract, restated for AT, has four obligations:

  1. Construct the payload with a memory manager. new tlm_generic_payload(mm) (or set_mm(mm) after default construction) arms reference counting. A default-constructed payload has no manager and acquire()/release() on it is undefined behavior.
  2. acquire() before you launch the transaction. The initiator acquire()s before sending BEGIN_REQ, establishing "this transaction holds a reference." The Intermediate gen() does exactly this.
  3. release() when you are done, exactly once per acquire(). The last party done with the payload release()s it; conventionally the target releases when it processes END_RESP (or when the initiator completes via return). The reference count hits zero, the payload calls its manager's free(), and the pool recycles it. Our pool's free() calls reset() to drop stale extensions and byte-enable pointers before reuse — the same hygiene Part 2's pool used.
  4. Never touch a payload after your last release(). Once you release, the object may be recycled into another transaction at any time. Reading it after release is a use-after-free even if the bytes happen to still be there.

The full memory-manager and pool code is in the Intermediate example's Pool class — re-read it now with these four obligations in mind, and notice that the initiator acquires while the target releases, so the reference count correctly spans the entire two-party lifetime of the transaction. For the deeper treatment of tlm_mm_interface, acquire/release internals, and why a stack payload is conformant for blocking transport, see the Advanced section of Part 2 — The Generic Payload & Blocking Transport; this post is the place that memory management stops being optional and starts being load-bearing.

Key takeaway The single rule that prevents the most AT bugs: under multi-phase AT the payload must outlive every nb_transport call that references it. That forces a heap payload with a memory manager, acquire() before the first phase, and a single balanced release() after the last. A stack payload is legal only for the early-completion shortcut, where the transaction never outlives the call.

Corner 3: peq_with_cb_and_phase is the production scheduling tool

Every AT model needs to schedule phase transitions for future simulation times without blocking — END_REQ 2 ns from now, BEGIN_RESP 10 ns after that. You could hand-roll this with raw sc_events and a payload-tagged data structure, but you would reinvent, badly, what tlm_utils::peq_with_cb_and_phase already does correctly. It is a payload event queue with phase: you give it a callback at construction, and thereafter peq.notify(trans, phase, delay) arranges for callback(trans, phase) to fire at now + delay, handling both delta-cycle (SC_ZERO_TIME) and timed notifications, and correctly queuing multiple pending transactions. Both the Beginner and Intermediate targets use it; here is the minimal shape in isolation:

#include <tlm_utils/peq_with_cb_and_phase.h>

SC_MODULE(Target) {
  // The PEQ is templated on the module type and bound to a callback method.
  tlm_utils::peq_with_cb_and_phase<Target> m_peq;

  SC_CTOR(Target) : m_peq(this, &Target::peq_cb) { /* ... */ }

  // Somewhere on the forward path, defer a phase instead of answering inline:
  //   m_peq.notify(trans, tlm::END_REQ, sc_time(2, SC_NS));

  // The PEQ invokes this at now + delay, for each notified (trans, phase):
  void peq_cb(tlm::tlm_generic_payload& trans, const tlm::tlm_phase& phase) {
    if (phase == tlm::END_REQ)   { /* send END_REQ on bw path, then schedule BEGIN_RESP */ }
    if (phase == tlm::BEGIN_RESP){ /* send BEGIN_RESP on bw path */ }
  }
};

The reason it is the standard tool, not just a convenience: getting deferred phase scheduling right by hand is subtle. You must coalesce notifications at the same time correctly, preserve per-transaction ordering, and avoid the classic bug of notifying an sc_event that is shared across transactions (which would lose all but the last). The PEQ solves all of that, which is why essentially every production AT target — and every TLM-2.0 example in the SystemC distribution — is built on it. Use it; do not hand-roll event scheduling for AT phases.

Corner 4: who advances the transaction, and the TLM_UPDATED path we did not take

This post's examples used TLM_ACCEPTED plus deferred backward calls for every multi-phase transition — the target says "accepted" and later calls back with the next phase. The base protocol also permits the synchronous TLM_UPDATED path: a peer that can advance the transaction immediately, in the same call, returns TLM_UPDATED with the new phase and delay in the by-reference arguments, and the caller acts on them at once instead of waiting for a callback. Both are conformant; the choice is about whether the next phase is ready synchronously (TLM_UPDATED) or needs to wait for simulated time or an internal condition (TLM_ACCEPTED + later call). Mixing them up is a classic AT bug: returning TLM_ACCEPTED while also having modified phase (the caller will ignore your phase, as the rules require, and your intended transition is silently lost), or returning TLM_UPDATED without setting a meaningful phase (the caller acts on garbage). The discipline: if you return TLM_ACCEPTED, the phase/delay you leave behind are meaningless and a future call carries the transaction; if you return TLM_UPDATED, set phase/delay to the genuine next state; if you return TLM_COMPLETED, the transaction is over regardless of phase.

Version differences

The non-blocking transport interface — nb_transport_fw/nb_transport_bw signatures, the tlm_sync_enum enum, the four base-protocol phases, the request/response exclusion rules, and the memory-management protocol — is identical across SystemC 2.3.0, 2.3.1, 2.3.3, 2.3.4, and 3.0.x. The tlm, tlm_utils, and peq_with_cb_and_phase headers ship bundled with the kernel from 2.3.0 onward. SystemC 3.0.x tracks IEEE 1666-2023, which deprecates the SC_HAS_PROCESS(name) macro form (the SC_CTOR/SC_MODULE forms in these examples avoid it) and emits a deprecation note if you use it, but no nb_transport or base-protocol semantics changed. Every example here compiles and runs identically on any 2.3.x or 3.0.x build; the -DSC_ALLOW_DEPRECATED_IEEE_API flag in the build line silences benign deprecation warnings from the bundled headers on 3.0.x and is harmless on older versions.

When AT is worth it versus when to stay in LT

Everything in this post costs something. AT models are several times the code of their LT equivalents, run slower per transaction (more calls, scheduling, and bookkeeping), and are far easier to get subtly wrong. So the engineering question is never "AT or LT in the abstract" but "does this model need what AT buys?" The decision table:

Concern Use LT (b_transport) Use AT (nb_transport)
Booting an OS / running software Yes — speed dominates No — needless overhead
Functional correctness only Yes No
Transaction throughput matters but not overlap Yes (annotate latency) No
Pipelining / multiple outstanding transactions Cannot express it Yes — this is the reason AT exists
Backpressure / finite request buffers / flow control Cannot express it Yes
Contention timing on a shared interconnect Approximate at best Yes — phases model it
Arbitration latency, response reordering No Yes
Coding effort / maintenance Low High
Simulation speed Fast (esp. with temporal decoupling) Slower

The practical rule: start in LT, and move a component to AT only when you have a specific timing question that requires modeling overlap, backpressure, or contention — and only for the components on the path of that question. A virtual platform is almost never all-AT; it is LT everywhere except the one interconnect or memory controller whose timing you are actually studying, which is modeled at AT and bridged to the LT world by adapters (the simple_target_socket even provides a built-in b/nb conversion so an LT initiator can talk to an AT target and vice versa). AT is a precision instrument, not a default. Reach for it deliberately, scope it narrowly, and keep the rest of the platform in the fast, simple LT style you learned in Part 2.

Hands-on exercise

Extend the pipelined AT initiator/target from the Intermediate section to add response-side backpressure, exercising the one phase the example treats as instantaneous.

In the Intermediate model, the initiator accepts every BEGIN_RESP immediately and returns END_RESP in the same backward call (zero response-acceptance latency). Real initiators have finite response buffers and cannot always accept a response the instant it arrives — they apply backpressure on the response side, exactly mirroring the request-side backpressure the target already does. Your task is to build that.

Modify the initiator so that:

  1. It models a response buffer of depth 1: it can hold at most one response awaiting acceptance at a time.
  2. When BEGIN_RESP arrives and the buffer is free, the initiator does not send END_RESP immediately. Instead it parks the response and schedules END_RESP for 3 ns later (model response-acceptance latency), then sends END_RESP from a deferred callback.
  3. While a response is parked (the buffer is occupied), a second BEGIN_RESP must be made to wait — the initiator withholds acceptance until the parked one is retired, mirroring the target's request-side parking.
  4. Print the simulation time at each BEGIN_RESP arrival and each END_RESP send, so you can see responses queuing behind the 3 ns acceptance latency.

Predict, before you run: with the target producing responses 2 ns apart (as the Intermediate trace shows responses at 12 ns and 14 ns) and the initiator taking 3 ns to accept each, the second response must wait for the first to clear. Sketch the expected END_RESP times on paper, then check against the run.

When that works, extend it further: make the target react to the slower response drain. Because END_RESP now arrives later, request slots free later, so the request-side backpressure should kick in sooner and more often. Confirm from the trace that adding response-side backpressure increases the request-side stalling — the two flow-control mechanisms compose, exactly as they would in real hardware where a slow consumer eventually stalls the producer.

Hints

  • You need a second peq_with_cb_and_phase in the initiator to schedule the deferred END_RESP, just as the target uses one to schedule END_REQ/BEGIN_RESP. The initiator's PEQ callback sends nb_transport_fw(trans, END_RESP, ...).
  • Track a bool resp_busy (or an int count capped at 1) in the initiator. Set it when you park a BEGIN_RESP, clear it when the deferred END_RESP fires, and at that point admit any BEGIN_RESP you queued while busy.
  • When a second BEGIN_RESP arrives while resp_busy, push its payload onto an initiator-side backlog std::queue and return TLM_ACCEPTED without scheduling its END_RESP yet — that withholding is the response backpressure.
  • Do not call wait() anywhere in the nb_transport methods — defer with the PEQ, exactly as the target does. The non-blocking contract still holds.
  • Watch your acquire()/release() balance: the payload must now stay alive through the deferred END_RESP, which is later than before. The target still release()s on END_RESP; make sure that END_RESP is the deferred one, so the payload is not recycled while still parked in the initiator.
  • Remember the request/response exclusion is per transaction: parking a response does not stop other transactions' requests from flowing, so the request and response backlogs are independent queues.

No solution is provided. Getting the second PEQ, the response backlog, and the acquire/release lifetime right yourself is the entire point — it is the same machinery as the request side, applied to the response side, and building it is how AT stops being theory.

Common mistakes

  • Treating TLM_COMPLETED as "the transaction succeeded." It means "no more phase transitions," not "it worked." Success or failure lives in get_response_status(), independently. A target can return TLM_COMPLETED with TLM_ADDRESS_ERROR_RESPONSE — finished and failed. Fix: always check the response status separately from the sync enum, exactly as you did in LT.
  • Using a stack payload for a multi-phase transaction. Under multi-phase AT the payload is dereferenced by callbacks (the PEQ, the backward path) that fire after the launching call returned, so a stack local is a use-after-free / segfault. Fix: heap-allocate from a pool, construct with a memory manager, acquire() before BEGIN_REQ, and release() exactly once after END_RESP. A stack payload is legal only for the early-completion shortcut, where the transaction never outlives the call.
  • Mismatched acquire()/release(). An extra acquire() leaks the payload (it never returns to the pool); an extra release() recycles it while another party still references it (use-after-free). Fix: exactly one release() per acquire(); conventionally the initiator acquires before the request and the target releases on END_RESP, so the count spans the whole two-party lifetime.
  • Calling wait() inside an nb_transport method. Non-blocking means non-blocking: these methods must return immediately and never consume simulation time. Calling wait() inside one is a runtime error and breaks the entire phase model. Fix: express all latency by annotating delay and scheduling future phases with a peq_with_cb_and_phase; never block.
  • Confusing TLM_ACCEPTED with TLM_UPDATED. TLM_ACCEPTED says "ignore my returned phase/delay; a future call carries the transaction." TLM_UPDATED says "my returned phase/delay are the real next state; act on them now." Returning TLM_ACCEPTED after modifying phase silently loses your intended transition; returning TLM_UPDATED without a meaningful phase makes the caller act on garbage. Fix: match the return code to whether the next phase is ready synchronously.
  • Violating the request/response exclusion / phase ordering. Sending BEGIN_RESP before END_REQ, or issuing a new BEGIN_REQ on a hop before the previous request's END_REQ came back, breaks the base protocol and confuses interconnects and checkers. Fix: hold the one legal order — BEGIN_REQ → END_REQ → BEGIN_RESP → END_RESP per transaction — and gate the initiator's next request on receiving END_REQ (as the Intermediate gen() does with wait(can_send)).
  • Hand-rolling phase scheduling with a shared sc_event. Notifying one sc_event for multiple in-flight transactions loses all but the last notification, silently dropping transactions. Fix: use peq_with_cb_and_phase, which queues per-transaction notifications correctly. Do not reinvent it.
  • Reaching for AT when LT would do. AT is several times the code and slower per transaction; using it where you do not need overlap, backpressure, or contention modeling is pure cost. Fix: stay in LT by default; move only the specific component whose timing you are studying to AT, and bridge it to the LT world with the socket's built-in b/nb conversion.

Recap

After working through this post you can now:

  • State the core idea — non-blocking transport splits the single b_transport call into phases so control can return to the initiator before the transaction completes — and explain why that split is what enables overlap, pipelining, and backpressure that b_transport structurally cannot express.
  • Name the four base-protocol phases (BEGIN_REQ, END_REQ, BEGIN_RESP, END_RESP), state their one legal ordering, and explain the request/response exclusion rule (the request phase closes before the response phase opens, per transaction).
  • Distinguish the three tlm_sync_enum return codes — TLM_ACCEPTED ("ignore my return args, a callback follows"), TLM_UPDATED ("my return args are the next state, act now"), and TLM_COMPLETED ("the transaction is over") — and say what each obligates the caller to do.
  • Explain why TLM_COMPLETED is a protocol statement, not a status statement, and check the response status independently of the sync enum.
  • Write a pipelined AT initiator and target that keep multiple transactions in flight, and apply backpressure by withholding END_REQ when the request pipeline is full — and read a trace to prove the overlap and the stall.
  • Use the early-completion shortcut (TLM_COMPLETED on BEGIN_REQ) to collapse all four phases into one call, and recognize it as the bridge between AT and LT.
  • Manage payload memory correctly under AT: construct with a memory manager, acquire() before the first phase, release() exactly once after the last, recycle through a pool, and never touch a payload after release — and explain why a stack payload that was safe in LT is a use-after-free under multi-phase AT.
  • Use peq_with_cb_and_phase to schedule deferred phase transitions without blocking, instead of hand-rolling sc_event plumbing.
  • Decide when AT is worth its cost versus when to stay in fast, simple LT, and scope AT narrowly to the components whose timing you are actually studying.

Further reading

Standards

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §15.2 (base protocol: phases §15.2.3, request/response exclusion and ordering §15.2.4, payload memory management under AT §15.2.5, timing annotation §15.2.6) and §15.4 (nb_transport_fw/nb_transport_bw signatures and tlm_sync_enum semantics).

Vendor and consortium documents

  • Aynsley (Doulos), OSCI TLM-2.0 Language Reference Manual (JA32) — the canonical narrative treatment of the base protocol, the four phases, the sync-enum return values, and the AT coding style with pipelining and backpressure.
  • Doulos TLM-2.0 Tutorial — worked AT initiator/target examples and the base-protocol phase-ordering checklist.
  • Accellera Systems Initiative, SystemC 3.0.0 distribution, include/tlm_core/tlm_2/tlm_2_interfaces/tlm_fw_bw_ifs.h, tlm_core/tlm_2/tlm_generic_payload/tlm_phase.h, and include/tlm_utils/peq_with_cb_and_phase.h — authoritative source for the nb interfaces, the phase enum, and the payload-event-queue.

Textbooks

  • Grötker, Liao, Martin, and Swan, System Design with SystemC — transaction-level modeling chapters; the AT-versus-LT trade-off and the role of phases in timing-accurate modeling.

Next in this section

→ Part 5: Temporal Decoupling & the Quantum Keeper — how loosely-timed models run ahead of simulation time to go fast: the global quantum, tlm_quantumkeeper, the local-time offset that the delay argument has been hinting at since Part 2, and the DMI fast path for direct memory access. After the precision of AT, Part 5 returns to LT and makes it fast. Read it here: 20. SystemC Tutorial — Temporal Decoupling & the Quantum Keeper.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment