20. SystemC Tutorial - Quantum Keeper & Temporal Decoupling

Why this matters

Written 2026-06-05 for the concept-first series.

Every loosely-timed virtual platform you build has a speed ceiling, and most engineers hit it long before they understand what is actually holding the simulation back. They assume the cost is the modeled work — the bytes copied, the instructions decoded, the arithmetic the model performs — and they spend days optimizing that. It is almost never where the time goes. The dominant cost of a loosely-timed (LT) SystemC simulation is context switching: every single time a process calls wait() to let simulation time advance, the SystemC kernel has to suspend that process's coroutine, run its scheduler to decide who goes next, and resume some other process's coroutine. That suspend-schedule-resume cycle is pure overhead — none of it models anything — and a naive LT model pays it on every transaction. A processor model that issues a hundred million memory accesses, each followed by a wait() for the access latency, performs a hundred million context switches whose only product is moving the simulation clock forward by a few nanoseconds at a time. The accesses themselves — the actual b_transport calls that do the modeled work — are cheap by comparison. The kernel machinery around them is what makes the simulation crawl.

Temporal decoupling is the technique that removes almost all of that overhead, and it is the single most important performance idea in transaction-level modeling. The mechanism is disarmingly simple once you have the timing-annotation material from the generic-payload post — the sc_time& delay argument that a target annotates and a caller pays. There we showed two idioms: pay the delay immediately with wait(delay), or accumulate it across several calls and pay the sum once. Temporal decoupling is the second idiom taken to its logical conclusion: an initiator runs ahead of the kernel's clock on a private local time offset, banking up annotated latency without ever calling wait(), and only synchronizing — settling its offset into real simulation time with one wait() — when the banked offset reaches a budget called the quantum. A quantum of one microsecond on a model that takes ten nanoseconds per access turns a hundred wait() calls into one. That is a hundredfold reduction in the kernel overhead that was dominating your runtime, for the cost of a few lines of arithmetic and a defensible loss of timing observation resolution that — handled correctly — does not change a single timing number your model computes.

This post teaches that idea from first principles and then hands you the standard tool, tlm_utils::tlm_quantumkeeper, that packages it correctly. You will first build temporal decoupling by hand — accumulate the delay, compare against a threshold, wait() only when you cross it — and you will measure the speedup with real context-switch counts at three different quanta, watching the sync count collapse from 100 to 10 to 1 while the final simulated time stays exactly the same. That invariant — the timing math does not change, only the observation granularity — is the load-bearing insight of the whole post, and most engineers get it backwards on their first encounter. Then you will watch two decoupled initiators reorder their observable prints as the quantum grows, and we will reason carefully about what that reordering does and does not mean for causality. You will rewrite the manual machinery with the quantum keeper and confirm it reproduces your hand-built numbers exactly. And in the Advanced section you will meet the Direct Memory Interface (DMI) — the other half of LT performance — which lets an initiator bypass b_transport entirely for plain memory, why it gives another large speedup, and the hard rule that makes it safe: DMI is legal only for storage with no side effects, because the whole point is that it skips the code where side effects would live. By the end you will know exactly which knob to turn to make an LT model fast, how far you can turn it before accuracy degrades, and how to reason about the accuracy you are trading away.

Prerequisites

  • Part 1 — Why TLM Exists. You need the motivating picture: why transaction-level modeling trades pin-and-cycle accuracy for simulation speed, and the loosely-timed (LT) versus approximately-timed (AT) coding-style distinction. Temporal decoupling is the LT speed technique, so the LT mindset from Part 1 is the ground this post stands on.
  • Part 2 — The Generic Payload & Blocking Transport. This post builds directly on that one's timing-annotation material: the sc_time& delay argument, the rule that targets annotate latency and initiators pay it, and the contrast between the immediate-wait idiom and the accumulate-and-pay idiom. If "the target adds to delay, the caller decides when to consume it" is not second nature, re-read that section before continuing — temporal decoupling is exactly the accumulate idiom industrialized.
  • Foundations Part 2 — Simulation Time. You need to be fluent in what sc_time_stamp() reports, how wait(sc_time) advances the simulation clock, and the relationship between the kernel's notion of "now" and the delta/timed scheduling cycles. Temporal decoupling is precisely the art of letting an initiator's modeled time run ahead of the kernel's sc_time_stamp(), so you must be clear on what that stamp means before you start desynchronizing from it.

One idea from the generic-payload post does all the work here: the delay argument is the initiator's local clock, not a per-call scratch value. Everything below is the consequence of taking that seriously.

Mental model (first principles)

Start with the most reductive correct statement of the problem, then build the solution on it.

A loosely-timed simulation's cost is dominated by synchronization, not by work. When an initiator does "access memory, then advance time by the access latency," the access is a cheap function call but the advance time is a wait(), and wait() is expensive: it suspends the calling process's coroutine, returns control to the SystemC scheduler, which examines its run queues, picks the next runnable process, and resumes its coroutine. That round trip through the kernel is the context switch. If every transaction ends in a wait(), the simulation performs one context switch per transaction, and for a model that issues transactions by the million, the kernel scheduling — not the modeled behavior — is where the wall-clock seconds go.

Now the fix, stated just as plainly. An initiator does not have to advance the shared clock after every access. It can keep its own private clock, run ahead on it, and only reconcile with the shared clock occasionally. Picture two field engineers, Ana and Ben, each doing a sequence of timed tasks, who are required to agree on a shared wall clock. The naive protocol is: after every task, both walk to the central clock, sync their watches, and walk back. They spend more time walking to the clock than doing tasks. The decoupled protocol is: each keeps a wristwatch, does a whole batch of tasks advancing only their own wristwatch, and they meet at the central clock only at scheduled intervals — say, every simulated microsecond — to reconcile. The tasks each performs, and the time each task takes, are identical under both protocols. What changes is how often they walk to the clock (the context switches) and how far each one's wristwatch can drift ahead of the other's before they meet (the observation skew). That is temporal decoupling in one image: a private wristwatch (the local time offset), batches of work advancing only that watch, and scheduled reconciliations (syncs) at a budgeted interval (the quantum).

The mechanical realization is the timing-annotation delay from the generic-payload post, reinterpreted. There, delay was an sc_time you passed into b_transport, the target added its latency to, and you then paid with wait(delay). The reinterpretation is: delay is the initiator's wristwatch — its local time offset from the kernel clock. Each access, the initiator passes its current offset in as delay; the target adds its latency; the initiator stores the bigger offset back. No wait() happens. The initiator's modeled "now" is sc_time_stamp() + local_offset — the kernel clock plus how far the wristwatch has run ahead. The offset just keeps growing as the initiator does more work, while sc_time_stamp() stays frozen. Only when the offset reaches the quantum — the budgeted maximum drift — does the initiator call wait(local_offset), which advances the kernel clock by the whole accumulated amount in one context switch and resets the offset to zero. A quantum of 1 us over 10 ns accesses means the initiator does 100 accesses on its wristwatch and pays them all with a single wait(), replacing 100 context switches with 1.

Three facts about this picture are non-negotiable, and getting any of them wrong produces a broken model:

  • The timing math is unchanged. Each access annotates the same latency regardless of the quantum. The initiator's modeled "now" after N accesses is the same number whether it synced after each one or banked them all. The quantum changes when the kernel clock catches up, not what the total elapsed modeled time is. We will prove this with measured runs: identical final time at every quantum.
  • What the quantum changes is observation granularity and ordering. Between syncs, the kernel clock is frozen, so anything that reads sc_time_stamp() (a logger, a peer process) sees the initiator's progress only in quantum-sized jumps. And because each decoupled initiator runs ahead independently, two of them can reach the same modeled time at different real scheduling moments, so their events get observed in an order that the quantum determines. Their individual timelines are exact; the interleaving the kernel witnesses is quantized.
  • The target still only annotates; it never waits. The "targets annotate, initiators pay" rule from the generic-payload post is what makes decoupling possible. A target that called wait() internally would force a context switch on every access and collapse the quantum to zero. Decoupling lives entirely on the initiator side; the target is unchanged.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#dbeafe', 'primaryTextColor': '#1e293b', 'primaryBorderColor': '#3b82f6', 'lineColor': '#64748b', 'secondaryColor': '#f1f5f9'}}}%% sequenceDiagram participant I as Initiator (local offset) participant T as Target (annotate only) participant K as Kernel clock loop until offset >= quantum I->>T: b_transport(trans, offset) T->>T: offset += latency (NO wait) T-->>I: return (kernel clock frozen) end I->>K: wait(offset) -- one context switch K->>K: advance by whole offset I->>I: offset = 0

Read the diagram as the wristwatch story: many accesses bank latency into the offset with the kernel clock frozen, then one wait() settles the whole batch into real time. The quantum is the loop's exit condition — make it bigger and the loop runs longer before the single expensive wait().

The rest of this post makes each of those three facts concrete and measured: the Beginner section builds the wristwatch by hand and instruments it; the Intermediate section swaps in the standard tlm_quantumkeeper and confronts the accuracy trade-off; the Advanced section adds DMI, the other half of LT speed.

Beginner: First Principles

We will build temporal decoupling with no library help at all — just the delay argument, a threshold compare, and a wait() — so that the quantum keeper in the next section is obviously just packaging and not magic. The system is one initiator issuing a hundred writes to a 256-byte memory that charges 10 ns per access. The initiator accumulates the annotated delay in a local offset and calls wait() only when the offset crosses the quantum. We instrument two counters: how many b_transport calls happen (the modeled work) and how many wait() calls happen (the context switches — the cost we are trying to cut).

// file: qk_manual.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
//          -I$SYSTEMC_HOME/include qk_manual.cpp \
//          -L$SYSTEMC_HOME/lib -lsystemc -o qk_manual
// Run:   DYLD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./qk_manual <quantum_ns>

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>

static unsigned long g_b_transport_calls = 0; // payload-handling work
static unsigned long g_context_switches  = 0; // wait() calls = sync points

// Memory target: annotates 10 ns per access, NEVER waits.
SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  unsigned char mem[256];
  SC_CTOR(Memory) : socket("socket") {
    for (int i = 0; i < 256; i++) mem[i] = 0;
    socket.register_b_transport(this, &Memory::b_transport);
  }
  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    g_b_transport_calls++;
    sc_dt::uint64  addr = trans.get_address() & 0xFF;
    unsigned int   len  = trans.get_data_length();
    unsigned char* ptr  = trans.get_data_ptr();
    if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
      for (unsigned i = 0; i < len; i++) mem[(addr + i) & 0xFF] = ptr[i];
    else
      for (unsigned i = 0; i < len; i++) ptr[i] = mem[(addr + i) & 0xFF];
    delay += sc_time(10, SC_NS);            // 10 ns latency, ANNOTATED only
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

// Initiator that runs ahead on a local clock, syncing only at the quantum.
SC_MODULE(Cpu) {
  tlm_utils::simple_initiator_socket<Cpu> socket;
  sc_time quantum;
  SC_HAS_PROCESS(Cpu);
  Cpu(sc_module_name n, sc_time q)
    : sc_module(n), socket("socket"), quantum(q) {
    SC_THREAD(run);
  }
  void run() {
    tlm::tlm_generic_payload trans;
    unsigned int data = 0;
    sc_time local = SC_ZERO_TIME;          // the wristwatch: local time offset
    for (int n = 0; n < 100; n++) {        // 100 accesses
      trans.set_command(tlm::TLM_WRITE_COMMAND);
      trans.set_address((n * 4) & 0xFF);
      trans.set_data_ptr(reinterpret_cast<unsigned char*>(&data));
      trans.set_data_length(4);
      trans.set_streaming_width(4);
      trans.set_byte_enable_ptr(nullptr);
      trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
      socket->b_transport(trans, local);   // target ADDS to the offset
      if (local >= quantum) {              // threshold reached: synchronize
        wait(local);                       // pay the WHOLE accumulated offset
        g_context_switches++;
        local = SC_ZERO_TIME;              // wristwatch reset
      }
    }
    if (local > SC_ZERO_TIME) { wait(local); g_context_switches++; } // settle tail
  }
};

int sc_main(int argc, char* argv[]) {
  double q_ns = (argc > 1) ? atof(argv[1]) : 100.0;
  sc_time q(q_ns, SC_NS);
  Cpu    cpu("cpu", q);
  Memory mem("mem");
  cpu.socket.bind(mem.socket);
  sc_start();
  std::cout << "quantum = " << q
            << "  final time = " << sc_time_stamp()
            << "  b_transport calls = " << g_b_transport_calls
            << "  context switches (wait) = " << g_context_switches << "\n";
  return 0;
}

Compile it once, then run it three times with different quanta — 10 ns, 100 ns, and 1 us — and predict the three lines before you read them.

Expected output:

quantum = 10 ns  final time = 1 us  b_transport calls = 100  context switches (wait) = 100
quantum = 100 ns  final time = 1 us  b_transport calls = 100  context switches (wait) = 10
quantum = 1 us  final time = 1 us  b_transport calls = 100  context switches (wait) = 1

Read those three lines slowly, because they contain the entire thesis of the post. The final simulated time is 1 us in all three runs — one hundred accesses at 10 ns each is 1000 ns, period, regardless of the quantum. The b_transport call count is 100 in all three — the modeled work is identical; every access still happens and still copies its bytes. The only thing that changes is the context-switch count: 100, then 10, then 1. At quantum 10 ns the offset hits the threshold after every single access (each access banks exactly 10 ns), so the initiator waits 100 times — this is the naive, fully-synchronized LT model, the one with no decoupling benefit. At quantum 100 ns the offset only crosses the threshold every tenth access (10 accesses × 10 ns = 100 ns), so the initiator waits 10 times — a 10x reduction in kernel overhead for free. At quantum 1 us the offset never crosses until all 100 accesses are banked, so the initiator waits exactly once — a 100x reduction. Same work, same modeled time, 100x fewer context switches. That is temporal decoupling, and the measured numbers say it plainly: the quantum is a pure performance knob over the synchronization cost, and it does not touch the timing result.

It is worth being precise about why the final time is invariant, because the intuition "I'm skipping waits, so surely time comes out different" is exactly the misconception to kill. The latency is annotated into local on every access no matter what. When the initiator finally calls wait(local), it advances the kernel clock by the whole banked amount — it does not lose the nanoseconds it skipped waiting for, it pays them in a lump. Banking 10 ns ten times and paying 100 ns once advances the clock by exactly the same 100 ns as paying 10 ns ten times; addition is associative. The wait() calls you removed were never adding time — they were only reconciling the offset into the kernel clock more frequently than necessary. Remove the redundant reconciliations and the arithmetic is untouched.

Note The target in this example never calls wait() — it only does delay += sc_time(10, SC_NS). That is what makes decoupling work. If the target waited internally, the initiator's local offset could never grow past one access before the kernel clock advanced, the local >= quantum test would fire every time, and you would be back to 100 context switches no matter how large you set the quantum. "Targets annotate, initiators pay" from the generic-payload post is the precondition for everything here.

The other half: observation reordering between two initiators

The single-initiator sweep shows the speedup but hides the cost, because with one initiator there is nothing to observe out of order. Add a second decoupled initiator and the quantum's effect on observation becomes visible. The next program runs two initiators, A and B, each doing four accesses to its own port of a two-port memory, each printing its modeled time (sc_time_stamp() + local) after every access. (We give the memory two target sockets so each initiator binds its own; routing multiple initiators through one shared port is a bus-fabric topic for a later post. Here we only want to watch how the two initiators' prints interleave.)

// file: qk_two_initiators.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
//          -I$SYSTEMC_HOME/include qk_two_initiators.cpp \
//          -L$SYSTEMC_HOME/lib -lsystemc -o qk_two_initiators
// Run:   DYLD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./qk_two_initiators <quantum_ns>

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>

// Two-port memory so each initiator binds its own target socket.
SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> sock0;
  tlm_utils::simple_target_socket<Memory> sock1;
  SC_CTOR(Memory) : sock0("sock0"), sock1("sock1") {
    sock0.register_b_transport(this, &Memory::b_transport);
    sock1.register_b_transport(this, &Memory::b_transport);
  }
  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    delay += sc_time(10, SC_NS);
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

SC_MODULE(Cpu) {
  tlm_utils::simple_initiator_socket<Cpu> socket;
  sc_time quantum;
  const char* tag;
  SC_HAS_PROCESS(Cpu);
  Cpu(sc_module_name n, sc_time q, const char* t)
    : sc_module(n), socket("socket"), quantum(q), tag(t) { SC_THREAD(run); }
  void run() {
    tlm::tlm_generic_payload trans;
    unsigned int data = 0;
    sc_time local = SC_ZERO_TIME;
    for (int n = 0; n < 4; n++) {
      trans.set_command(tlm::TLM_WRITE_COMMAND);
      trans.set_address(0);
      trans.set_data_ptr(reinterpret_cast<unsigned char*>(&data));
      trans.set_data_length(4);
      trans.set_streaming_width(4);
      trans.set_byte_enable_ptr(nullptr);
      trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
      socket->b_transport(trans, local);
      // modeled "now" = kernel clock + how far the wristwatch ran ahead
      std::cout << tag << " access " << n
                << " models time " << (sc_time_stamp() + local) << "\n";
      if (local >= quantum) { wait(local); local = SC_ZERO_TIME; }
    }
    if (local > SC_ZERO_TIME) wait(local);
  }
};

int sc_main(int argc, char* argv[]) {
  double q_ns = (argc > 1) ? atof(argv[1]) : 10.0;
  sc_time q(q_ns, SC_NS);
  Cpu    a("A", q, "A");
  Cpu    b("B", q, "B");
  Memory mem("mem");
  a.socket.bind(mem.sock0);
  b.socket.bind(mem.sock1);
  std::cout << "=== quantum = " << q << " ===\n";
  sc_start();
  return 0;
}

Run it at quantum 10 ns and then at quantum 100 ns.

Expected output:

=== quantum = 10 ns ===
A access 0 models time 10 ns
B access 0 models time 10 ns
A access 1 models time 20 ns
B access 1 models time 20 ns
A access 2 models time 30 ns
B access 2 models time 30 ns
A access 3 models time 40 ns
B access 3 models time 40 ns
=== quantum = 100 ns ===
A access 0 models time 10 ns
A access 1 models time 20 ns
A access 2 models time 30 ns
A access 3 models time 40 ns
B access 0 models time 10 ns
B access 1 models time 20 ns
B access 2 models time 30 ns
B access 3 models time 40 ns

Look at what stayed the same and what changed. The modeled-time stamps are identical between the two runs — A's accesses are at 10, 20, 30, 40 ns and B's at 10, 20, 30, 40 ns in both quanta. No timing number moved. What changed is the order the kernel printed them in. At quantum 10 ns, each initiator banks only 10 ns before it hits the threshold and wait()s, handing control back to the kernel, which then runs the other initiator; so the two interleave tightly — A0, B0, A1, B1, … — and the print order tracks modeled time closely. At quantum 100 ns, each initiator's four accesses bank only 40 ns total, never reaching the 100 ns threshold, so neither one ever wait()s until it finishes all four. The scheduler runs A's entire batch first (A0–A3, all on A's wristwatch with the kernel clock frozen at 0), then B's entire batch. The prints reorder: all of A, then all of B.

This is the cost of temporal decoupling, and it is essential to name it precisely. Each initiator's own timeline is exact and unchanged — A's accesses really are modeled at 10/20/30/40 ns; B's really are at 10/20/30/40 ns; nothing about either one's computed timing is wrong. What the large quantum changes is the interleaving that the kernel — and therefore any observer — witnesses. With a coarse quantum, A runs a long way ahead of B in real scheduling order even though they occupy the same modeled-time window, so events that a fine-grained simulation would have interleaved get observed in a batched, reordered sequence.

Whether that reordering matters depends entirely on whether the two initiators interact within a quantum. If A and B never touch shared state between syncs — two independent cores each pounding their own private memory — the reordering is invisible and harmless; the final state is identical. The danger is when they do interact: if A writes a shared location at modeled 25 ns and B reads it at modeled 30 ns, a fine quantum interleaves them correctly (A's write is observed before B's read), but a coarse quantum that lets B run its whole batch before A runs at all would have B read the stale value, because in real scheduling order B's read executed before A's write — even though in modeled time the write is earlier. This is the causality hazard of temporal decoupling, and it is why the rule "sync before you touch state another decoupled process is racing to produce" exists. We return to it in the accuracy discussion. For now, fix the two observations in your mind: the quantum never changes a modeled timestamp; it changes the order events are observed, which is harmless for independent initiators and a correctness hazard for interacting ones.

Intermediate: How It Really Works

The manual machinery in the Beginner section — declare an sc_time local, pass it as delay, compare against the quantum, wait() and reset when it crosses — is correct, but it has three problems in real models. First, you have to repeat that arithmetic in every initiator, and a copy-paste error (forgetting the reset, comparing > instead of >=, forgetting the tail sync) silently corrupts timing. Second, the quantum is usually a global property of the simulation — you want every initiator on the same budget, set in one place — and a hand-rolled sc_time quantum member per module makes that awkward. Third, the local time offset interacts with the global quantum in a subtle way at the boundaries (an initiator should not overshoot the quantum boundary), and the careful version of the arithmetic is fiddly. TLM-2.0 packages all of this in one small class, tlm_utils::tlm_quantumkeeper, and using it is the standard, expected, reviewable way to write a decoupled initiator. This section rewrites the manual sweep with the keeper and confirms it produces the same measured numbers.

The quantum keeper idiom

The keeper owns one initiator's local time offset and answers the question "is it time to sync?" Its method set maps directly onto the manual code you already wrote:

  • tlm_quantumkeeper::set_global_quantum(t) — a static call that sets the shared quantum for the whole simulation. Call it once in sc_main. Every keeper reads this global value.
  • reset() — zero the local offset and recompute where the next quantum boundary falls. Call it once before the initiator's loop.
  • get_local_time() — return the current local offset. You pass this into b_transport as the delay, because the offset is the delay.
  • set(t) — store an absolute offset back into the keeper. After b_transport returns with the target's latency added to delay, you call qk.set(delay) to absorb the new offset. (There is also inc(t) to add an increment; set taking the post-call delay is the idiom that matches "the target accumulated into the delay argument.")
  • need_sync() — return true once the local offset has reached the quantum boundary. This replaces your local >= quantum test.
  • sync() — call wait(local_offset) to advance the kernel clock by the banked offset, then reset(). This replaces your wait(local); local = SC_ZERO_TIME; pair.

The canonical per-access loop body is therefore: take the current offset as the delay, do the transaction, absorb the new offset, and sync if the budget is spent.

// file: qk_keeper.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
//          -I$SYSTEMC_HOME/include qk_keeper.cpp \
//          -L$SYSTEMC_HOME/lib -lsystemc -o qk_keeper
// Run:   DYLD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./qk_keeper <quantum_ns>

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <tlm_utils/tlm_quantumkeeper.h>
#include <iostream>

static unsigned long g_calls = 0;
static unsigned long g_syncs = 0;

SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  unsigned char mem[256];
  SC_CTOR(Memory) : socket("socket") {
    socket.register_b_transport(this, &Memory::b_transport);
  }
  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    g_calls++;
    delay += sc_time(10, SC_NS);           // annotate; never wait
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
};

SC_MODULE(Cpu) {
  tlm_utils::simple_initiator_socket<Cpu> socket;
  tlm_utils::tlm_quantumkeeper qk;         // owns this initiator's offset
  SC_CTOR(Cpu) : socket("socket") { SC_THREAD(run); }
  void run() {
    tlm::tlm_generic_payload trans;
    unsigned int data = 0;
    qk.reset();                            // local offset = 0
    for (int n = 0; n < 100; n++) {
      sc_time delay = qk.get_local_time(); // offset IS the delay
      trans.set_command(tlm::TLM_WRITE_COMMAND);
      trans.set_address((n * 4) & 0xFF);
      trans.set_data_ptr(reinterpret_cast<unsigned char*>(&data));
      trans.set_data_length(4);
      trans.set_streaming_width(4);
      trans.set_byte_enable_ptr(nullptr);
      trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
      socket->b_transport(trans, delay);   // target adds latency to delay
      qk.set(delay);                       // absorb the new offset
      if (qk.need_sync()) { qk.sync(); g_syncs++; }  // wait() past the quantum
    }
    qk.sync();                             // settle the tail
  }
};

int sc_main(int argc, char* argv[]) {
  double q_ns = (argc > 1) ? atof(argv[1]) : 100.0;
  // ONE place sets the budget for the whole simulation:
  tlm_utils::tlm_quantumkeeper::set_global_quantum(sc_time(q_ns, SC_NS));
  Cpu    cpu("cpu");
  Memory mem("mem");
  cpu.socket.bind(mem.socket);
  sc_start();
  std::cout << "quantum = " << sc_time(q_ns, SC_NS)
            << "  final time = " << sc_time_stamp()
            << "  b_transport calls = " << g_calls
            << "  syncs (wait) = " << g_syncs << "\n";
  return 0;
}

Run it at the same three quanta.

Expected output:

quantum = 10 ns  final time = 1 us  b_transport calls = 100  syncs (wait) = 100
quantum = 100 ns  final time = 1 us  b_transport calls = 100  syncs (wait) = 10
quantum = 1 us  final time = 1 us  b_transport calls = 100  syncs (wait) = 1

These are exactly the manual numbers — final time 1 us at every quantum, 100 b_transport calls, and syncs collapsing 100 → 10 → 1. That is the point: the quantum keeper is not a different algorithm, it is the same accumulate-and-sync arithmetic you wrote by hand, packaged so you cannot get it wrong and so the quantum is set globally in one place. The mapping is one-to-one — get_local_time()/set() are your local offset, need_sync() is your local >= quantum, sync() is your wait(local); local = SC_ZERO_TIME;. Once you have seen the manual version, the keeper holds no mystery; it is the reviewable form of code you already understand.

Two idiom details are worth stating because they are the parts people get wrong. First, you pass qk.get_local_time() in as the delay and qk.set(delay) the result back on every access. The keeper does not intercept the b_transport call; it just holds the offset, and you are responsible for threading that offset through the delay argument each time. Forgetting the qk.set(delay) after the call means the keeper never sees the latency the target added, need_sync() never fires, and the model runs the entire simulation on a frozen kernel clock — a classic decoupling bug. Second, the trailing qk.sync() after the loop is mandatory. When the loop ends, the keeper may still hold an unspent offset (the last few accesses that did not reach a quantum boundary). Without the final sync, that banked time is never paid into the kernel clock, and your final sc_time_stamp() is short. The manual version had the same if (local > SC_ZERO_TIME) wait(local); tail for exactly this reason. Always settle the tail.

Tip tlm_quantumkeeper also offers set_and_sync(t) (absorb an offset and sync in one call) and inc(t) (add an increment rather than set absolutely). The get_local_time() → b_transport → set(delay) → need_sync()/sync() sequence shown here is the form that lines up with "the target accumulated latency into the delay argument," which is why it is the one you will see in Doulos examples and most production initiators. Learn this one first; the others are conveniences over it.

Accuracy trade-offs: choosing the quantum

The measured runs make the speedup undeniable and the timing invariance undeniable, which raises the only real engineering question this topic poses: how big should the quantum be? The answer is "as large as your accuracy budget allows, and no larger," and making that concrete means naming exactly what a large quantum degrades — because it is not the timing numbers.

What a large quantum does not break: the total modeled time, any single initiator's own access timeline, the final memory state of independent initiators, and the latency each access annotates. All of these are quantum-invariant, as the measured runs show. If your model is a single initiator, or several non-interacting initiators each on private memory, you can set the quantum enormous and lose nothing but the granularity of your logs.

What a large quantum does break, in order of how often it bites:

  • Interrupt and event observation latency. Suppose a peripheral raises an interrupt at modeled time 25 ns, and a decoupled CPU initiator is running 100 ns ahead on its wristwatch. The CPU will not observe the interrupt until it next syncs — potentially up to a full quantum late. The interrupt is not lost (the event is still there when the CPU finally reconciles), but the measured interrupt latency — the gap between "interrupt asserted" and "CPU started its handler" — is now inflated by up to a quantum. If you are characterizing interrupt response, a 1 us quantum makes every interrupt look up to 1 us slow. The fix is to keep the quantum below the interrupt latency you need to resolve, or to force a sync when servicing time-critical events.
  • Ordering between interacting initiators. As the two-initiator demo showed, a coarse quantum lets one initiator run its whole batch before another runs at all. If they share state, the one that runs "first" in scheduling order sees stale data from the other, because the other's modeled-earlier writes have not been executed yet. This is the genuine correctness hazard — not just inaccuracy but wrong results — and it is why decoupled initiators must sync before touching state another decoupled initiator is concurrently producing. Shared-memory communication, lock acquisition, and producer/consumer FIFOs between two decoupled cores are the danger zones.
  • Spin-loop and polling artifacts. A CPU polling a status register in a tight loop, fully decoupled, will spin for a whole quantum of modeled time without ever giving the peripheral a chance to update the register (the peripheral's update happens in another process that only runs when the CPU syncs). The decoupled spin can appear to hang or to take a quantum longer than it should. Polling loops are a classic place to insert an explicit qk.sync() per iteration to let the rest of the model advance.

The practical rule of thumb, then, is a layered one. Start with the quantum your accuracy requirement implies: if you must resolve interrupt latency to within 100 ns, your quantum cannot exceed ~100 ns. If you are booting an OS and only care about throughput and final state, megahertz-scale quanta (microseconds of modeled time) are fine and give the biggest speedup. Then force syncs at the specific points where the coarse quantum would be wrong: before an access to state shared with another decoupled initiator, when servicing a time-critical interrupt, and inside polling loops. The combination — a large quantum for the common case, targeted syncs for the rare cross-process interactions — gives you most of the speed of a huge quantum with the accuracy of a small one exactly where it matters. The mistake to avoid is treating the quantum as a single global accuracy dial; it is a default that you override locally where correctness demands.

Warning A bigger quantum is not unconditionally faster, and the assumption that it is leads to over-sized quanta that cost more than they save. Past the point where the quantum already spans your typical run-ahead, increasing it further removes no more syncs (you have already eliminated them) but increases the cross-process observation skew, forcing you to add more corrective sync() calls to stay correct — and those forced syncs are context switches you put back. The fast quantum is the largest one your accuracy budget tolerates without needing a forest of corrective syncs, not the largest one representable.

Advanced: Edge Cases & LRM Corners

Temporal decoupling attacks one of the two big LT costs — the context switches. The other big cost is the per-access function-call machinery itself: even with a huge quantum, every memory access still flows through a virtual b_transport call, decodes a generic payload, copies bytes, and returns. For a CPU model executing instructions out of a code memory, that per-access overhead, repeated billions of times, becomes the bottleneck once the syncs are gone. The Direct Memory Interface (DMI) is TLM-2.0's answer: for plain memory, let the initiator obtain a raw pointer to the target's storage once and then read and write it directly, skipping b_transport entirely. This is the Advanced corner of this post not because the API is hard — it is small — but because the rule that makes it safe is the kind of thing that bites senior engineers who reach for the speed without respecting the constraint.

The DMI protocol

DMI is a negotiation on top of the normal socket. The initiator asks; the target may grant or refuse:

  1. The initiator requests DMI by calling socket->get_direct_mem_ptr(trans, dmi), passing a payload (whose address and command indicate the region and direction it wants) and an empty tlm::tlm_dmi descriptor. The call returns bool.
  2. The target decides. If the addressed region is plain, side-effect-free storage, the target fills in the tlm_dmi descriptor — set_dmi_ptr(raw_pointer), set_start_address(lo) / set_end_address(hi) (the contiguous region the pointer covers), set_read_latency(t) / set_write_latency(t) (the per-access latency the initiator should still annotate even though it bypasses b_transport), and allow_read() / allow_write() / allow_read_write() (which directions are granted) — and returns true. If the region is not DMI-able (a peripheral with side effects, or simply a target that does not implement DMI), it returns false.
  3. The initiator caches and uses the pointer. On a grant, the initiator stores the descriptor. For any subsequent access whose address falls in [start, end] and whose direction is allowed, it computes dmi_ptr + (addr - start) and reads or writes the raw bytes directly — no payload, no b_transport, no virtual call — while still annotating the DMI latency onto its local time offset to keep timing honest. On a refusal, or for any address outside the granted region, it falls back to the normal b_transport path.
  4. The target invalidates when the grant stops being valid by calling invalidate_direct_mem_ptr(start, end) on the backward path (for example, when memory is remapped, when another master writes a region the initiator had cached, or when a model reconfigures). The initiator must drop any cached descriptor overlapping that range and revert to b_transport until it re-requests and is re-granted.

The following program makes the speed difference concrete. A 4 KB plain-memory target implements both b_transport and get_direct_mem_ptr. The initiator does two million writes through b_transport, then requests DMI once and does two million writes through the cached raw pointer, timing both phases with a wall clock.

// file: qk_dmi.cpp
// Build: g++ -std=c++17 -DSC_ALLOW_DEPRECATED_IEEE_API \
//          -I$SYSTEMC_HOME/include qk_dmi.cpp \
//          -L$SYSTEMC_HOME/lib -lsystemc -o qk_dmi
// Run:   DYLD_LIBRARY_PATH=$SYSTEMC_HOME/lib ./qk_dmi

#include <systemc.h>
#include <tlm.h>
#include <tlm_utils/simple_initiator_socket.h>
#include <tlm_utils/simple_target_socket.h>
#include <iostream>
#include <chrono>
#include <cstring>

static unsigned long g_b_transport_calls = 0;
static unsigned long g_dmi_grants        = 0;

SC_MODULE(Memory) {
  tlm_utils::simple_target_socket<Memory> socket;
  static const unsigned int SIZE = 4096;
  unsigned char mem[SIZE];
  SC_CTOR(Memory) : socket("socket") {
    for (unsigned i = 0; i < SIZE; i++) mem[i] = 0;
    socket.register_b_transport(this, &Memory::b_transport);
    socket.register_get_direct_mem_ptr(this, &Memory::get_dmi);
  }
  // Slow path: a virtual call + payload decode per access.
  void b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    g_b_transport_calls++;
    sc_dt::uint64  addr = trans.get_address();
    unsigned int   len  = trans.get_data_length();
    unsigned char* ptr  = trans.get_data_ptr();
    if (trans.get_command() == tlm::TLM_WRITE_COMMAND)
      for (unsigned i = 0; i < len; i++) mem[addr + i] = ptr[i];
    else
      for (unsigned i = 0; i < len; i++) ptr[i] = mem[addr + i];
    delay += sc_time(10, SC_NS);
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
  }
  // DMI request: grant a raw pointer to the whole plain-memory region.
  // LEGAL ONLY because this target has NO side effects -- it is just storage.
  bool get_dmi(tlm::tlm_generic_payload& /*trans*/, tlm::tlm_dmi& dmi) {
    g_dmi_grants++;
    dmi.allow_read_write();
    dmi.set_dmi_ptr(mem);
    dmi.set_start_address(0);
    dmi.set_end_address(SIZE - 1);
    dmi.set_read_latency(sc_time(10, SC_NS));
    dmi.set_write_latency(sc_time(10, SC_NS));
    return true;                           // DMI granted
  }
};

SC_MODULE(Cpu) {
  tlm_utils::simple_initiator_socket<Cpu> socket;
  bool          dmi_valid = false;
  tlm::tlm_dmi  dmi;
  SC_CTOR(Cpu) : socket("socket") { SC_THREAD(run); }

  void run() {
    const int N = 2000000;
    unsigned int data = 0xABCD;

    // Phase 1: pure b_transport -- one virtual call per access.
    auto t0 = std::chrono::steady_clock::now();
    {
      tlm::tlm_generic_payload trans;
      sc_time delay = SC_ZERO_TIME;
      for (int n = 0; n < N; n++) {
        trans.set_command(tlm::TLM_WRITE_COMMAND);
        trans.set_address((n * 4) & 0xFFC);
        trans.set_data_ptr(reinterpret_cast<unsigned char*>(&data));
        trans.set_data_length(4);
        trans.set_streaming_width(4);
        trans.set_byte_enable_ptr(nullptr);
        trans.set_dmi_allowed(false);
        trans.set_response_status(tlm::TLM_INCOMPLETE_RESPONSE);
        socket->b_transport(trans, delay);
      }
    }
    auto t1 = std::chrono::steady_clock::now();

    // Phase 2: request DMI once, then hit the raw pointer directly.
    {
      tlm::tlm_generic_payload trans;
      trans.set_command(tlm::TLM_READ_COMMAND);
      trans.set_address(0);
      if (socket->get_direct_mem_ptr(trans, dmi) && dmi.is_write_allowed())
        dmi_valid = true;

      for (int n = 0; n < N; n++) {
        sc_dt::uint64 addr = (n * 4) & 0xFFC;
        if (dmi_valid && addr >= dmi.get_start_address()
                      && addr <= dmi.get_end_address()) {
          unsigned char* p = dmi.get_dmi_ptr() + (addr - dmi.get_start_address());
          std::memcpy(p, &data, 4);        // direct: no call, no payload
        }
      }
    }
    auto t2 = std::chrono::steady_clock::now();

    double slow_ms = std::chrono::duration<double, std::milli>(t1 - t0).count();
    double dmi_ms  = std::chrono::duration<double, std::milli>(t2 - t1).count();

    std::cout << "accesses per phase   : " << N << "\n";
    std::cout << "b_transport calls    : " << g_b_transport_calls
              << " (slow phase), DMI grants: " << g_dmi_grants << "\n";
    std::cout << "transport wall time  : " << slow_ms << " ms\n";
    std::cout << "DMI wall time        : " << dmi_ms  << " ms\n";
    std::cout << "DMI speedup          : " << (slow_ms / dmi_ms) << "x\n";
  }
};

int sc_main(int, char*[]) {
  Cpu    cpu("cpu");
  Memory mem("mem");
  cpu.socket.bind(mem.socket);
  sc_start();
  return 0;
}

The wall-clock numbers depend on your machine, but the structure is fixed and reproducible. On the reference build (SystemC 3.0.1, the standard recipe with no -O2), the run reports:

accesses per phase   : 2000000
b_transport calls    : 2000000 (slow phase), DMI grants: 1
transport wall time  : 40.4 ms
DMI wall time        : 8.66 ms
DMI speedup          : 4.67x

The load-bearing numbers are the counts, which are deterministic: two million b_transport calls in the slow phase, exactly one DMI grant for the fast phase. The fast phase does the same two million writes but issues zero b_transport calls — after the single get_direct_mem_ptr negotiation, every access is a memcpy into the raw pointer. The wall-clock ratio (~4–5x here; it lands around 4.6x consistently on this machine, and substantially higher under -O2 where the memcpy path optimizes to almost nothing) will vary with compiler, flags, and hardware, but the direction never does: eliminating the per-access call machinery is a large, real speedup on top of whatever temporal decoupling already bought you. In a CPU model fetching instructions, DMI on the code memory is often the difference between a usable and an unusable model.

Why DMI is restricted to side-effect-free memory

Now the rule that makes DMI safe, and the misconception that makes it dangerous. A DMI access never enters b_transport. That is the entire source of the speedup — and the entire source of the constraint. Whatever behavior a target codes inside its b_transport body is skipped on a DMI access. For plain memory, the b_transport body does nothing but copy bytes, so skipping it and copying the bytes directly is semantically identical — the raw memcpy and the b_transport produce the same memory state. That is why plain memory is DMI-able.

A peripheral is different. Consider a UART status register that clears its "data ready" bit on read, or a FIFO that pops an element on read, or a control register whose write arms a timer or fires a doorbell interrupt. For all of these, the meaningful behavior lives in the b_transport body — the read side effect, the pop, the timer arm. If such a target granted DMI, the initiator's direct pointer reads and writes would skip every one of those side effects: the status bit would never clear, the FIFO would never pop, the timer would never arm. The model would be silently, grossly wrong, and nothing would flag it, because the raw memory access succeeds — it just does the wrong thing. This is exactly why DMI is restricted to side-effect-free memory: the optimization's mechanism is "skip the side-effect code," so it is only legal where there is no side-effect code to skip. A correct peripheral model refuses DMI — its get_direct_mem_ptr returns false — forcing every access onto the b_transport path where its side effects actually execute. The discipline is one line (return false;) and forgetting it is one of the nastiest bugs in TLM modeling, because it manifests as a peripheral that "mostly works" until the access pattern happens to hit the DMI fast path.

Warning Never grant DMI from a target whose reads or writes have side effects. The DMI fast path bypasses b_transport by design, so any clear-on-read, pop-on-read, write-to-arm, or interrupt-on-access behavior is silently skipped. The rule is mechanical: if your b_transport body does anything other than copy bytes to and from plain storage, return false from get_direct_mem_ptr. DMI is a memory optimization, not a general transport accelerator.

Corner: invalidation and the validity window

The other DMI subtlety is invalidation. A granted pointer is valid only until the target says otherwise. A model that remaps memory, a cache that flushes, or another master that writes a region one initiator had cached must call invalidate_direct_mem_ptr(start, end) on the affected range, and every initiator that cached an overlapping pointer must discard it and fall back to b_transport until it re-requests. The misconception "once I have the pointer it is good forever" produces use-after-remap bugs: the initiator keeps writing through a stale pointer into storage that no longer means what it did. In the simple single-memory example above, no invalidation ever happens (the memory is static and singly-owned), but any model with reconfigurable memory maps or multiple masters of the same storage must implement the invalidation half of the protocol, and an initiator must always be prepared for its cached DMI to be revoked between accesses.

Version differences

tlm::tlm_global_quantum, tlm_utils::tlm_quantumkeeper, and the DMI API (tlm_dmi, get_direct_mem_ptr, invalidate_direct_mem_ptr) are identical across SystemC 2.3.0, 2.3.1, 2.3.3, 2.3.4, and 3.0.x, bundled with the kernel since 2.3.0. The keeper method set used here (set_global_quantum, reset, get_local_time, set, need_sync, sync) and the DMI descriptor setters (set_dmi_ptr, set_start_address/set_end_address, set_read_latency/set_write_latency, allow_read_write) are unchanged across these releases. Every example here compiles and runs identically on any 2.3.x or 3.0.x build; the SC_HAS_PROCESS form used in the manual examples emits a deprecation note under SystemC 3.0.x's IEEE 1666-2023 tracking but is functionally unchanged.

Hands-on exercise

Build a quantum-sweep instrument that measures and tabulates the speedup of temporal decoupling, so the relationship between quantum, sync count, and modeled time stops being something you read about and becomes something you produced.

Start from the manual qk_manual.cpp skeleton (one initiator, a 10 ns/access memory, a local-offset accumulate-and-sync loop), and turn it into a parameter sweep:

  1. Parameterize the workload and the quantum. Make the access count N and the quantum both settable (command-line arguments or a loop in sc_main). Keep the per-access latency fixed at 10 ns so the expected final time is always N × 10 ns — that gives you a known-correct value to check the measured one against.
  2. Instrument three quantities per run: the b_transport call count (modeled work — should be exactly N every time), the wait()/sync count (the context switches — the thing the quantum controls), and the final sc_time_stamp() (the modeled time — should be exactly N × 10 ns every time). Print them as one row per quantum.
  3. Sweep the quantum across at least five values spanning from "one access per quantum" to "the whole workload in one quantum" — for N = 1000 accesses at 10 ns, sweep quanta of 10 ns, 50 ns, 100 ns, 500 ns, and 10 us. Collect the rows into a table.
  4. Add a wall-clock column. Time each sweep run with std::chrono around sc_start() and add the measured wall-clock milliseconds to each row, so the table shows quantum, syncs, modeled time, and real time side by side.

Then read your own table and confirm three things in writing: (a) the modeled time column is constant across every quantum (the timing math is quantum-invariant); (b) the sync count falls as ceil(N × 10ns / quantum) (the context switches drop with the quantum); and (c) the wall-clock time falls with the sync count up to the point where syncs are nearly eliminated, then flattens (bigger quantum stops helping once the syncs are gone). That flattening is the "bigger is not always faster" result made empirical on your machine.

When that works, extend it in two directions. First, rewrite the swept initiator with tlm_quantumkeeper and confirm every row of your table is reproduced exactly — same syncs, same modeled time — proving the keeper is just packaging. Second, add a second decoupled initiator on a second memory port and a shared counter that both increment via a (DMI-refusing, side-effecting) target; sweep the quantum again and watch the shared counter's final value go wrong once the quantum grows large enough that the two initiators stop interleaving — then add a forced qk.sync() before each shared-counter access and watch correctness return at the cost of some of the speedup. That second extension is the accuracy/speed trade-off, measured on your own bench.

Hints

  • The expected final time is sc_time(N * 10, SC_NS). Print it alongside the measured sc_time_stamp() and assert they match — a mismatch means a missing tail sync() or a dropped qk.set(delay).
  • The expected sync count for the manual loop is ceil(N * 10ns / quantum) when each access banks exactly 10 ns; for N = 1000 at quantum 100 ns that is 100 syncs. Use this to check your instrumentation before trusting the wall-clock numbers.
  • For the wall-clock column, wrap only sc_start() in the std::chrono::steady_clock timing — not the elaboration — and run each quantum a few times to see the variance before reporting a number.
  • For the keeper rewrite, remember the four-line idiom: sc_time delay = qk.get_local_time(); before the call, qk.set(delay); after, if (qk.need_sync()) qk.sync(); to reconcile, and one trailing qk.sync(); after the loop. Drop any of them and your modeled time comes out short or your sync count comes out wrong.
  • For the shared-counter extension, the side-effecting target must return false; from get_direct_mem_ptr (no DMI) and its b_transport must do the read-modify-write; the wrong final value comes from the coarse quantum letting one initiator run its whole batch on a stale counter value before the other runs at all.
  • To plot the table without a plotting library, just print it as columns (quantum syncs modeled_time wall_ms) and paste it into a spreadsheet; the shape (syncs and wall-time falling together, then wall-time flattening) is the deliverable, not a pretty graph.

No solution is provided. The understanding lives in producing the table yourself and reading the invariance of the modeled-time column against the collapse of the sync column.

Common mistakes

  • Believing a larger quantum changes the simulated-time result. It does not. Every access annotates the same latency regardless of quantum, and the banked offset is paid in full at each sync, so the final modeled time is quantum-invariant — the measured sweeps show identical 1 us at every quantum. The quantum changes the number of context switches and the granularity of cross-process observation, never the timing arithmetic. Fix the mental model: the quantum is a performance-and-observation knob, not an accuracy-of-timing knob.
  • Forgetting qk.set(delay) after b_transport (or the trailing qk.sync()). If you do not absorb the post-call delay back into the keeper, it never learns the latency the target added, need_sync() never fires, and the whole simulation runs on a frozen kernel clock. If you skip the final sync(), the last unspent offset is never paid and your final timestamp is short. Fix: thread get_local_time() in and set(delay) out on every access, and always settle the tail with one sync().
  • Letting the target call wait() to model latency under decoupling. A target that waits internally forces a context switch on every access, collapsing the quantum to zero and erasing the entire benefit. Fix: targets annotate (delay += latency) and return; initiators (or their keepers) decide when to pay. This is the same "targets annotate, initiators pay" rule from the generic-payload post, and decoupling depends on it absolutely.
  • Setting the quantum larger than your accuracy budget allows. A coarse quantum inflates observed interrupt latency by up to a quantum and lets interacting initiators read each other's stale state. Fix: size the quantum to the finest timing you must resolve (interrupt latency, cross-core ordering), and override it with targeted sync() calls before shared-state accesses, time-critical interrupts, and inside polling loops — large default quantum, local syncs where correctness needs them.
  • Assuming a bigger quantum is always faster. Past the point where the quantum already spans your typical run-ahead, increasing it removes no further syncs but increases cross-process skew, forcing corrective syncs that claw the savings back. Fix: pick the largest quantum your accuracy budget tolerates without needing a forest of corrective syncs, and measure — the wall-clock-versus-quantum curve flattens, and pushing past the flat is pure risk for no speed.
  • Granting DMI from a side-effecting target. DMI bypasses b_transport, so any clear-on-read, pop-on-read, write-to-arm, or interrupt-on-access behavior is silently skipped, producing a model that "mostly works" until an access hits the DMI fast path. Fix: a target whose b_transport does anything beyond copying plain bytes must return false; from get_direct_mem_ptr. DMI is legal only for side-effect-free storage.
  • Caching a DMI pointer forever and ignoring invalidation. A granted pointer is valid only until the target calls invalidate_direct_mem_ptr. A model with reconfigurable maps or multiple masters of the same storage that never invalidates — or an initiator that never honors invalidation — produces use-after-remap corruption. Fix: implement the invalidation half of the protocol on the target, and on the initiator always be prepared for a cached DMI to be revoked between accesses, falling back to b_transport.

Recap

After working through this post you can now:

  • Explain why loosely-timed simulation cost is dominated by context switches (wait() calls returning control to the kernel scheduler), not by the modeled work, and why removing redundant synchronizations is the highest-leverage LT optimization.
  • State the core idea of temporal decoupling — an initiator runs ahead of the kernel clock on a private local time offset, banking annotated latency, and synchronizes with one wait() only when the offset reaches the quantum — and reinterpret the b_transport delay argument as that local offset.
  • Build temporal decoupling by hand (accumulate delay, compare against the quantum, wait() and reset on crossing) and measure the result: identical final modeled time at every quantum, with the context-switch count collapsing in proportion to the quantum (100 → 10 → 1 in the worked sweep).
  • Articulate the load-bearing invariant — the quantum changes observation granularity and ordering, never the timing math — and prove it from the measured runs.
  • Explain and demonstrate the cross-initiator observation reordering a coarse quantum produces, distinguish it (harmless for independent initiators) from the genuine causality hazard (wrong results for interacting initiators reading each other's stale state), and know to sync before shared-state accesses.
  • Use tlm_utils::tlm_quantumkeeper with the canonical idiom — set_global_quantum, reset, get_local_time (in), set (out), need_sync, sync, trailing sync — and recognize it as the reviewable packaging of the manual arithmetic, reproducing the same measured numbers.
  • Choose a quantum per use case: large for throughput/boot workloads, bounded by the interrupt latency or cross-core ordering you must resolve, with targeted forced syncs where a coarse quantum would be incorrect — and know that bigger is not unconditionally faster.
  • Use the Direct Memory Interface (DMI): request with get_direct_mem_ptr, fill a tlm_dmi descriptor, cache and use the raw pointer for a large additional speedup, and honor invalidate_direct_mem_ptr — and explain the hard rule that DMI is legal only for side-effect-free memory, because the optimization works by skipping the b_transport body where side effects would live.

Further reading

Standards

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §11.2 (temporal decoupling and the global quantum), §11.3 (the quantum keeper and the tlm_quantumkeeper utility), and §11.4 (the Direct Memory Interface: tlm_dmi, get_direct_mem_ptr, invalidate_direct_mem_ptr, and the grant/invalidate protocol).

Vendor and consortium documents

  • Aynsley (Doulos), OSCI TLM-2.0 Language Reference Manual (JA32) — the canonical narrative treatment of temporal decoupling, the quantum-keeper idiom, and the DMI grant/invalidate handshake.
  • Doulos TLM-2.0 Tutorial — worked loosely-timed initiators using tlm_quantumkeeper and DMI, and the rationale for keeping side-effecting peripherals off the DMI path.
  • Accellera Systems Initiative, SystemC 3.0.0 distribution, include/tlm_utils/tlm_quantumkeeper.h and include/tlm_core/tlm_2/tlm_2_interfaces/tlm_dmi.h — authoritative source for the keeper method set and the DMI descriptor API.

Textbooks

  • Grötker, Liao, Martin, and Swan, System Design with SystemC — the loosely-timed coding style and the simulation-speed argument that motivates temporal decoupling.

Next in this section

→ Part 6: Bus Fabric — Routing & Decoding — how one initiator reaches many targets: an interconnect module that routes a transaction from an initiator to the right target based on a memory map, performs address decode and localization (subtracting the target’s base so each target sees offsets from zero), and returns TLM_ADDRESS_ERROR_RESPONSE for accesses that fall outside every mapped region. Read it here: 21. SystemC Tutorial — Bus Fabric: Routing & Decoding.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment