3. SystemC Tutorial - Delta Cycles & Event-Driven Semantics

Why this matters

Rewritten 2026-05-23 with deeper first-principles material.

Every multi-process SystemC model you will ever write — a virtual platform of an ARM Cortex-M class SoC, a SiFive RISC-V pre-silicon model used at Intel for early firmware bring-up, a cycle-approximate model of a SoC interconnect being tuned for an automotive ADAS chip — is correct only because the kernel's delta-cycle machinery preserves a specific property: writes made by one process during an evaluation become visible to all other processes simultaneously, exactly one delta later, with no observable interleaving. Get the delta-cycle mental model right and your models match real hardware. Get it wrong and your model either deadlocks on the first multi-process handshake or, far worse, runs and reports nonsense numbers that boot Linux and pass smoke tests for weeks before someone notices the cycle count is off by a factor of two.

Engineers at virtual-platform houses — Wind River Simics, Synopsys Virtualizer, ARM Fast Models, Cadence Helium — debug delta-cycle bugs every release. The bugs look like "my interrupt model fires one cycle late when the timer overflow lands on the same delta as the GIC priority update." That sentence is incomprehensible without a working mental model of delta cycles, and unfixable without understanding why notify() and notify(SC_ZERO_TIME) are not the same thing. RTL engineers come into SystemC assuming <= semantics work the same way — they do, but the SystemC kernel makes the mechanism explicit in a way SystemVerilog's NBA scheduler hides. Pure software engineers come in assuming signal.write(x); int y = signal.read(); returns x in y — it does not, and the moment they have a four-stage producer-consumer chain they have lost a day chasing the wrong value through their code.

This post fixes that. By the end you will be able to predict the exact number of delta cycles a small SystemC program takes to converge, before running it, and to read the simulator's trace output and explain every line. That is the bar.

Prerequisites

  • Part 1 — Modules, Ports & Signals. You need to be comfortable declaring SC_MODULE, binding sc_signal channels to sc_in / sc_out ports, and registering a process with SC_METHOD or SC_THREAD inside SC_CTOR.
  • Part 2 — Simulation Time & Clocks. You need to understand sc_time, SC_NS, the difference between sc_start() and sc_start(t), and how sc_time_stamp() reports the current simulated instant.
  • SystemC 2.3.x installed. The installer posts (referred to as P1 and P2 in this series — the original installation tutorials for Linux and macOS) cover environment setup, SYSTEMC_HOME, and the g++ command-line invocation. All examples in this post compile with a C++17 compiler (g++ 9 or newer, or clang++ 10 or newer).

Three concepts from earlier parts get heavy use here: sc_signal<T> (the buffered channel), SC_METHOD (the non-blocking process kind), and SC_THREAD (the suspendable process kind, used for the wait()-based examples). If any of those is hazy, re-read the relevant section of Parts 1 and 2 before continuing.

Mental model (first principles)

To predict what a SystemC program will do, you need to model the kernel itself as a state machine. Most tutorials skip this step and dive into syntax. The result is that readers can copy and modify examples but cannot reason about novel designs. The mental model that follows is a deliberately small, faithful abstraction of sc_simcontext::crunch() in the Accellera reference simulator. If you internalize it, you can read any SystemC program and predict its output.

The kernel runs a 3-phase loop, per delta cycle, at the current simulation time:

  1. Evaluate phase. The kernel takes the set of runnable processes — processes that were made runnable in the previous delta's post-update notification phase, or by the previous time-advance's timed-notification step — and runs each one to completion (for SC_METHOD) or until its next wait() (for SC_THREAD). The order in which the kernel picks processes off the runnable set is implementation-defined: the standard does not specify it. During evaluation, processes read signals (which return the current, not pending, value) and may call .write() to schedule a new value, fire events with .notify(), or call wait() (threads only) to suspend. Each .write() on an sc_signal<T> is a side effect of evaluation: it invokes request_update(), which stashes the new value in a "next" slot and registers the signal as having an update pending; the signal's current value does not change yet. A subsequent .read() in the same evaluate phase — even from inside the same process — returns the old current value. There is no exception to this rule.
  1. Update phase. When no more processes are runnable, the kernel walks the set of signals with pending updates and copies each "next" value into the "current" slot. For each signal whose new current value differs from the old current value, the kernel fires the signal's internal value_changed_event(). This event-firing is what determines who runs next.
  1. Post-update notification phase. Any process statically sensitive to a signal whose value changed in phase 2, or dynamically waiting on an event that fired in phase 2 or via an earlier notify(SC_ZERO_TIME), becomes runnable for the next delta cycle. The delta counter advances by one. Simulation time stays at its current value. The kernel loops back to phase 1.

When phase 1 begins with an empty runnable set, the system is stable at the current simulation time. The kernel then asks "is there a timed notification scheduled?" If yes, it advances sc_time_stamp() to that notification's time, fires the event (which may make processes runnable), and re-enters the delta loop. If no, simulation ends — either because sc_stop() was called or because nothing more can happen.

Counting convention used throughout this series. A "delta cycle" is one evaluate-update-notification pair after sc_start begins. Writes posted from sc_main before sc_start are pending writes that commit in an initialization update (sometimes called "delta 0") at the very start of sc_start, before any user delta cycles run. When we count "N delta cycles" for a program in this post and the rest of the series, we mean N evaluate-update pairs after that initialization update has already happened.

              ┌───────────────────────────────┐
       ┌────► │  Runnable set is empty?       │
       │      └───────────────┬───────────────┘
       │                      │
       │       no ◄───────────┴──────────► yes
       │       │                            │
       │       ▼                            ▼
       │  ┌──────────────────────────┐    ┌─────────────────────────────┐
       │  │ 1. EVALUATE              │    │  Any timed notification     │
       │  │   pick runnable process  │    │  scheduled for a future t?  │
       │  │   (impl-defined order)   │    └───────────────┬─────────────┘
       │  │   run to completion or   │                    │
       │  │   next wait()            │      no ◄──────────┴────────► yes
       │  │   .write() => buffer     │      │                         │
       │  │   .notify() => may make  │      ▼                         ▼
       │  │   runnable               │    STOP                ┌────────────────────┐
       │  └────────────┬─────────────┘                        │  TIME ADVANCE      │
       │               │ all runnable done                    │  sc_time_stamp() = t │
       │               ▼                                      │  fire timed events │
       │  ┌──────────────────────────┐                        │  (may make runnable)│
       │  │ 2. UPDATE                │                        └──────────┬──────────┘
       │  │   pending signals:       │                                   │
       │  │     current <- next      │                                   │
       │  │     if changed: fire     │                                   │
       │  │     value_changed_event  │                                   │
       │  └────────────┬─────────────┘                                   │
       │               ▼                                                 │
       │  ┌──────────────────────────┐                                   │
       │  │ 3. POST-UPDATE NOTIFY    │                                   │
       │  │   sensitive processes    │                                   │
       │  │   => runnable set        │                                   │
       │  │   delta_count++          │                                   │
       │  └────────────┬─────────────┘                                   │
       │               │                                                 │
       └───────────────┴─────────────────────────────────────────────────┘
                       (back to top — next delta or next time step)

Two consequences fall out of this loop and matter for everything in the rest of this post:

  • Delta cycles take zero simulation time. No matter how many delta cycles the kernel runs at simulation time T, sc_time_stamp() returns T the whole way through. Time advances only after the runnable set is empty.
  • The order in which two writes to the same signal in the same evaluate phase resolve is implementation-defined. The standard says nothing about which process the kernel will run first. The "last writer wins" rule is real, but which writer is last is not portable across simulators.

Why does this loop exist? Because it preserves the illusion of concurrent execution on a sequential machine. Real hardware runs every flip-flop in parallel on the same clock edge. A simulator cannot, but it can fake it by running every clocked process at the same simulation timestamp and then resolving all the inter-process visibility in a single atomic update phase. The delta cycle is the cost of that illusion. Once you see it that way, every subsequent design rule — "don't read a signal you just wrote and expect the new value", "register your handshake events with notify() not notify(SC_ZERO_TIME) unless you mean to delay" — becomes obvious rather than memorized.

Beginner: First Principles

The simplest possible SystemC program that exhibits a delta cycle has exactly two processes: one that writes a signal, and one that reads it. We will build that program now, run it in our heads, predict the output line by line, and then verify by running it on a real kernel.

// file: delta_intro.cpp
// Build:  g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//             delta_intro.cpp -o delta_intro -lsystemc
// Run:    LD_LIBRARY_PATH=$SYSTEMC_HOME/lib-linux64 ./delta_intro

#include <systemc.h>
#include <iostream>

SC_MODULE(writer) {
  sc_out<int> out;

  void drive() {
    std::cout << "[writer] before write, out.read() = "
              << out.read() << "\n";
    out.write(42);
    std::cout << "[writer] just  wrote 42; out.read() = "
              << out.read() << "  <-- still the OLD value\n";
  }

  SC_CTOR(writer) {
    SC_METHOD(drive);
    // No sensitivity list => runs once during initialisation.
  }
};

SC_MODULE(reader) {
  sc_in<int> in;

  void observe() {
    std::cout << "[reader] sensitivity fired; in.read() = "
              << in.read() << "\n";
  }

  SC_CTOR(reader) {
    SC_METHOD(observe);
    sensitive << in;            // fire whenever 'in' changes
    dont_initialize();          // skip the initial activation
  }
};

int sc_main(int /*argc*/, char* /*argv*/[]) {
  sc_signal<int> wire;
  wire.write(0);                 // give the signal a known starting value

  writer w("w");  w.out(wire);
  reader r("r");  r.in(wire);

  std::cout << "[main] starting sc_start(SC_ZERO_TIME)\n";
  sc_start(SC_ZERO_TIME);
  std::cout << "[main] sc_start returned; wire.read() = "
            << wire.read() << "\n";
  return 0;
}

Compile and run, and you will see the following output. Read the expected output before the walkthrough — try to explain each line to yourself first.

[main] starting sc_start(SC_ZERO_TIME)
[writer] before write, out.read() = 0
[writer] just  wrote 42; out.read() = 0  <-- still the OLD value
[reader] sensitivity fired; in.read() = 42
[main] sc_start returned; wire.read() = 42

Walk through it line by line, with the 3-phase kernel loop from the previous section in mind.

[main] starting sc_start(SC_ZERO_TIME). Before sc_start was called, we called wire.write(0) from sc_main. That call does not immediately set wire's current value — it registers a pending update. The current value of wire at that moment is its default-constructed value (also 0 for int). When sc_start begins, the kernel runs an initialization update that commits all pending writes posted from sc_main — so wire's current slot becomes 0 (last write wins, and both the default and our explicit write were 0; if we had written wire.write(99) from sc_main, wire's current value would become 99 here). Then the kernel runs the initialization delta: every process that did not call dont_initialize() becomes runnable. writer::drive qualifies; reader::observe does not.

[writer] before write, out.read() = 0. The kernel enters delta 1's evaluate phase. It picks writer::drive (the only runnable process) and runs it. Inside, out.read() returns the current value of the bound signal — which is 0, the value committed by the initialization update. Crucially, read() here is not reading from out; it is reading from the channel out is bound to. The port forwards.

[writer] just wrote 42; out.read() = 0 <-- still the OLD value. This is the lesson. out.write(42) calls sc_signal::request_update() under the hood. The new value 42 is stashed in the "next" slot. The current slot still holds 0. When the very next line calls out.read(), the kernel returns the current slot — 0. The write has not taken effect because the evaluate phase has not ended.

If you expected out.read() to show 42 after the write, you are feeling the update-phase delay first-hand. That delay is not a bug; it is the only way the simulator can pretend that an arbitrary number of processes ran at the same instant. Burn this in. Every multi-process bug you ever debug in SystemC starts with the question "did this read see the pre-update or post-update value?", and the answer is always "pre-update, until the next delta".

(End of [writer] drive(). The writer was an SC_METHOD, so it returns. The kernel's runnable set is now empty for delta 1's evaluate phase.)

The kernel enters delta 1's update phase. It walks pending signals — there is one, wire, with a pending value of 42. The kernel copies 42 into wire's current slot and fires wire's internal value_changed_event. Any process sensitive to wire (that is, reader::observe) is added to the runnable set for the next delta. Delta counter advances to 2. Simulation time is still 0.

[reader] sensitivity fired; in.read() = 42. Delta 2 evaluate phase begins. reader::observe is the only runnable process. It calls in.read(), which forwards to wire. wire's current value is now 42. The reader prints 42. The reader returns. Runnable set is empty.

Delta 2 update phase: no pending writes. Post-update notification: no new processes runnable. Runnable set still empty. No timed notifications pending. sc_start returns.

[main] sc_start returned; wire.read() = 42. Back in sc_main, the signal's current value is 42.

This six-line program ran two delta cycles after sc_start began (plus the initialization update before delta 1). The writer ran in delta 1; the buffered write committed in delta 1's update phase; the reader ran in delta 2 because the buffered write changed the signal; nothing was left to do; the kernel returned. Total simulated time consumed: zero. Total delta cycles consumed: two. Memorize this trace. Every program in the rest of this post is just a larger version of it.

A common confusion at this point is to look at the output and conclude that sc_signal writes are "asynchronous". They are not. They are precisely scheduled: every write made during an evaluate phase commits during the next update phase, period. There is no race. Two writers writing to the same signal in the same evaluate phase resolve in a well-defined update phase; we will see what "well-defined" means when both writers are independent in the Intermediate section.

Two operational details worth knowing before you continue.

First, the dont_initialize() call on the reader matters. Without it, the reader's SC_METHOD would run once during the initialization delta with in.read() returning 0 (the initial value of wire), producing an extra [reader] sensitivity fired; in.read() = 0 line at the top of the output. That extra line is not wrong — it just confuses the lesson. dont_initialize() is your friend when you want the reader to only react to changes, not to the initial state.

Second, the choice of SC_METHOD rather than SC_THREAD for the writer was deliberate. An SC_METHOD runs to completion on each activation; there is no wait(). That makes the kernel state at any line of drive() trivially deterministic — there is one process running, and it runs through to its return. We will introduce threads with wait() in the Intermediate section, where the delta-cycle interaction gets richer.

That is the smallest interesting SystemC program. One signal, one write, one delta of latency. Everything else is layering.

Intermediate: How It Really Works

The Beginner example was deliberately tiny so the kernel trace fit in your head. Real SystemC designs are not that small. The interesting questions are: what happens when two processes both want to react to the same change, when three writers hit the same signal in the same evaluation, when a chain of dependent processes forms, when threads suspend on wait() mid-delta, and when the simulation engineer needs to know whether a chain of delta cycles is a feature or a performance problem. This section walks through three substantial worked examples and then gives you a decision table you can refer back to when designing.

Worked example 1: ping-pong between two processes

This is the smallest interesting two-process design. Process A is sensitive to signal sig_a and writes to sig_b. Process B is sensitive to sig_b and writes to sig_a. From sc_main we seed one of the signals and watch what happens.

// file: ping_pong.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            ping_pong.cpp -o ping_pong -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(ping_pong) {
  sc_in<int>  a_in;
  sc_out<int> a_out;
  sc_in<int>  b_in;
  sc_out<int> b_out;

  void on_a() {
    int v = a_in.read();
    if (v < 4) {
      std::cout << "  [on_a]  sees a=" << v << "; writing b=" << v + 1 << "\n";
      b_out.write(v + 1);
    } else {
      std::cout << "  [on_a]  sees a=" << v << "; guard false, no write\n";
    }
  }

  void on_b() {
    int v = b_in.read();
    if (v < 4) {
      std::cout << "  [on_b]  sees b=" << v << "; writing a=" << v + 1 << "\n";
      a_out.write(v + 1);
    } else {
      std::cout << "  [on_b]  sees b=" << v << "; guard false, no write\n";
    }
  }

  SC_CTOR(ping_pong) {
    SC_METHOD(on_a); sensitive << a_in; dont_initialize();
    SC_METHOD(on_b); sensitive << b_in; dont_initialize();
  }
};

int sc_main(int /*argc*/, char* /*argv*/[]) {
  sc_signal<int> sig_a, sig_b;
  sig_a.write(0);
  sig_b.write(0);

  // Each ping_pong module has symmetric in/out pairs.
  // We bind a_in and a_out to the same external signal (sig_a):
  // on_a reads via a_in; on_b writes via a_out. Same wire, two directions.
  ping_pong pp("pp");
  pp.a_in(sig_a); pp.a_out(sig_a);
  pp.b_in(sig_b); pp.b_out(sig_b);

  std::cout << "[seed] sig_a <- 1\n";
  sig_a.write(1);
  sc_start(SC_ZERO_TIME);
  std::cout << "[done] sig_a=" << sig_a.read() << " sig_b=" << sig_b.read() << "\n";
  return 0;
}

Predict the output before reading on. The seed sig_a.write(1) kicks the chain; each on_* only writes back when its read value is less than 4. Where does the chain stop, and how many delta cycles after sc_start does it take?

Run it. The output:

[seed] sig_a <- 1
  [on_a]  sees a=1; writing b=2
  [on_b]  sees b=2; writing a=3
  [on_a]  sees a=3; writing b=4
  [on_b]  sees b=4; guard false, no write
[done] sig_a=3 sig_b=4

Count the delta cycles, applying the convention from the Mental Model section: deltas are evaluate-update pairs after the initialization update commits the seed write. The sig_a.write(1) call in sc_main is a pending write. When sc_start begins, the initialization update commits sig_a = 1, fires value_changed_event, and makes on_a runnable for delta 1. (Both processes used dont_initialize(), so nothing else fires in the initialization delta.) Delta 1 evaluate runs on_a: reads a=1, guard true, buffers b = 2. Delta 1 update commits sig_b = 2, makes on_b runnable. Delta 2 evaluate runs on_b: reads b=2, guard true, buffers a = 3. Delta 2 update commits sig_a = 3, makes on_a runnable. Delta 3 evaluate runs on_a again: reads a=3, guard true, buffers b = 4. Delta 3 update commits sig_b = 4, makes on_b runnable. Delta 4 evaluate runs on_b: reads b=4, guard false, prints "no write". Delta 4 update: no pending writes, no events fired, runnable set stays empty. sc_start returns.

Four delta cycles after sc_start began (deltas 1, 2, 3, 4), preceded by the initialization update that committed the seed. Simulation time stayed at 0 the whole way. This is the canonical pattern: a chain of process-to-process dependencies through signals creates a delta cycle per "hop" in the chain. If you have a three-stage combinational pipeline (A → B → C), expect three deltas to settle. A ten-stage chain takes ten. The kernel handles them all inside one sc_start(SC_ZERO_TIME) call — you do not have to step through them.

Two things to notice. First, the guard v < 4 is what makes the chain finite. Without it, this design would loop forever, the runnable set never empties, simulation time never advances, and sc_start never returns. That is called a delta-cycle livelock and it is one of the easier-to-diagnose SystemC bugs because the process simply hangs. Second, the values you see on sig_a and sig_b at the end (3 and 4) are the values after all four user deltas have committed. You cannot observe the intermediate values from outside the kernel — they exist only inside their respective delta's evaluate phase, and the std::cout lines are the only window into them.

Worked example 2: three writers, one signal, same evaluation

Now we deliberately set up the case the standard warns about: multiple processes writing different values to the same signal in the same evaluate phase. The LRM says the result is implementation-defined. The question is: what does that mean in practice, and how can a real design accidentally land in this case?

// file: multi_writer.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            multi_writer.cpp -o multi_writer -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(triple_writer) {
  sc_in<bool> trigger;
  sc_out<int> out;

  void w1() { if (trigger.read()) { std::cout << "  [w1] writes 10\n"; out.write(10); } }
  void w2() { if (trigger.read()) { std::cout << "  [w2] writes 20\n"; out.write(20); } }
  void w3() { if (trigger.read()) { std::cout << "  [w3] writes 30\n"; out.write(30); } }

  SC_CTOR(triple_writer) {
    SC_METHOD(w1); sensitive << trigger; dont_initialize();
    SC_METHOD(w2); sensitive << trigger; dont_initialize();
    SC_METHOD(w3); sensitive << trigger; dont_initialize();
  }
};

int sc_main(int /*argc*/, char* /*argv*/[]) {
  sc_signal<bool> trig;
  sc_signal<int>  data;
  trig.write(false);
  data.write(0);

  triple_writer tw("tw");
  tw.trigger(trig);
  tw.out(data);

  std::cout << "[seed] trig <- true\n";
  trig.write(true);
  sc_start(SC_ZERO_TIME);
  std::cout << "[done] data = " << data.read() << "\n";
  return 0;
}

On a typical Accellera 2.3.4 build the output looks like this:

[seed] trig <- true
  [w1] writes 10
  [w2] writes 20
  [w3] writes 30
[done] data = 30

The "30" at the end is not portable. It happens because this kernel's runnable set is implemented as a FIFO and the methods were registered in the order w1, w2, w3. On a kernel that uses a LIFO, the order might be w3, w2, w1 and the final value would be 10. On a kernel that randomises for stress testing (a debugging feature some commercial simulators offer), every run could be different. The LRM's "implementation-defined" wording exists precisely so kernel implementers can change this without breaking the standard.

The practical takeaway: never design a system where the final value of a signal depends on the order three writers fire in the same evaluation phase. If you need to model multiple drivers (a tri-state bus, a wired-OR), use sc_signal_resolved or sc_signal_rv, which carry a resolution function that handles multiple writes deterministically. The Advanced section walks through this in more detail. For now: if you see std::cout in your design code showing two or three writers all firing on the same event and all targeting the same signal, treat that as a design bug — even if today's output happens to be the value you wanted.

How does a real design accidentally land in this case? The most common path is shared instance reuse: a designer builds a generic register_block and instantiates two of them, then realises both blocks need to expose a status flag to the same upstream sc_signal<bool> error_flag. Both blocks bind to the same signal, both update it during evaluation, and the design "works" on the engineer's machine. The build farm finds it three months later when one stage of CI uses a different kernel build.

Worked example 3: SC_THREAD with wait() and the delta-cycle interaction

So far every process has been an SC_METHOD. Threads change the story because wait() can suspend a process mid-execution, which interacts with delta cycles in ways SC_METHOD cannot. Here is a small clocked producer.

// file: clocked_producer.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            clocked_producer.cpp -o clocked_producer -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(producer) {
  sc_in<bool>  clk;
  sc_out<int>  data;

  void drive() {
    int n = 0;
    while (n < 3) {
      wait();                                  // wait for next clk posedge
      data.write(n);
      std::cout << "[" << sc_time_stamp() << "] producer wrote " << n
                << "; data.read() inside producer = " << data.read() << "\n";
      n++;
    }
    sc_stop();
  }

  SC_CTOR(producer) {
    SC_THREAD(drive);
    sensitive << clk.pos();
  }
};

SC_MODULE(observer) {
  sc_in<int> data;

  void watch() {
    std::cout << "[" << sc_time_stamp() << "] observer saw data = "
              << data.read() << "\n";
  }

  SC_CTOR(observer) {
    SC_METHOD(watch);
    sensitive << data;
    dont_initialize();
  }
};

int sc_main(int /*argc*/, char* /*argv*/[]) {
  sc_clock     clk("clk", 10, SC_NS);
  sc_signal<int> data;
  data.write(-1);

  producer p("p"); p.clk(clk); p.data(data);
  observer o("o"); o.data(data);

  sc_start();   // run until sc_stop() or runnable set drained with no future events
  return 0;
}

Note the bare sc_start() call. sc_start() with no argument runs until sc_stop() is called by some process, or until the runnable set is empty and no future timed events exist. Use sc_start(t) when you want a bounded simulation window — that is what we used in the Beginner example to step exactly through the delta chain at time 0.

Output:

[10 ns] producer wrote 0; data.read() inside producer = -1
[10 ns] observer saw data = 0
[20 ns] producer wrote 1; data.read() inside producer = 0
[20 ns] observer saw data = 1
[30 ns] producer wrote 2; data.read() inside producer = 1
[30 ns] observer saw data = 2

Walk through one clock edge. At t = 10ns, the clock's positive edge fires. The producer thread wakes from its wait(), executes data.write(0), then prints. Inside the producer's own evaluation, data.read() returns -1 — the initialised value, because the write to 0 has not committed yet. The producer loops to wait() and suspends. The runnable set is empty. The update phase runs: data = 0, value-changed-event fires, observer becomes runnable for the next delta. Delta cycle: observer runs at t = 10ns (still 10ns, no time advance), reads data = 0, prints. Runnable set empty. No timed events at 10ns. Kernel advances time to the next clock edge at 20ns.

This is the standard producer-observer pattern, and it is the prototype for every signal-based testbench you will ever build. The key point: even though both the producer and the observer print "10 ns" in their time stamps, they are not running simultaneously. The producer ran in delta N at time 10; the observer ran in delta N+1 at time 10. Their output lines are interleaved by simulation order, not by any wall-clock concurrency. If you watch the writes, the producer always sees the previous cycle's value when it reads back its own output. That is the update-phase delay biting again, just in a clocked context.

A small variation: replace wait() with wait(SC_ZERO_TIME). This is not the same as omitting the wait — it explicitly suspends the thread for one delta cycle. It is occasionally useful for handshake protocols where you need to release the simulator long enough for another process to react. We will see it in the Advanced section.

Decision table: which construct for which intent

What you want                          Use this
─────────────────────────────────────  ────────────────────────────────────
Notify a process right now, same       sc_event::notify()
delta, no waiting                      (immediate notification)

Notify a process at the start of       sc_event::notify(SC_ZERO_TIME)
the next delta — give the current      (delta notification)
evaluate phase a chance to finish

Notify a process at simulation time    sc_event::notify(t, SC_NS)
t from now                             (timed notification)

Carry a data value AND signal a        sc_signal<T>
change                                 (value-changed-event built in)

Carry a data value, but multiple       sc_signal_resolved or
processes will drive it                sc_signal_rv

Suspend a thread until a specific      wait(some_event)
event fires

Suspend a thread for one delta         wait(SC_ZERO_TIME)
without an explicit event

Suspend a thread until time t          wait(t, SC_NS)
elapses

Suspend a thread until the next        wait()       (in an SC_THREAD whose
clock edge declared in its static      sensitivity is sensitive << clk.pos())
sensitivity

React to a value change without        SC_METHOD + sensitive << signal
suspension semantics                   + dont_initialize()

Performance note: delta chains cost real CPU

Every delta cycle is bookkeeping: a pass through the pending-write list, a pass through the value-changed events, a pass through the runnable-process set. In a tight model where most clock edges trigger five or ten deltas of combinational propagation, the bookkeeping cost is significant. On a 64-core SoC virtual platform that simulates a billion target cycles in a CI pipeline, halving the average delta count per clock edge halves the simulation wall-clock time.

The two most common sources of unnecessary deltas in a real design are (1) chains of combinational SC_METHOD processes that could be folded into one (each method-to-method hop is a delta); and (2) signals that get written even when their value has not changed, because the kernel still has to walk them in the update phase to discover they did not change. If you suspect delta-cycle bloat, the simplest diagnostic is to add std::cout prints at the start of each SC_METHOD and watch how many fire per clock edge. Five per edge is normal. Fifty per edge in a small model means you have a propagation chain that wants to be one expression.

This concludes the Intermediate section. You now have the kernel mental model, a clean two-process trace, the failure mode of multi-writer collisions, and an idiomatic clocked producer-observer. The Advanced section drills into the LRM corners.

Advanced: Edge Cases & LRM Corners

The Beginner and Intermediate sections cover what 95% of SystemC code does. This section is the other 5% — the LRM corners that tutorials skip, the bugs that take a senior engineer half a day to track down, and the constructs you should know exist even if you do not reach for them weekly.

Corner 1: immediate vs delta vs timed notification, demonstrated

The single most important LRM corner in this post. The standard (§5.10) defines three forms of sc_event::notify. The forms are not interchangeable, and confusing them is the most common cause of "the testbench passes but the model is one cycle off" bugs in production.

// file: notify_forms.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            notify_forms.cpp -o notify_forms -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(showcase) {
  sc_event ev_imm;
  sc_event ev_delta;
  sc_event ev_timed;

  void driver() {
    wait(10, SC_NS);                       // let consumers register on events
    std::cout << "[t=" << sc_time_stamp() << "] driver: firing all three\n";
    ev_imm.notify();                       // immediate
    ev_delta.notify(SC_ZERO_TIME);         // delta
    ev_timed.notify(5, SC_NS);             // timed
    wait(100, SC_NS);
    sc_stop();
  }

  void consume_imm() {
    while (true) {
      wait(ev_imm);
      std::cout << "[t=" << sc_time_stamp() << "] consume_imm fired\n";
    }
  }

  void consume_delta() {
    while (true) {
      wait(ev_delta);
      std::cout << "[t=" << sc_time_stamp() << "] consume_delta fired\n";
    }
  }

  void consume_timed() {
    while (true) {
      wait(ev_timed);
      std::cout << "[t=" << sc_time_stamp() << "] consume_timed fired\n";
    }
  }

  SC_CTOR(showcase) {
    SC_THREAD(driver);
    SC_THREAD(consume_imm);
    SC_THREAD(consume_delta);
    SC_THREAD(consume_timed);
  }
};

int sc_main(int /*argc*/, char* /*argv*/[]) {
  showcase s("s");
  sc_start();
  return 0;
}

Output:

[t=10 ns] driver: firing all three
[t=10 ns] consume_imm fired
[t=10 ns] consume_delta fired
[t=15 ns] consume_timed fired

Three observations.

First, consume_imm printed 10 ns and so did consume_delta. Both printed at simulation time 10ns, which is why people are tempted to call them equivalent. They are not. An immediate notification makes the waiting thread runnable for the current evaluation step — it runs in the same delta cycle as the notify() call, but only after the driver has yielded at its next wait(). The driver thread continues executing through all three notify() calls without interruption, then suspends on wait(100, SC_NS). Only then does the scheduler service the runnable consumers: consume_imm runs first (made runnable by the immediate notify), followed by consume_delta in the next delta cycle (still at time 10ns). If you have other processes that run between deltas, they will see the world differently when triggered by ev_imm versus ev_delta — that one-delta gap is observable.

Second, consume_timed printed 15 ns. The notify(5, SC_NS) advanced simulation time by 5ns relative to the moment the notification was scheduled (10ns), and the consumer ran at the resulting 15ns timestamp.

Third — and this is the subtle one — if the driver had fired the events before the consumers had a chance to register on them, the immediate notification would have been lost. sc_event is not buffered. A notification fired while no one is waiting evaporates. The 10ns delay at the top of driver() exists precisely to let the consumers reach their wait(ev_…) calls before the notifications go out. This is one of the most common bugs in event-driven SystemC code: the producer fires immediately at initialization, the consumer has not yet reached its wait(), and the event is silently lost. Delta notifications can be lost the same way: if a consumer has not yet reached its wait(ev_delta) call before the next delta cycle arrives, the notification fires into nothing. Only timed notifications scheduled for a future time give the consumer a guarantee of arriving in a "waiting" state — and even then, only if it reaches the wait before the scheduled time.

The practical rule: use notify() (immediate) for synchronous handshakes where the producer knows the consumer is already waiting. Use notify(SC_ZERO_TIME) (delta) when you want to release the current delta back to the scheduler before the consumer reacts. Use notify(t) for any timed delay greater than zero. Do not substitute one for another because "they look the same in this case".

Corner 2: sc_signal_resolved and multi-driver buses

The Intermediate section's three-writer example showed why sc_signal cannot safely model a multi-driver bus. The LRM defines sc_signal_resolved (for single bits) and sc_signal_rv<N> (for N-bit vectors) precisely for this case. Both carry a resolution function that combines all current drivers into one current value at the end of each update phase.

The resolution table is four-valued: SC_LOGIC_0 (driven low), SC_LOGIC_1 (driven high), SC_LOGIC_Z (high-impedance / not driven), SC_LOGIC_X (conflict / unknown). The rules: any driver writing SC_LOGIC_Z does not contribute. If exactly one driver contributes a defined value, that value wins. If two drivers contribute different defined values, the result is SC_LOGIC_X. Two drivers contributing the same defined value (both SC_LOGIC_1, both SC_LOGIC_0) agree and produce that value — this is how wired-AND, wired-OR, and bidirectional bus patterns work. This is the same semantics as VHDL's std_logic and SystemVerilog's tri net.

A short complete program that exhibits the resolution rule:

// file: resolved_bus.cpp
// Build: g++ -std=c++17 -I$SYSTEMC_HOME/include -L$SYSTEMC_HOME/lib-linux64 \
//            resolved_bus.cpp -o resolved_bus -lsystemc

#include <systemc.h>
#include <iostream>

SC_MODULE(driver_a) {
  sc_inout_resolved bus;
  void drive() { bus.write(SC_LOGIC_1); }
  SC_CTOR(driver_a) { SC_METHOD(drive); }
};

SC_MODULE(driver_b) {
  sc_inout_resolved bus;
  sc_in<bool> mode;
  void drive() { bus.write(mode.read() ? SC_LOGIC_0 : SC_LOGIC_Z); }
  SC_CTOR(driver_b) { SC_METHOD(drive); sensitive << mode; }
};

int sc_main(int /*argc*/, char* /*argv*/[]) {
  sc_signal_resolved bus;
  sc_signal<bool>    mode;

  driver_a a("a"); a.bus(bus);
  driver_b b("b"); b.bus(bus); b.mode(mode);

  mode.write(false);                        // B drives Z, A drives 1
  sc_start(SC_ZERO_TIME);
  std::cout << "mode=false -> bus = " << bus.read() << "\n";  // expect 1

  mode.write(true);                         // B drives 0, A drives 1
  sc_start(SC_ZERO_TIME);
  std::cout << "mode=true  -> bus = " << bus.read() << "\n";  // expect X
  return 0;
}

Expected output:

mode=false -> bus = 1
mode=true  -> bus = X

Use sc_signal_resolved whenever your design genuinely has multiple drivers on the same wire — open-drain interrupt lines, shared tri-state buses, wired-OR signals. Do not use it as a general substitute for sc_signal<bool>: the resolution function makes it slower, and the four-valued logic propagates X aggressively in ways that can mask real bugs in unrelated code.

Corner 3: unspecified evaluation order makes some "working" code non-portable

Consider a process whose behavior depends on the order in which two other processes ran in the same delta. The LRM (§4.2.1.3) says the runnable set is processed in an order chosen by the implementation. The PoC simulator happens to use insertion order. Commercial simulators do not all agree.

The trap: a designer writes a model whose behavior is correct on their machine, ships it to a colleague using a different simulator, and the behavior changes. Common variants of this bug:

  • An SC_METHOD that increments a shared counter on each invocation; the final counter value depends on how many times the kernel re-runs it inside one delta, which depends on the order other writers in that delta touched its sensitivity signal.
  • A multi-writer sc_signal whose "last write wins" outcome depends on the order the writers were scheduled (the Intermediate section's example 2 is exactly this).
  • A debug print that interleaves output from multiple processes; the print order is not standardised and varies even between runs on the same kernel (rare, but possible if the implementation uses any nondeterminism).

The defence: never write code whose correctness depends on the order independent runnable processes are picked from the runnable set. If you need ordering, model it explicitly with a sequence of events. If you need atomic accumulation, use a single arbiter process. If you suspect a delta-order dependency, run the program under at least two different SystemC builds (e.g., Accellera 2.3.4 and a commercial kernel) and confirm the output matches. If it does not, that is a bug in your model, not the kernels.

Corner 4: dynamic sensitivity in SC_METHOD (forward link to Part 4)

SC_METHOD processes are statically sensitive by default. They can also use next_trigger() inside their body to set a one-shot dynamic sensitivity for the next activation. The catch: next_trigger() only changes the trigger for the next run. The static sensitivity stays in effect for subsequent runs unless next_trigger() is called every time. Misuse pattern: a method calls next_trigger(some_event) thinking it has switched to event-driven behaviour permanently, only to discover the next clock-edge activation still fires it via the original static sensitivity. This trap, and the corresponding wait() semantics for SC_THREAD, are the main subject of Part 4 (Processes & Sensitivity: SC_METHOD, SC_THREAD, SC_CTHREAD). Worth knowing it exists; do not work around it here.

Corner 5: debugging a delta-cycle livelock

When a kernel hangs because the runnable set never empties, the symptom is identical to any other infinite loop — sc_start does not return — but the diagnosis is different. The simulator is not stuck in your code; it is faithfully running an unbounded chain of delta cycles, each fully draining its runnable set into a new runnable set of the same size. Standard debuggers show you parked inside sc_simcontext::crunch(), which is not helpful.

The first diagnostic is to bound sc_start. Replace sc_start() (run forever) with sc_start(1, SC_NS) and add a print before and after. If the print after never appears even though no timed events should be pending in that 1ns window, you are looking at a delta livelock — by definition the kernel will exhaust its time budget on delta cycles at the same simulation timestamp without ever advancing time. The second diagnostic is to instrument every SC_METHOD and SC_THREAD body with a print at entry. Run for 100ms of wall-clock time then Ctrl+C and look at the tail of the output. If the same process names recur over and over with no time advance, those processes are your livelock loop. Trace the signal dependencies between them; one of them has an unguarded write that always changes the signal value.

The third diagnostic is more invasive but more definitive: instrument every sc_signal write with a print of the old value, new value, and writer's process name. The first cluster of writes that show non-stop oscillation (signal A changing, then B, then A, then B, repeatedly) identifies the unguarded chain. Add a prev_value check to the offending writer (if (next == prev_value) return;) and the livelock resolves.

A subtler variant of the same bug: a process that writes to a signal it is itself sensitive to, with no guard, causes the kernel to add the process back to the runnable set every update phase, forever. This is a particularly easy bug to introduce when copy-pasting code, and it is invisible until you actually run the simulation.

Corner 6: sc_signal initial values and the initialization delta

The very first delta cycle of any simulation — sometimes called the initialization delta — runs every process that does not have dont_initialize() set, exactly once, regardless of sensitivity. This is a special case in the LRM (§4.2.1.2): the initial set of runnable processes is "all processes that did not request dont_initialize". The intent is to give every process one chance to set up its outputs and event subscriptions before the simulation starts reacting to changes.

The trap: an SC_METHOD registered with sensitive << some_signal will run once at initialization regardless of whether some_signal has changed. It will see whatever value some_signal was constructed with, which is the default-constructed value of its T (zero for arithmetic types, false for bool, etc.) unless sc_main initialised it first. This first activation often produces output that does not appear in any sensitivity-based mental model of the design, surprising new users.

The Beginner section's example used dont_initialize() on the reader for exactly this reason: without it, the reader would have run once at initialization with in.read() = 0, producing an extra trace line that did not match the lesson. In production code, the choice between calling dont_initialize() and letting the initialization activation happen is genuinely design-dependent. Combinational logic models almost always want the initialization activation (so outputs settle to the correct function of inputs immediately). Edge-triggered or event-reactive logic almost always wants dont_initialize() so spurious activations do not leak into the trace.

Version differences

Kernel scheduling behaviour — the evaluate / update / delta loop, the three notify forms, the value-changed event for sc_signal, the implementation-defined runnable-set ordering — is identical across SystemC 2.3.1, 2.3.3, and 2.3.4. No changes have been made to the schedule in these releases. Differences between these releases are in build system, C++17 conformance, TLM-2.0 utilities, and some sc_vector helpers. Anything in this post compiles and behaves the same on any 2.3.x release. If you are reading this on a pre-2.3 release (which would be 2.2.0 from 2007), the event-binding semantics were different in ways that do affect the sc_event examples above, and the recommendation is to upgrade rather than work around the differences.

Hands-on exercise

Build a three-process design and predict the number of delta cycles before running.

Design: a module containing three SC_METHOD processes — A, B, and C — and three sc_signal<int> channels — s1, s2, s3. Wire them so that:

  • Process A is sensitive to s1 and writes to s2.
  • Process B is sensitive to s2 and writes to s3.
  • Process C is sensitive to s3 and prints the value it sees.

In sc_main, initialise all three signals to 0. Then write a value (say 42) to s1 and call sc_start(SC_ZERO_TIME). Before running, write down on a piece of paper: how many delta cycles will the kernel run? Which delta will each process fire in? What will the simulation time be when C prints?

Then run it with std::cout prints at the start of each process showing the value the process saw, and add sc_time_stamp() to each print. Compare what you predicted to what the kernel did.

Hint: count both evaluate phases and update phases — they alternate. And remember the seed write from sc_main is its own initial buffered write that takes one update phase to commit.

When you finish, modify the design so that C also writes back to s1 (creating a loop), guarded by a counter so the loop terminates after, say, 5 iterations. Predict and verify the delta count again. This is the canonical pattern for a small event-driven state machine, and being able to count its delta cycles in your head is the skill you should walk out of this post with.

Common mistakes

  • Reading a signal right after writing it and expecting the new value. The write is buffered. signal.write(x); int y = signal.read(); puts the old value of signal into y. Fix: read after a wait() (if you are in a thread), or wait for the value-changed-event on the consumer side (if you are in a method).
  • Using notify(SC_ZERO_TIME) when you wanted notify() (or vice versa). The two differ by one delta cycle. In a synchronous handshake where both producer and consumer run on the same clock edge, the one-delta difference can shift a sample by one cycle. Fix: pick the form that matches your intent and document why. If the producer knows the consumer is already waiting, immediate. If not, delta.
  • Letting a delta-cycle chain run forever. If process A writes to signal s that triggers process B, which writes back to a signal that triggers A, with no terminating guard, the runnable set never empties. sc_start never returns. Simulation hangs. Fix: every delta-cycle loop must have a guard that eventually stops a write from firing the chain. Verify by limiting sc_start(t, SC_NS) to a small time and watching whether it returns.
  • Multiple SC_METHOD processes writing to the same sc_signal in the same evaluation. "Last writer wins" is true, but which writer is last is implementation-defined. Fix: use sc_signal_resolved if you genuinely have multiple drivers; otherwise restructure so only one process writes the signal.
  • Forgetting dont_initialize() on a reactive process. Without it, every SC_METHOD runs once during the initialization delta with whatever values the signals had at construction. This often produces a spurious early activation that confuses the trace. Fix: call dont_initialize() on any process that should only react to subsequent changes, not the initial state.
  • Confusing simulation time with delta count. Two events that both happen "at 10ns" may be in different delta cycles. If you correlate trace output by sc_time_stamp() alone, you may conclude two unrelated events were simultaneous when they were one delta apart. Fix: in any debug-heavy code, print both sc_time_stamp() and a delta-tag (a manual counter you increment per wait(SC_ZERO_TIME)).

Recap

After working through this post you can now:

  • State the 3 phases of the SystemC kernel's per-delta loop from memory.
  • Predict the exact number of delta cycles a small SystemC program will run before sc_start(SC_ZERO_TIME) returns.
  • Explain why sc_signal::write() does not immediately change the value sc_signal::read() returns inside the same evaluate phase.
  • Distinguish immediate, delta, and timed notification of an sc_event and pick the right one for a given handshake.
  • Recognise a multi-writer collision on a single sc_signal and choose sc_signal_resolved when multi-driver semantics are actually wanted.
  • Explain why the order independent runnable processes execute within a delta is implementation-defined, and avoid designs that depend on it.
  • Recognise and fix a delta-cycle livelock (an unguarded chain that prevents the runnable set from emptying).
  • Read a kernel trace and tell apart effects that happened in the same evaluate phase, in the same delta but different evaluate steps, in different deltas at the same time, and at different simulation times.
  • Use dont_initialize() correctly to suppress spurious initial activations.
  • Diagnose delta-cycle bloat in a slow-running model by instrumenting SC_METHOD activations per clock edge.

Further reading

Standards

  • IEEE Std 1666-2011, IEEE Standard for Standard SystemC® Language Reference Manual, §4.2.1 (scheduling algorithm), §4.2.1.3 (update phase and delta cycles), §5.10 (event and notify semantics), §6.4 (sc_signal).

Vendor and consortium documents

  • Accellera Systems Initiative, SystemC 2.3.x User Guide, sections on the scheduler and channels.
  • Accellera Systems Initiative, SystemC Proof-of-Concept Implementation, source file src/sysc/kernel/sc_simcontext.cpp — function sc_simcontext::crunch(). This is the kernel scheduling loop in code form.

Textbooks

  • Doulos, SystemC Golden Reference Guide — the simulation-kernel and event-notification sections.
  • Black, Donovan, and Tahar, SystemC: From the Ground Up (2nd ed.) — the simulation-kernel chapter and the events / dynamic-sensitivity chapter.
  • Bhasker, A SystemC Primer — sections on processes, events, and the simulation kernel.
  • Grötker, Liao, Martin, and Swan, System Design with SystemC — scheduler and channel chapters.

Training notes

  • EuroPractice SystemC training material, kernel mental-model module.
  • Doulos online SystemC tutorials, scheduler and event-handling segments.

Next in this section

→ Part 4: Processes & Sensitivity — SC_METHOD, SC_THREAD, SC_CTHREAD — the three process kinds in terms of their kernel-level implementation (callback vs coroutine), how static and dynamic sensitivity interact with wait() and next_trigger(), and the decision table for choosing the right process kind for a given modelling task.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment