4. SystemC Tutorial - Processes & Sensitivity: SC_METHOD, SC_THREAD, SC_CTHREAD

Why this matters

Rewritten 2026-05-23 with deeper first-principles material.

Two engineers are debugging the same SystemC bug. One has written the testbench as a series of SC_METHOD processes, each pulling state forward in tiny reactive callbacks. The other has written the equivalent test as a single SC_THREAD that drives stimulus sequentially using wait() calls. They have the same observable behaviour and the same bug. The first engineer reads the trace, sees a method fire when they did not expect it, and spends an hour tracing the sensitivity list to figure out which signal change triggered it. The second engineer reads their own code top-to-bottom, follows the thread's wait() calls, and finds the bug in fifteen minutes. The first model is not wrong. The bug is not subtle. What separates the two debugging sessions is the kind of process each engineer chose, and how much of the simulation's behaviour is reactive (event-driven) versus sequential (control-flow-driven). Knowing when to reach for SC_METHOD and when to reach for SC_THREAD — and why SC_CTHREAD is almost never the right answer in modern code — is the single most important architectural decision in SystemC modelling.

That is what this post fixes. You will leave able to explain the three process kinds in terms of their kernel-level implementation (callback vs coroutine), to predict how sensitivity lists interact with wait() and next_trigger(), to choose the right process kind for a given modelling task without thinking about it, and — in the Advanced section — to debug the kernel-level errors that follow from mis-registering a process or forgetting dont_initialize().

Prerequisites

Mental model (first principles)

The SystemC kernel is a scheduler. Its job is to maintain a runnable set of processes, run them in the evaluate phase, commit their writes in the update phase, then re-evaluate based on new sensitivity triggers. The kernel does not care about the language-level distinction between SC_METHOD, SC_THREAD, and SC_CTHREAD in the abstract — it cares about how each process is invoked and suspended. Those two implementation details differ between the three kinds, and from those two details everything else follows.

                Process kinds at the kernel level
                ─────────────────────────────────

   SC_METHOD                   SC_THREAD                   SC_CTHREAD
   ─────────                   ─────────                   ──────────
   function pointer            coroutine (own stack)       coroutine (own stack)
   invoked by kernel           kernel switches to it       kernel switches to it
   runs to completion          runs until wait()           runs until wait()
   cannot wait()               can wait(time/event/...)    wait(N) = N clock cycles
   re-triggers from            wait() with no args         hardwired to one clock
   static sensitivity          uses static sensitivity     edge specified at
   every delta                 once, then dynamic          construction

Three kinds, three different relationships with the kernel.

SC_METHOD is the simplest. The kernel holds a function pointer. When something in the method's sensitivity list fires, the kernel adds the method to the runnable set. In the evaluate phase, the kernel calls the function. The function runs straight-line C++ and returns. That's it — no stack to manage, no suspension state, no context to save. The method is re-triggerable by definition: every invocation is independent. If you need state to persist between invocations, that state lives in the module's members (since the method is a member function), not in any local variable of the function itself.

SC_THREAD is more involved. The kernel maintains a separate user-space stack for the thread (typically allocated on the heap, default size 64 KB but configurable). When the kernel decides to run the thread, it context-switches: saves the kernel's registers, restores the thread's registers, jumps to the thread's instruction pointer. The thread runs until it calls wait(...), at which point the thread saves its registers and the kernel takes back control. The kernel notes the wait condition; when that condition is satisfied, the thread becomes runnable again and the kernel will context-switch back into it on the next evaluation step. The thread's local variables and call stack persist across wait() calls — that is the whole point. Sequential behaviour with persistent state between scheduling points.

SC_CTHREAD is a specialised SC_THREAD whose static sensitivity is hardwired to a clock-edge event at construction. Inside an SC_CTHREAD, wait() with no arguments means "wait for the next event on the clock edge I was constructed with"; wait(5) means "wait for 5 occurrences of that event." This counts clock cycles. It is convenient when you are writing pure clocked synchronous behaviour and want to count cycles like wait(3); count <= count + 1; wait(1); — but it is also rigid: an SC_CTHREAD cannot be sensitised to a non-clock event, cannot mix clock-cycle waits with event waits cleanly, and is structurally different enough from the other two kinds that most modern code prefers SC_THREAD(fn) { wait(clk.posedge_event()); ... } for the same result with more flexibility.

The kernel's scheduling decisions are identical for all three kinds: a runnable process is dispatched in the evaluate phase, and the kind only changes how the dispatch happens. The choice you make at registration time encodes "this code is reactive logic" (method), "this code is sequential behaviour with persistent state" (thread), or "this code is clocked HLS-style synchronous logic" (cthread). Pick wrong and the code still works, but you fight the framework.

Beginner: First Principles

The smallest example that exercises all three process kinds in one module:

// file: process_kinds.cpp
#include <systemc.h>

SC_MODULE(three_kinds) {
    sc_in<bool>  clk;
    sc_in<bool>  trigger;

    int method_count = 0;
    int thread_count = 0;
    int cthread_count = 0;

    void on_trigger() {
        ++method_count;
        std::cout << "[" << sc_time_stamp()
                  << "] METHOD fired (count = " << method_count << ")\n";
    }

    void thread_body() {
        while (true) {
            wait(trigger.posedge_event());
            ++thread_count;
            std::cout << "[" << sc_time_stamp()
                      << "] THREAD resumed (count = " << thread_count << ")\n";
        }
    }

    void cthread_body() {
        while (true) {
            wait();   // waits for clk.pos() (set at construction)
            ++cthread_count;
            std::cout << "[" << sc_time_stamp()
                      << "] CTHREAD on clk edge (count = " << cthread_count << ")\n";
            wait(2);  // wait for 2 more clock cycles before next iteration
        }
    }

    SC_CTOR(three_kinds) {
        SC_METHOD(on_trigger);
        sensitive << trigger;
        dont_initialize();

        SC_THREAD(thread_body);

        SC_CTHREAD(cthread_body, clk.pos());
    }
};

int sc_main(int, char**) {
    sc_clock          clk("clk", 10, SC_NS);
    sc_signal<bool>   trig;

    three_kinds dut("dut");
    dut.clk(clk);
    dut.trigger(trig);

    sc_start(5, SC_NS);
    trig.write(true);
    sc_start(5, SC_NS);
    trig.write(false);
    sc_start(5, SC_NS);
    trig.write(true);
    sc_start(50, SC_NS);
    return 0;
}

Expected output (kernel evaluation order between processes at the same time is implementation-defined; this is the Accellera-PoC ordering):

[5 ns] METHOD fired (count = 1)
[5 ns] THREAD resumed (count = 1)
[10 ns] CTHREAD on clk edge (count = 1)
[10 ns] METHOD fired (count = 2)
[15 ns] METHOD fired (count = 3)
[15 ns] THREAD resumed (count = 2)
[30 ns] CTHREAD on clk edge (count = 2)
[40 ns] CTHREAD on clk edge (count = 3)
[60 ns] CTHREAD on clk edge (count = 4)

(Note: the relative ordering of [10 ns] CTHREAD and [10 ns] METHOD is implementation-defined — both become runnable in the same delta (the clock posedge and trig going false both commit at 10 ns). The Accellera PoC runs them in registration order. The salient structure — method/thread firing on trigger posedges and the cthread counting clock cycles with wait(2) gaps — is stable.)

Walk through what just happened. Three processes registered in one constructor body, three different relationships with the kernel.

The SC_METHOD(on_trigger) plus sensitive << trigger and dont_initialize() registers a method process that triggers on every change of trigger. We add dont_initialize() because otherwise the method runs once at simulation start with trigger = false (default) and pollutes the output with [0 s] METHOD fired. With dont_initialize(), the first firing is at 5 ns when sc_main writes trig = true (which becomes visible in the next delta after the write commits). The method runs to completion, returns, and the kernel forgets about it until the next change.

The SC_THREAD(thread_body) registers a thread. The kernel context-switches into thread_body at simulation start; the thread runs to its first wait(trigger.posedge_event()) and suspends. When trig rises at 5 ns, the kernel resumes the thread; ++thread_count runs; the print runs; the loop iterates back to wait(...) and the thread suspends again. The local variable while (true) loop persists because the thread has its own stack.

The SC_CTHREAD(cthread_body, clk.pos()) registers a clocked thread sensitive to the clock's positive edge. wait() with no argument inside this CTHREAD means "wait for the next clk.pos() event." wait(2) means "wait for 2 more clk.pos() events." So the body executes: wait for first clk.pos (at 10 ns), increment to count=1, print, then wait(2) — two more edges at 20 ns and 30 ns — then resumes at 30 ns (count=2). The next wait() fires at 40 ns (count=3), then wait(2) passes 50 and 60 ns, resuming at 60 ns (count=4). The final sc_start(50, SC_NS) ends at 65 ns, so the next wait(2) endpoints at 70 and 80 ns fall outside the simulation window.

A common confusion here: why does the method fire at 5 ns rather than at the next clk.pos? Because the method is sensitive to trigger, not to clk. The method fires whenever trigger changes. sc_main wrote trig = true between sc_start calls, which is a pending write that commits when the next sc_start resumes the kernel — at time 5 ns. The method's "first firing time" has nothing to do with the clock.

Intermediate: How It Really Works

The Beginner section showed one example of each kind. The Intermediate section exposes what makes each kind shine — and what makes each kind a poor fit when used wrong.

When SC_METHOD is the right choice

SC_METHOD is the right choice for reactive logic — code that needs to recompute a derived value or take a small action whenever its inputs change. The canonical pattern is combinational logic in a model: given inputs A and B, compute output Y. Every change of A or B triggers a re-evaluation; the evaluation is straight-line code; the result lands in an output signal.

// file: mux21.cpp
#include <systemc.h>

SC_MODULE(mux21) {
    sc_in<bool>          sel;
    sc_in<sc_uint<8>>    a, b;
    sc_out<sc_uint<8>>   y;

    void eval() {
        y.write(sel.read() ? b.read() : a.read());
    }

    SC_CTOR(mux21) {
        SC_METHOD(eval);
        sensitive << sel << a << b;
    }
};

The eval method runs whenever any of sel, a, b changes. It has no internal state — every invocation produces a result purely from its inputs. This is exactly what RTL combinational logic does, and the model maps to RTL one-to-one.

When SC_THREAD is the right choice

SC_THREAD is the right choice for sequential logic — code whose behaviour is naturally expressed as a sequence of steps separated by waits. The canonical example is stimulus generation:

// file: stimulus.cpp
#include <systemc.h>

SC_MODULE(stimulus) {
    sc_in<bool>          clk;
    sc_out<sc_uint<8>>   data;
    sc_out<bool>         valid;

    void drive() {
        valid.write(false);
        data.write(0);

        wait(20, SC_NS);   // initial settling

        for (int i = 1; i <= 10; ++i) {
            wait(clk.posedge_event());
            data.write(i);
            valid.write(true);
            std::cout << "[" << sc_time_stamp()
                      << "] drove " << i << "\n";

            wait(clk.posedge_event());
            valid.write(false);
        }

        wait(10, SC_NS);
        sc_stop();
    }

    SC_CTOR(stimulus) {
        SC_THREAD(drive);
    }
};

Trying to write this as an SC_METHOD would force you to encode the loop counter and the "which step am I on" state as module members, with an FSM-style switch statement inside the method. That works, but it is harder to read, harder to maintain, and fights the framework. The thread version reads like the testbench script it is.

When SC_CTHREAD is the right choice (rarely)

SC_CTHREAD is the right choice for clocked HLS-style synchronous behaviour where the dominant cost of every wait is "wait one or more clock cycles." It was the canonical pattern in early high-level-synthesis flows (the Cynthesizer / Forte days) where the synthesis tool expected the input model to have one clocked thread per RTL block.

// file: cthread_counter.cpp
#include <systemc.h>

SC_MODULE(counter4) {
    sc_in_clk     clk;
    sc_in<bool>   rst_n;
    sc_out<sc_uint<4>> count;

    void run() {
        while (true) {
            if (!rst_n.read()) {
                count.write(0);
                wait();   // next clk edge
                continue;
            }
            sc_uint<4> v = count.read();
            count.write(v + 1);
            wait();
        }
    }

    SC_CTOR(counter4) {
        SC_CTHREAD(run, clk.pos());
        reset_signal_is(rst_n, false);   // optional: kernel calls reset on rst_n low
    }
};

In modern non-HLS code, the same module written with SC_METHOD and sensitive << clk.pos() is shorter, equally clear, and more easily integrated with mixed-flavour code. Unless your project's existing convention is SC_CTHREAD, prefer SC_METHOD for synchronous logic.

Dynamic sensitivity inside SC_THREAD

The most powerful pattern in SystemC is SC_THREAD with wait(...) overloads that take events, event lists, time, or combinations. The kernel supports five forms:

wait call Suspends until
wait() The next event on the static sensitivity list
wait(time) The given simulation time has elapsed
wait(event) The given event fires
`wait(event_a \ event_b)` Either event fires (OR)
wait(event_a & event_b) Both events have fired (AND)
wait(time, event) Either time elapses OR event fires, whichever first

The OR and AND forms are written with overloaded | and & on sc_event_or_list and sc_event_and_list respectively. Example:

// file: dynamic_sensitivity.cpp
#include <systemc.h>

SC_MODULE(observer) {
    sc_in<bool>  done;
    sc_in<bool>  error;

    void watch() {
        std::cout << "[" << sc_time_stamp()
                  << "] observer started; waiting for done or error\n";
        wait(done.value_changed_event() | error.value_changed_event());
        if (done.read()) {
            std::cout << "[" << sc_time_stamp() << "] saw done\n";
        }
        if (error.read()) {
            std::cout << "[" << sc_time_stamp() << "] saw error\n";
        }
        wait(50, SC_NS);
        std::cout << "[" << sc_time_stamp() << "] observer exits\n";
    }

    SC_CTOR(observer) { SC_THREAD(watch); }
};

The wait(... | ...) returns as soon as either signal changes. The thread then continues with whatever logic comes after. Dynamic sensitivity is what lets you write threads whose wakeup conditions vary at runtime — impossible to express cleanly with SC_METHOD.

next_trigger() in SC_METHOD

The method-process equivalent of dynamic sensitivity is next_trigger(...). Calling next_trigger(some_event) inside the method changes the trigger condition for the next invocation only; subsequent invocations revert to the static sensitivity list.

// file: next_trigger.cpp
#include <systemc.h>

SC_MODULE(armed) {
    sc_in<bool> arm;
    sc_in<bool> trigger;

    bool armed_state = false;

    void on_event() {
        if (!armed_state && arm.read()) {
            armed_state = true;
            std::cout << "[" << sc_time_stamp() << "] armed; waiting on trigger\n";
            next_trigger(trigger.posedge_event());   // override for next call
        } else if (armed_state && trigger.read()) {
            armed_state = false;
            std::cout << "[" << sc_time_stamp() << "] fired!\n";
            // No next_trigger -> reverts to static sensitivity (arm).
        }
    }

    SC_CTOR(armed) {
        SC_METHOD(on_event);
        sensitive << arm;
        dont_initialize();
    }
};

next_trigger lets you build a method that behaves like a small state machine without writing an explicit FSM. The kernel still treats the method as run-to-completion — next_trigger only modifies the trigger condition the kernel watches after the current invocation returns.

Decision table: which process kind?

You want to model Use
Combinational logic (output is pure function of inputs) SC_METHOD, sensitive << inputs
Registered logic (output = clocked state machine) SC_METHOD, sensitive << clk.pos()
Stimulus / test sequencer SC_THREAD with wait(...) calls
FIFO / queue consumer that waits on enqueue events SC_THREAD with wait(event)
HLS-targeted synchronous block (rare in new code) SC_CTHREAD(fn, clk.pos())
Reactive logic that takes occasional self-modifying steps SC_METHOD with next_trigger(...)
Observer that waits for one of several conditions SC_THREAD with `wait(ev1 \ ev2 \ ev3)`
Periodic process running every fixed time interval SC_THREAD with wait(time) loop, OR SC_METHOD clocked

Performance comparison

A rough rule of thumb on a modern x86 host:

  • SC_METHOD invocation cost: ~50–100 ns of host CPU per invocation.
  • SC_THREAD resumption cost: ~500–1000 ns of host CPU per resumption (the context switch dominates).

For most models, this difference is irrelevant — both are dwarfed by the actual work the process does. For very hot loops in cycle-accurate models — a CPU instruction-fetch model that runs millions of times per simulated second — preferring SC_METHOD over SC_THREAD can give a 2–5× simulation speedup. For testbench code, the difference vanishes.

Walking through one delta with three process kinds

A concrete trace to cement the kernel-level picture. Consider this hierarchy:

SC_MODULE(trace_demo) {
    sc_in<bool>      clk;
    sc_in<bool>      go;
    sc_event         done;
    sc_signal<int>   data;

    // Method: sensitive to data
    void on_data() {
        std::cout << "  [METHOD] data=" << data.read() << "\n";
    }

    // Thread: waits on go, then writes data
    void seq_thread() {
        wait(go.posedge_event());
        std::cout << "  [THREAD] go seen; writing data=42\n";
        data.write(42);
        wait(SC_ZERO_TIME);
        std::cout << "  [THREAD] after yield, data is " << data.read() << "\n";
        done.notify();
    }

    // CTHREAD: counts clocks, fires done after 3
    void clocked_run() {
        for (int i = 0; i < 3; ++i) {
            wait();
            std::cout << "  [CTHREAD] tick " << (i+1) << "\n";
        }
        std::cout << "  [CTHREAD] notifying done\n";
        done.notify();
    }

    SC_CTOR(trace_demo) {
        SC_METHOD(on_data);
        sensitive << data;
        dont_initialize();

        SC_THREAD(seq_thread);

        SC_CTHREAD(clocked_run, clk.pos());
    }
};

What does the kernel actually do across the first few deltas of sc_start(50, SC_NS) with go going high at 15 ns? Step by step:

  • Time 0, initialization phase:
  • on_data is NOT initialized (we called dont_initialize). So it does not run.
  • seq_thread is initialized: kernel context-switches into it, runs until wait(go.posedge_event()), suspends.
  • clocked_run is initialized: kernel context-switches in, runs until wait() (the implicit clock-edge wait), suspends.
  • Time 0 to 10 ns: runnable set empty. Time advances to next pending event — the clock's first edge at 10 ns (assuming a default clock with posedge_first = true and period 10 ns, this is the first rising-edge transition at 10 ns; the falling edge at 5 ns is not relevant for the pos() sensitivity).
  • Time 10 ns, evaluate phase: clocked_run is resumed (its wait() for the clock edge completes). Prints [CTHREAD] tick 1. Calls wait() again, suspends.
  • Time 15 ns: sc_main calls trig.write(true) (or however go is driven). Pending. Next sc_start resumes the kernel.
  • Time 15 ns, evaluate phase: seq_thread resumes (its wait(go.posedge_event()) is satisfied). Prints [THREAD] go seen; writing data=42. Calls data.write(42) — pending. Calls wait(SC_ZERO_TIME) — suspends, scheduled to resume next delta.
  • Time 15 ns, update phase: data commits from 0 to 42. value_changed_event fires.
  • Time 15 ns, post-update notification: on_data is sensitive to data — becomes runnable. seq_thread is also runnable (its wait(SC_ZERO_TIME) completes).
  • Time 15 ns, next delta evaluate: order is implementation-defined. In the Accellera PoC, registration-order ties: on_data runs first, prints [METHOD] data=42. Then seq_thread resumes, prints [THREAD] after yield, data is 42, calls done.notify() (immediate notification — but nothing is currently waiting on done, so it has no effect).
  • Time 20 ns: next clock edge. clocked_run resumes, prints [CTHREAD] tick 2, suspends.
  • Time 30 ns: [CTHREAD] tick 3. Then loop exits; prints [CTHREAD] notifying done; calls done.notify(). The thread function returns; the CTHREAD process terminates.

The exercise demonstrates how all three process kinds participate in the same scheduler with no special-case logic. The kernel does not "know" the process kind in any deep sense — it just dispatches according to the registered callback (for methods) or context-switches according to the saved stack (for threads). The kind only changes the implementation of dispatch, not its scheduling priority.

The "first useful pattern" — a clocked FSM as SC_METHOD

The most common pattern in real SystemC RTL-level models is a clocked finite-state machine implemented as a single SC_METHOD:

// file: fsm.cpp
#include <systemc.h>

SC_MODULE(simple_fsm) {
    sc_in<bool>       clk;
    sc_in<bool>       rst_n;
    sc_in<bool>       start;
    sc_in<bool>       done;
    sc_out<bool>      busy;

    enum state_t { IDLE, BUSY, FINISH };
    state_t state = IDLE;

    void seq() {
        if (!rst_n.read()) {
            state = IDLE;
            busy.write(false);
            return;
        }
        switch (state) {
            case IDLE:
                if (start.read()) {
                    state = BUSY;
                    busy.write(true);
                }
                break;
            case BUSY:
                if (done.read()) {
                    state = FINISH;
                }
                break;
            case FINISH:
                state = IDLE;
                busy.write(false);
                break;
        }
    }

    SC_CTOR(simple_fsm) {
        SC_METHOD(seq);
        sensitive << clk.pos();
        dont_initialize();
    }
};

A clocked state machine in 30 lines. The state is a module member (persistent across invocations). The transitions are an explicit switch. The output busy is updated synchronously with state changes. This pattern scales to any size FSM and maps 1:1 to RTL — the same FSM in SystemVerilog would be always_ff @(posedge clk) with the same switch logic.

If the same FSM were written as SC_THREAD, you would either need an outer while (true) loop with wait(clk.posedge_event()) at the top and a switch inside (essentially the same code), or you would need to express the state machine as a sequence of waits (wait until start; busy = true; wait until done; busy = false;), which only works if the state transitions are linear. For state machines with cycles or multiple paths, SC_METHOD is the clearer expression.

Cross-process synchronisation patterns (preview of Part 5)

Within a module, processes share state through member variables. Across modules — or even within a module when you want explicit synchronisation rather than incidental sharing — sc_event is the primary primitive. The full mechanics are covered in Part 5 (Events & Notifications), but knowing the patterns here will help you read the code in subsequent posts.

Pattern 1: producer notifies, consumer waits.

sc_event item_available;

// In producer thread:
queue.push_back(item);
item_available.notify();   // immediate notification

// In consumer thread:
while (queue.empty()) {
    wait(item_available);
}

Pattern 2: handshake via two events.

sc_event request, acknowledge;

// Initiator:
request.notify();
wait(acknowledge);

// Responder:
wait(request);
do_work();
acknowledge.notify();

Pattern 3: deadline with timeout.

// In responder thread (LRM-correct form):
sc_time deadline(100, SC_NS);
sc_time before = sc_time_stamp();
wait(deadline, request);          // wakes on request OR 100 ns elapse
if (sc_time_stamp() - before < deadline) {
    // Real request arrived before timeout
} else {
    // Timeout
}

(wait(time, event) is the LRM §5.2.17 form — it returns on whichever comes first: the time elapses, or the event fires. Part 5 covers the full semantics.)

Pattern 4: barrier — wait for N processes to all reach a checkpoint.

sc_event_and_list arrival_barrier;
arrival_barrier &= worker_1_arrived;
arrival_barrier &= worker_2_arrived;
arrival_barrier &= worker_3_arrived;

// Each worker thread:
some_event.notify();   // its own "arrived" event

// Coordinator thread:
wait(arrival_barrier);

All four patterns are bread-and-butter in real SystemC models — testbenches, scoreboards, monitors, and any model where multiple processes need to coordinate. Part 5 covers the full semantics, the three notification forms, and the subtle traps (the most common: "I sent an immediate notification but nobody was waiting yet"). For now, knowing the patterns exist is enough to read the testbench code in Part 6.

Quick reference: process registration syntax

A condensed cheat sheet for the syntactic forms:

// In SC_CTOR(name) { ... }:

SC_METHOD(member_function);
sensitive << signal_or_port;     // static sensitivity
sensitive << port.posedge_event();
sensitive << custom_event;
dont_initialize();               // optional — skip init-time run

SC_THREAD(member_function);
sensitive << signal_or_port;     // optional — used only by bare wait()
// (most threads do not need static sensitivity; they use wait(...) with args)

SC_CTHREAD(member_function, clk.pos());
reset_signal_is(rst_n, false);   // optional declarative reset
async_reset_signal_is(rst_n, false);  // optional async reset

Inside the process body, the dynamic-sensitivity APIs are:

// Inside SC_METHOD (each call sets behaviour for NEXT invocation):
next_trigger();                           // revert to static sensitivity
next_trigger(some_event);                 // one event
next_trigger(ev_a | ev_b);                // OR list
next_trigger(ev_a & ev_b);                // AND list
next_trigger(sc_time(100, SC_NS));        // time-based
next_trigger(sc_time(100, SC_NS), event); // time OR event

// Inside SC_THREAD or SC_CTHREAD (suspends the coroutine):
wait();                                   // wait on static sensitivity
wait(some_event);
wait(ev_a | ev_b);
wait(ev_a & ev_b);
wait(sc_time(100, SC_NS));
wait(100, SC_NS);                         // shorthand
wait(sc_time(100, SC_NS), event);

// Inside SC_CTHREAD only:
wait();                                   // 1 clock cycle (the construction clock)
wait(N);                                  // N clock cycles

That's the entire surface area of the process-management API. Everything else in SystemC builds on these primitives.

Mixing process kinds in one module

A single module can register any number of any kind of process. They all share the module's member variables (which is the canonical way to communicate between processes inside a module) and they all run in the same kernel scheduler — no synchronisation primitives needed beyond sc_event for cross-process signalling.

// file: mixed.cpp
#include <systemc.h>

SC_MODULE(producer_consumer) {
    sc_in<bool>     clk;
    sc_in<bool>     enq;
    sc_in<sc_uint<8>> enq_data;

    sc_event new_item_event;
    std::deque<sc_uint<8>> queue;

    // Producer: SC_METHOD on enq edge
    void on_enq() {
        if (enq.read()) {
            queue.push_back(enq_data.read());
            new_item_event.notify(SC_ZERO_TIME);   // delta notify; consumer wakes
        }
    }

    // Consumer: SC_THREAD that processes queue items one per clock
    void consume() {
        while (true) {
            if (queue.empty()) {
                wait(new_item_event);
            }
            sc_uint<8> v = queue.front();
            queue.pop_front();
            std::cout << "[" << sc_time_stamp()
                      << "] consumed " << v << "\n";
            wait(clk.posedge_event());
        }
    }

    SC_CTOR(producer_consumer) {
        SC_METHOD(on_enq);
        sensitive << enq.posedge_event();
        dont_initialize();

        SC_THREAD(consume);
    }
};

Two processes in one module: a method that reacts to enqueues, a thread that consumes one item per clock. They share the queue deque (since they're members of the same module) and the new_item_event (the method notifies, the thread waits). This is the bread-and-butter pattern for SystemC models where you want reactive behaviour AND sequential behaviour coexisting.

What happens when you pick the wrong kind

A small case study. The same behaviour — "every 10 ns, increment a counter and print it" — implemented three ways, with commentary on the trade-offs.

Attempt 1: SC_METHOD with self-scheduling via next_trigger.

// file: ticker_method.cpp (annotated)
SC_MODULE(ticker_m) {
    int count = 0;
    void tick() {
        std::cout << "[" << sc_time_stamp() << "] tick " << ++count << "\n";
        next_trigger(10, SC_NS);
    }
    SC_CTOR(ticker_m) {
        SC_METHOD(tick);
        // no static sensitivity — first call by initialization,
        // subsequent calls via next_trigger().
    }
};

This works. The first invocation happens at simulation start (initialization phase); each invocation schedules the next 10 ns out. Cost: zero stack overhead, lowest possible per-invocation cost. Readability: medium — the temporal structure (every 10 ns) is hidden inside a single line at the bottom of the function body.

Attempt 2: SC_THREAD with a while (true) loop.

// file: ticker_thread.cpp (annotated)
SC_MODULE(ticker_t) {
    void tick_loop() {
        int count = 0;
        while (true) {
            std::cout << "[" << sc_time_stamp() << "] tick " << ++count << "\n";
            wait(10, SC_NS);
        }
    }
    SC_CTOR(ticker_t) { SC_THREAD(tick_loop); }
};

This also works. The temporal structure is now explicit and reads top-to-bottom. The counter is a local variable (persistent because thread stacks persist). Cost: 64 KB of stack per instance, ~1 μs per resumption. Readability: high.

Attempt 3: SC_CTHREAD with a clock.

// file: ticker_cthread.cpp (annotated)
SC_MODULE(ticker_c) {
    sc_in<bool> clk;
    void tick_loop() {
        int count = 0;
        while (true) {
            wait();   // one clock cycle
            std::cout << "[" << sc_time_stamp() << "] tick " << ++count << "\n";
        }
    }
    SC_CTOR(ticker_c) { SC_CTHREAD(tick_loop, clk.pos()); }
};

// In sc_main: sc_clock clk("clk", 10, SC_NS); ticker_c t("t"); t.clk(clk);

This works if there is a clock to drive it. The temporal structure is "every clock cycle" — and the period is set by the external clock, not by the module. Cost: stack overhead (same as SC_THREAD) plus a clock channel. Readability: high if the model already has a clock, slightly worse if you have to instantiate a clock just for this.

The differences shake out as:

Concern SC_METHOD SC_THREAD SC_CTHREAD
Lines of code Slightly fewer Slightly more Slightly more
Stack overhead per instance 0 ~64 KB ~64 KB
Per-event CPU cost lowest medium medium
Readability of temporal structure Worst (hidden in next_trigger) Best (explicit loop with wait) Best
Couples module to a clock No No Yes
Cycle-accurate counting Awkward Awkward (need explicit count) Native (wait(N))

The right answer depends on what you value. For a periodic ticker that runs millions of times in a long simulation, SC_METHOD wins on performance. For a single ticker that runs occasionally and you want easy maintenance, SC_THREAD wins on readability. For HLS-targeted clocked logic, SC_CTHREAD is the framework's expected pattern.

In practice, the choice is rarely about performance — it is about which model expresses the intent of the code most clearly. Pick the one that reads like the thing you are modelling.

Advanced: Edge Cases & LRM Corners

Corner 1: wait() with no argument means different things in SC_THREAD vs SC_CTHREAD

Inside an SC_THREAD, wait() with no argument suspends until the next event on the static sensitivity list. If the thread has no static sensitivity (no sensitive << ... after registration), wait() will hang forever — the kernel marks the thread as runnable when something in its (empty) sensitivity list fires, which is never.

Inside an SC_CTHREAD, wait() with no argument suspends until the next occurrence of the clock edge that the CTHREAD was constructed with — which IS its static sensitivity, by construction. wait(N) for integer N means "wait for N occurrences of that clock edge." This is the conveniences of CTHREAD: cycle counting becomes natural.

The pitfall is mixing them up. If you mentally substitute SC_THREAD for SC_CTHREAD and try wait(3) expecting "3 clock cycles," you actually get "wait for 3 simulation-time units" — because wait(N) on an SC_THREAD with N an integer is implicitly converted to wait(sc_time(N, default_unit)). Three nanoseconds of wait, not three clock cycles. The bug is silent.

Corner 2: dont_initialize() semantics and its scope

dont_initialize() is a member of sc_module. When called, it sets a flag on the most recently registered process. The kernel checks this flag during initialization: if set, the process does not run during the initialization phase; if unset, the process runs once.

The scope rule means order matters:

// CORRECT: dont_initialize applies to method_a
SC_METHOD(method_a);
sensitive << some_signal;
dont_initialize();

SC_METHOD(method_b);
sensitive << other_signal;
// method_b will run at initialization
// WRONG: dont_initialize applies only to method_b
SC_METHOD(method_a);
sensitive << some_signal;

SC_METHOD(method_b);
sensitive << other_signal;
dont_initialize();

If you want both methods to skip initialization, you need two dont_initialize() calls, one after each registration.

Corner 3: SC_METHOD calling wait() — what actually happens

Calling wait(...) inside an SC_METHOD is an LRM violation but not a compile error — the function exists in scope. The kernel detects it at runtime and raises:

Error: wait() is only allowed in SC_THREAD and SC_CTHREAD processes;
       in SC_METHOD processes use next_trigger() instead.

The error fires the first time the method attempts to wait. The fix is mechanical: if you need wait(), change SC_METHOD to SC_THREAD and the registration macro at the top of the constructor.

Corner 4: SC_THREAD with no static sensitivity that calls bare wait()

SC_THREAD(my_thread);
// Note: no `sensitive << ...` after registration

If my_thread calls wait() with no argument, it suspends forever — the kernel watches an empty sensitivity list. This is a real bug pattern; the thread silently never wakes after the first wait, and the simulation may appear to deadlock with no error message. The fix is either to add static sensitivity (sensitive << some_event; after the SC_THREAD line) or to always use wait(...) with explicit arguments inside the thread.

Corner 5: reset_signal_is and async_reset_signal_is for SC_CTHREAD

SC_CTHREAD supports declarative reset:

SC_CTHREAD(my_run, clk.pos());
reset_signal_is(rst_n, false);    // synchronous reset, active low
// or
async_reset_signal_is(rst_n, false);  // asynchronous reset, active low

When the named reset signal is at its active level, the kernel restarts the CTHREAD from its entry point. This works only with SC_CTHREAD; ordinary SC_THREAD does not have an analogous declarative reset.

For modern SC_THREAD and SC_METHOD code, implement reset explicitly: check the reset signal inside the process body and re-initialise state when active. This is more verbose but more portable across HLS-vs-simulation contexts.

Corner 6: Thread stack size

The kernel allocates a stack for each SC_THREAD (and SC_CTHREAD). Default size is 64 KB. For threads that recurse or have large local arrays, this can overflow silently — typically presenting as a segfault that points into the kernel's context-switch code, not into the thread's body.

To set a non-default stack size:

SC_THREAD(big_thread);
set_stack_size(1024 * 1024);   // 1 MB stack

set_stack_size applies to the most recently registered thread (same "scope rule" as dont_initialize). For most threads the default is plenty; this knob is for unusually heavy modelling code.

Corner 7: Ordering of processes at the same delta cycle

When multiple processes become runnable in the same delta cycle, the kernel's scheduling order is implementation-defined per IEEE 1666-2011 §4.2.1.2. The Accellera reference implementation processes them in registration order, but other simulators may differ. Cross-reference: Part 3 (Delta Cycles) Corner 3 covers this and its implications for race conditions in detail.

The practical advice: never write code whose correctness depends on the order in which two processes in the same delta evaluate. If you need to enforce order, introduce an sc_event or a sequence of delta cycles to serialise.

Corner 8: wait(0, SC_NS) is a delta-cycle yield

A subtle pattern: wait(0, SC_NS) (or equivalently wait(SC_ZERO_TIME)) inside an SC_THREAD yields to the kernel and resumes in the next delta cycle at the same simulation time. The thread is not added to any event-based wait list; it is scheduled to resume after one delta.

This is useful for cases where a thread has just made some writes to signals and wants to see the committed values (after the update phase) before continuing. Without the yield, the thread sees its own pending writes as not-yet-committed values via read(). With the yield, the kernel processes the update phase, the writes commit, and the thread resumes with the new values visible.

SC_MODULE(self_observer) {
    sc_signal<int> sig;
    void run() {
        sig.write(42);
        int before = sig.read();   // before == 0 (not yet committed)
        wait(SC_ZERO_TIME);
        int after = sig.read();    // after == 42 (committed)
        std::cout << "before=" << before << " after=" << after << "\n";
    }
    SC_CTOR(self_observer) { SC_THREAD(run); }
};

Output: before=0 after=42. This is the same delta-cycle update-phase semantics covered in Part 3, applied to threads.

Corner 9: sc_spawn for dynamic process creation during elaboration

sc_spawn (defined in <sysc/utils/sc_spawn.h>) lets you register a process at runtime, not just at module construction. It returns an sc_process_handle for manipulation. The signature:

sc_process_handle h = sc_spawn(
    /*body*/    [this]() { my_function(); },
    /*name*/    "spawned_proc",
    /*options*/ nullptr  // or &spawn_options for non-default behaviour
);

The spawned process can be a method or thread depending on spawn_options. The sc_process_handle can later be used to kill, suspend, resume, or query the process.

The use cases are narrow: most modelling needs are met by static process registration in the constructor. sc_spawn is the right tool for testbench frameworks that need to launch new test sequences on demand (e.g., UVM-SystemC's sequence mechanism uses spawned threads internally).

Restriction: sc_spawn calls must happen during elaboration (before sc_start) OR from within a running process. They cannot be called from outside the SystemC context.

Corner 10: disable(), suspend(), resume(), kill() on sc_process_handle

Once you have an sc_process_handle, you can manipulate the process's runtime state:

  • disable() — process becomes runnable but the kernel will not actually run it; its waits still complete, the kernel just skips its evaluation.
  • suspend() — process is paused; even if events fire that would wake it, it does not run until resume() is called.
  • resume() — undoes suspend().
  • kill() — process is terminated; no further invocations regardless of events. Stack is freed (for threads).
  • valid() — true if the handle refers to a real process.
  • terminated() — true if the process has exited (returned normally or been killed).
sc_process_handle h = sc_spawn(..., "h");
// Later in some control thread:
h.suspend();
wait(100, SC_NS);
h.resume();

The combination of sc_spawn and the handle API is what makes dynamic test orchestration possible. For most modelling code, this is overkill — but knowing it exists explains how UVM-SystemC, SCV, and other frameworks build their sequence-control machinery.

Corner 11: sc_event_or_list and sc_event_and_list — the underlying types

wait(ev_a | ev_b) returns when either event fires. wait(ev_a & ev_b) returns when both have fired (in any order, possibly across multiple delta cycles). The | and & operators on sc_event construct sc_event_or_list and sc_event_and_list objects respectively.

For longer lists, you can build them explicitly:

sc_event_or_list any_done;
any_done |= thread_a_done;
any_done |= thread_b_done;
any_done |= thread_c_done;
wait(any_done);   // returns when any of the three fires

sc_event_and_list all_ready;
all_ready &= scoreboard_ready;
all_ready &= driver_ready;
all_ready &= monitor_ready;
wait(all_ready);   // returns when all three have fired

The AND list's semantics — "all three have fired" — is latched: once all three have fired (possibly at very different times), the list is satisfied; the next wait(all_ready) call resets the latch and waits for all three to fire again. This is the right primitive for "wait until all of these one-shot conditions have happened" without writing explicit state tracking.

Hands-on exercise

Build a handshake protocol module that implements a four-phase req/ack handshake:

  • Inputs: sc_in<bool> req;, sc_in<bool> data_valid;
  • Outputs: sc_out<bool> ack;
  • Behaviour: when req rises, wait one clock cycle, set ack high; wait until req falls; set ack low; loop.

Write this twice: once as an SC_THREAD with wait(req.posedge_event()) style calls, and once as an SC_METHOD with a state machine and next_trigger. Compare the two implementations in terms of:

  • Lines of code
  • Readability
  • Where the state lives (module member vs implicit in control flow)
  • What happens if you need to add a third state to the handshake (e.g., a "request-acknowledged" intermediate state)

There is no "right" answer to which is better — but the exercise will sharpen your intuition for when to reach for which kind.

Hint: the SC_THREAD version will be roughly half the lines of the SC_METHOD version, but the SC_METHOD version may be easier to formally verify because all states are explicit.

Common mistakes

  • Calling wait(N) in SC_THREAD expecting N clock cycles. It is N units of sc_time in the default time unit, not N clock cycles. Use for (int i = 0; i < N; ++i) wait(clk.posedge_event()); instead, or switch to SC_CTHREAD if cycle counting is the dominant concern.
  • dont_initialize() in the wrong position. It applies only to the most recently registered process. Order matters; check the call sites carefully when you have multiple processes in one constructor.
  • Calling wait() inside SC_METHOD. Runtime error. The fix is either to convert to SC_THREAD or to use next_trigger(...) for the dynamic-sensitivity equivalent.
  • Static sensitivity inadvertently empty in SC_THREAD. A thread with no sensitive << ... calling bare wait() deadlocks silently. Either add explicit sensitivity or always use wait(...) with arguments.
  • Local variables in SC_METHOD expecting persistence across invocations. They do not persist — every invocation starts with fresh locals. Persistent state must be a module member.
  • Sharing module members between processes without considering ordering. If two processes in the same module touch the same member variable in the same delta, the result is implementation-defined. Add explicit sequencing or use a separate sc_event to serialise.
  • Allocating large arrays as locals in SC_THREAD. A 100 KB local array will overflow the default 64 KB stack. Either bump the stack with set_stack_size, or move the array to be a module member.
  • Forgetting that SC_CTHREAD wait(N) counts events of its construction clock. If you wrote SC_CTHREAD(run, clk.pos()) and inside run you call wait(3), you wait for 3 rising edges of clk — not 3 edges of any other clock you might also reference. CTHREADs are sensitive only to the one clock they were constructed with; if you need multi-clock behaviour, use SC_THREAD and explicit wait(other_clk.posedge_event()).
  • Returning from inside an SC_THREAD. When the thread's function returns, the thread terminates and its stack is freed. The kernel does not consider this an error — but the process is gone and any state it owned via stack-local variables is gone with it. Usually you want while (true) { ... } to keep the thread alive; if you intentionally want to exit, terminate by returning, but don't accidentally fall off the end of the function.
  • Using next_trigger() outside an SC_METHOD. It's a no-op (or runtime error in some kernels) if called from a thread. The thread equivalent is wait(...). Symmetrically, calling wait() inside an SC_METHOD produces a runtime error. The two APIs are not interchangeable.

Recap

After this post, you can:

  • Explain the kernel-level difference between SC_METHOD (callback), SC_THREAD (coroutine with stack), and SC_CTHREAD (coroutine hardwired to a clock edge).
  • Choose the right process kind for a given modelling task without thinking about it.
  • Use static sensitivity (sensitive << ...) and dynamic sensitivity (wait(...), next_trigger(...)) correctly in each kind.
  • Express AND-of-events and OR-of-events conditions in a thread's wait list.
  • Use dont_initialize() correctly and explain why its position matters.
  • Recognise and debug the deadlock pattern of an SC_THREAD with empty static sensitivity calling bare wait().
  • Predict the relative simulation speed of method-heavy vs thread-heavy models.
  • Mix process kinds in a single module to express a producer (method) and consumer (thread) cleanly.

Further reading

  • IEEE 1666-2011 §5.2 (the chapter on module member functions and process semantics). Read §5.2.6–§5.2.18 in order; the LRM's coverage of process kinds is unusually clear, and §5.2.11 on thread semantics is required reading before writing any non-trivial SC_THREAD.
  • Accellera SystemC User's Guide — the processes chapter.
  • Accellera SystemC PoC source: src/sysc/kernel/sc_method_process.h, sc_thread_process.h, sc_cthread_process.h. Reading these three together is the fastest way to demystify what each kind actually is at the implementation level.
  • Doulos SystemC Golden Reference Guide — the processes section. The Doulos guide's coverage of next_trigger is particularly good.
  • Black, Donovan, Tahar, SystemC: From the Ground Up (2nd ed.) — the processes chapter walks through the kernel implementation in pseudo-code, which complements the LRM well and clears up most lingering ambiguity about how each process kind is dispatched at runtime.

Next in this section

→ Part 5: Events & Notifications — sc_event, the three notification forms (immediate, delta, timed), and the patterns for cross-process signalling that follow naturally from the process kinds you now know. The cross-process patterns previewed at the end of the Intermediate section get their full LRM treatment there, along with the most common SystemC pitfall: notifying an event that has no waiters yet, and silently losing the notification. By the end of Part 5 you will have everything you need to read and write multi-process SystemC models with confidence.

Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment