Killing fork/join Threads Cleanly in SystemVerilog - disable fork, process::kill & the Isolation Fork

Reset asserts in the middle of a burst. The DUT drops everything, state machines snap to idle — and your driver keeps going. Its drive_transaction() task is forty lines into a wait-for-ready loop that will never see ready again, cheerfully wiggling address lines at a design that is no longer listening. Twelve cycles later reset deasserts, the DUT wakes up, and the first thing it sees is the stale back half of a transaction from a previous epoch. The checker melts down. The bug report says "DUT misbehaves after reset." The DUT is fine. Your thread outlived the world it was born into.

Killing threads cleanly is not an exotic skill — it is the difference between a testbench that survives reset testing and one that fails in ways that smear blame onto the design. SystemVerilog gives you three mechanisms (disable fork, process::kill(), disable label); this post covers all three, the isolation idiom that makes the first one safe, the honest story about the famous #0 — and then the deep end: random stability, kill-safe cleanup, and hung-thread forensics.

Note Originally published in 2018; rewritten in 2026. The original recommended a bare #0 after fork/join_none as "critical" — this version explains the race honestly and gives a handshake that doesn't depend on scheduler ordering. It also adds the two patterns real testbenches use most: the reset-aware driver loop and the isolation fork.
~14 min read · Intermediate body, Advanced tail · Pairs with UVM Heartbeat - Detecting Hung Components.

The Pattern You Actually Need: Watchdog with join_any + disable fork

Most thread-killing in real testbenches isn't "kill thread X on demand" — it's "run this until that happens." Reset-aware driving is the canonical case, and the canonical shape is a race between the work and the event that invalidates it:

// Reset-aware driver loop — the pattern behind every robust UVM driver
task run_phase(uvm_phase phase);
  forever begin
    @(posedge vif.rst_done);            // wait until out of reset
    fork
      begin : work
        forever begin
          seq_item_port.get_next_item(req);
          drive_transaction(req);        // may be interrupted mid-burst
          seq_item_port.item_done();
        end
      end
      begin : reset_watcher
        @(negedge vif.rst_n);            // reset! the work is now meaningless
      end
    join_any
    disable fork;                        // kill whichever branch lost the race
    cleanup_bus();                       // drive idle values — see Common Mistakes
  end
endtask

join_any resumes the parent when the first branch finishes; disable fork then kills every remaining child. The driver thread can be forty lines deep in a handshake — it dies instantly, the bus gets cleaned, and the outer forever waits for the next reset epoch. No zombie threads, no stale transactions crossing reset.

The Isolation Fork: Making disable fork Safe

disable fork has a blast radius: it kills all child processes of the current process — not just the two you raced. If your task previously spawned a background monitor thread with join_none, disable fork kills that too. The 2018 version of this post called the statement "risky" and moved on; the actual fix is a standard idiom — wrap the race in its own parent so the kill can't reach anything else:

task drive_with_timeout(input int unsigned timeout_cycles);
  fork begin                    // ISOLATION fork: one child = one new parent scope
    fork
      drive_transaction(req);                       // the work
      repeat (timeout_cycles) @(posedge vif.clk);   // the timeout
    join_any
    disable fork;               // kills only the loser INSIDE the isolation scope
  end join                      // isolation fork completes as a unit
endtask
flowchart TD
    P["Parent process
(has other children: monitor, heartbeat)"] --> ISO["Isolation fork
(new parent scope)"] ISO --> W["work branch"] ISO --> T["timeout branch"] P --> M["monitor thread
(SAFE — outside blast radius)"] ISO -.->|"disable fork
kills only these"| W style ISO fill:#fef3c7,stroke:#d97706 style W fill:#fee2e2,stroke:#ef4444 style T fill:#fee2e2,stroke:#ef4444 style M fill:#d1fae5,stroke:#059669
Key fork begin ... end join around a join_any / disable fork pair is not decoration — it is a firewall. The inner disable fork executes in the isolation fork's process, so its kill radius stops at the isolation boundary. Adopt it as a reflex: any disable fork in a task that could ever be called from a thread-owning context needs the wrapper.

Targeted Kills: the process Class

When you genuinely need to kill one specific thread from far away — a stimulus thread owned by another component, a per-channel worker in a pool — capture its process handle:

process worker[$];

task start_workers(int n);
  for (int i = 0; i < n; i++)
    fork
      automatic int idx = i;
      begin
        worker.push_back(process::self());   // FIRST statement: publish the handle
        run_channel(idx);
      end
    join_none
  wait (worker.size() == n);                 // handshake — not #0; see below
endtask

task stop_worker(int idx);
  if (worker[idx] != null && worker[idx].status() != process::FINISHED)
    worker[idx].kill();                      // kills the thread AND its children
endtask

The process API also gives you introspection — the basis of the forensics section later:

stateDiagram-v2
    [*] --> RUNNING : fork spawns
    RUNNING --> WAITING : blocking statement
    WAITING --> RUNNING : event fires
    RUNNING --> SUSPENDED : suspend()
    SUSPENDED --> RUNNING : resume()
    RUNNING --> FINISHED : completes
    WAITING --> KILLED : kill() / disable
    RUNNING --> KILLED : kill() / disable
    FINISHED --> [*]
    KILLED --> [*]

The #0 Confession

The 2018 version of this post — like most tutorials — told you a bare #0 after fork/join_none is "critical." Time to be honest about what that actually does. The race is real: join_none children don't start until the parent blocks, so the parent can reach code that reads thread_handle before the child's process::self() assignment ran. A #0 blocks the parent for one iteration of the scheduler's inactive region, the child runs, the handle is valid. It works.

It also encodes an assumption — "one round of zero-delay rescheduling is enough" — that quietly breaks when the code composes: nest the pattern inside another zero-delay construct, add a second #0-dependent handshake, or let a class library reorder its own zero-delay events, and you're debugging a null handle that only appears in the full environment. The robust version costs one line and assumes nothing about scheduling:

process h = null;
fork
  begin
    h = process::self();      // publish before ANY blocking statement
    do_work();
  end
join_none
wait (h != null);             // explicit handshake: sleeps until the fact is true
Tip wait (h != null) states the condition you need; #0 states a guess about when the scheduler will provide it. Prefer conditions to guesses, here and everywhere — the same instinct that replaces #100 settling delays with event waits. If you must synchronize on many spawns, count them (wait (workers.size() == n)) rather than stacking #0s.

disable label: the Legacy Option and Its Trap

fork
  begin : thread_A
    forever do_work_a();
  end
  begin : thread_B
    forever do_work_b();
  end
join_none

disable thread_A;    // thread_B keeps running

Targeted, readable — and carrying a hazard the other methods don't have: disable works on the named block, not on a thread. If that block is in a task with several concurrent activations (a reused driver task, a recursive call, multiple agent instances of the same module), disable my_label kills every active instance of the block, everywhere. In class-based, multi-instance testbenches that makes it close to unusable; treat it as a module-land tool and reach for process handles or isolation forks in class land.

Comparison

MethodKillsSafe in class-based TBsBest for
join_any + disable fork (isolated)Losers of the race, within the firewallYes — with the isolation wrapperReset handling, timeouts, watchdogs — 90% of real uses
process::kill()One specific thread + its childrenYesWorker pools, cross-component control
disable labelEvery active instance of the blockRisky — multi-activation hazardSimple module-level threads

Common Mistakes

  • Bare disable fork in shared code. Kills sibling threads you forgot you had — monitor gone, heartbeat gone, an hour of confusing debug. Isolation-fork wrapper, always.
  • Killing a thread that holds a lock. kill() does not unwind: a killed thread never executes its semaphore.put() / mailbox handshake / item_done(). The next thread to request that resource hangs forever — and the hang is three components away from the kill. Audit every kill point: what does this thread hold?
  • Killing mid-transaction and not cleaning the bus. The dead driver leaves whatever it last drove on the interface. The cleanup_bus() in the reset pattern isn't optional — drive idles explicitly after every kill.
  • Trusting #0 in composed code. Use a wait-condition handshake; guesses about scheduler rounds don't survive integration.
  • disable label on a multiply-activated task block. All instances die, not "the" instance. If you can't prove single activation, don't use it.
  • Forgetting that kill() is recursive. Children of the killed thread die too — usually what you want, occasionally a surprise when a "logger" child was expected to flush.

Interview Corner

Q: disable fork vs process::kill() — when is each the right tool?

A: disable fork is structural — it kills all children of the current process, which is exactly right for "the race is decided, everyone else stop" patterns (reset, timeout), and exactly wrong near unrelated sibling threads unless wrapped in an isolation fork. process::kill() is targeted — one handle, one thread (plus its children), killable from anywhere that holds the handle. Structure for races, handles for reach.

Q: What is an isolation fork and why does it exist?

A: fork begin ... end join wrapped around a join_any/disable fork pair. It creates an intermediate parent process, so the inner disable fork's "kill all my children" stops at the wrapper instead of reaching threads the enclosing task spawned earlier. It exists because disable fork's radius is defined by process ancestry, not by lexical proximity — the wrapper realigns the two.

Q: A regression hangs at time 80us. You suspect a killed thread. What's the connection?

A: Killed threads don't release what they hold. If the victim held a semaphore, a mailbox slot, or a sequencer grant (item_done() never called), the next requester blocks forever — at a place far from the kill. The forensic move: find what resources the killed code path acquires, then check which of them was in the acquired state at the hang.

Q: Why does adding a harmless fork/join_none debug thread change your random stimulus?

A: Every thread gets its own RNG seeded from its parent's at spawn time — so thread creation order and count participate in seeding. A new fork shifts the seeding sequence of every thread spawned after it, and the "same seed" now produces different stimulus. That's the random-stability problem; see the Advanced section for the containment rules.

Beyond the Basics: Advanced → Expert

Level 1 — wait fork: the forgotten fourth verb

Killing gets the attention, but its dual matters as much: wait fork blocks until all children of the current process finish — the graceful end-of-test counterpart to disable fork's guillotine. The end-of-test idiom: stop feeding stimulus, wait fork for in-flight threads to drain, then assert done. Testbenches that kill everything at end-of-test instead of draining are the ones whose "passing" final packets were never actually checked. (In UVM, objections play this role at the phase level — but inside a component, wait fork is still the tool.)

Level 2 — Random stability: kills, forks, and the seeds between them

SystemVerilog's random stability rules seed each thread's RNG from its parent at spawn, in spawn order. Consequences worth engineering around: adding, removing, or reordering forks — including debug-only threads and the reset-recovery re-forks in this post's main pattern — reshuffles downstream seeding, so "rerun with the same seed" reproduces nothing. Containment: give long-lived threads their own explicit seeds (process::self().srandom(stable_seed) as the thread's first act), derive stable_seed from stable identifiers (component path, channel index) rather than spawn order, and keep debug forks after the stimulus-relevant ones. A reset-heavy test that re-forks its driver loop per epoch should reseed per epoch from an epoch counter — otherwise the number of resets changes all post-reset stimulus.

Level 3 — Kill-safe code: designing threads that die well

SV has no finally, so cleanup after a kill is the killer's job — which means killable threads must be written to be cleaned up. The discipline: (1) narrow the kill window — acquire locks and start bus activity as late as possible, release as early as possible; (2) make cleanup idempotent and external — a cleanup_bus()/release_locks() that the parent calls after every disable fork, safe to run whether or not the victim got far; (3) for sequencer-facing drivers, guarantee the item_done() accounting in the parent's cleanup path if the child died between get_next_item and item_done — the alternative is a sequencer wedged forever; (4) prefer designs where the long-lived resource holder is never the killable thread — the watcher gets killed, the worker completes idempotent units. When you can restructure so that kills only ever hit stateless code, most of this checklist evaporates — that's the real goal.

Level 4 — How UVM itself plays this game

UVM's machinery is built from these exact primitives, and knowing where changes your debugging. Sequences: uvm_sequence_base captures its body's process handle, which is what sequence.kill() and sequencer.stop_sequences() use — and why a killed sequence exhibits every hazard in this post (held grants, no unwinding) if the driver side isn't reset-aware. Phasing: phase.jump() ends the phase's task processes non-gracefully — component run_phase code that holds resources across arbitrary await points has the same kill-safety obligations as any worker thread. The architecture that falls out for reset-capable agents: the monitor never dies (it observes reset as protocol, reporting through it), the driver dies freely (reset-aware loop, stateless transaction units), and the sequence layer is told, via stop_sequences() or a reset event the virtual sequence watches — three components, three different relationships with thread death, one coherent reset story.

Level 5 — Hung-thread forensics

The terminal skill: a regression that neither passes nor fails, just... continues. Timeout fires at 10ms; now what? Order of attack: (1) objection audit — +UVM_OBJECTION_TRACE or uvm_root's objection dump shows who never dropped; nine times out of ten the hang has a name at this step. (2) Heartbeat instrumentation — a uvm_heartbeat per agent converts "something hung" into "this component hung," turning a 10ms-timeout mystery into a targeted failure at first-missed-beat. (3) Process-status dump — keep your worker-pool process handles and print status() per thread on timeout: everything WAITING is a suspect, and the one waiting on a semaphore whose holder is KILLED is your smoking gun. (4) The scheduler's-eye view — simulator-specific thread viewers show every process's current source line at the hang. The pattern across all four: hangs are not mysterious; they are exactly one unmet condition, and the entire craft is instrumenting so the condition's name survives to the log.

Key Takeaways

  • The workhorse is join_any + disable fork inside an isolation fork — reset handling, timeouts, watchdogs. Bare disable fork in shared code is a sibling-killer.
  • process::kill() for targeted, cross-scope kills; disable label only where single activation is provable.
  • Replace #0 guesses with wait-condition handshakes — state the condition, not the scheduler round.
  • Kills don't unwind: released nothing, cleaned nothing. Kill-safe design + parent-side idempotent cleanup, or wedge a semaphore three components away.
  • Thread lifecycle touches random stability, end-of-test draining (wait fork), and hang forensics — the advanced ladder is the same primitives at bigger radius.
Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment