2. STA Fundamentals - Setup & Hold with Clock Skew
A chip comes back from the fab and goes on the shmoo plot. A setup problem has a signature everyone can live with: the part fails at 500 MHz, works at 450, ships with a slower speed grade. A hold problem has a different signature: the part fails at 500 MHz, fails at 100 MHz, fails at 1 MHz — fails at every frequency, because the race it loses happens inside a single clock edge, and no amount of slowing down changes it. One of these problems costs you a speed bin. The other costs you a metal respin.
This post — Part 2 of the STA Fundamentals series — is about why those two failures behave so differently, and how clock skew sits at the center of both. Along the way it fixes the mistake that most tutorials (including the 2015 version of this very post) make: using one combinational delay for both checks. Setup and hold are not the same check with different margins. They are two different races, run on two different paths.
What is Clock Skew?
The clock does not arrive everywhere at once. Between the clock source and each flip-flop sits a tree of buffers and wire, and no two branches of that tree are identical. The difference in arrival time between two flops is clock skew:
flowchart TD
CLK["Clock Source"] --> B1["Buffer"]
B1 --> B2["Buffer"]
B1 --> B3["Buffer"]
B2 --> B4["Buffer"]
B3 --> B5["Buffer"]
B4 --> FF1["FF1
(Launch)"]
B5 --> FF2["FF2
(Capture)"]
style CLK fill:#d1fae5,stroke:#10b981
style FF1 fill:#dbeafe,stroke:#3b82f6
style FF2 fill:#dbeafe,stroke:#3b82f6
style B1 fill:#f3f4f6,stroke:#9ca3af
style B2 fill:#f3f4f6,stroke:#9ca3af
style B3 fill:#f3f4f6,stroke:#9ca3af
style B4 fill:#f3f4f6,stroke:#9ca3af
style B5 fill:#f3f4f6,stroke:#9ca3af
Clock Skew = Tcapture − Tlaunch
where Tlaunch is the clock arrival time at the launching flop (FF1) and Tcapture at the capturing flop (FF2). Positive skew means the capture clock arrives later.
| Skew Type | Condition | Effect on Setup | Effect on Hold |
|---|---|---|---|
| Positive | Tcapture > Tlaunch | Helps (more time) | Hurts (less margin) |
| Negative | Tcapture < Tlaunch | Hurts (less time) | Helps (more margin) |
| Zero | Tcapture = Tlaunch | No effect | No effect |
Two Different Races
Here is the reframe that makes everything else in this post fall into place. Between FF1 and FF2, every clock edge starts two races:
| Setup check | Hold check | |
|---|---|---|
| The race | Data launched at edge N must arrive before capture edge N+1 | Data launched at edge N must not arrive until after capture edge N has safely closed |
| Edges involved | Two different edges — one clock period apart | The same edge at both flops |
| Path that matters | The slowest path (Tcomb,max) | The fastest path (Tcomb,min) |
| Clock period appears? | Yes — slowing the clock relaxes it | No — frequency-independent |
| Failure in silicon | Fails above some frequency → speed bin | Fails at every frequency → respin |
Tcomb into both equations — as the original version of this post did — is quietly assuming the longest and shortest paths are the same path, and will conclude that hold is never a problem. Silicon disagrees.Setup with Skew: the Next-Edge Race
Data leaves FF1 at its clock edge, takes Tc2q + Tcomb,max to arrive, and must beat FF2's next edge by Tsetup. Positive skew moves that next edge later — a gift of extra time:
{ "signal": [
{ "name": "CLK @ FF1", "wave": "p.....|p.....", "node": ".a" },
{ "name": "CLK @ FF2", "wave": "0.p...|.p....", "node": "..b....c" },
{ "name": "FF1.Q", "wave": "x..3..|......", "data": ["D launched"] },
{ "name": "FF2.D", "wave": "x....4|......", "data": ["D arrives"] }
], "edge": ["a->b +skew", "b<->c capture at next edge"], "head": { "text": "Positive Skew - Setup Analysis" }, "config": { "hscale": 1.5 } }
Setup Constraint
Tclk + Tskew ≥ Tc2q,max + Tcomb,max + Tsetup
Rearranged for the minimum clock period:
Tclk ≥ Tc2q,max + Tcomb,max + Tsetup − Tskew
Every term on the right is a max — worst case, slow corner, longest path. The check asks: even on the worst day, does the slowest data still make it?
Hold with Skew: the Same-Edge Race
The hold check involves no second clock edge. At edge N, FF2 is busy capturing the old data on its D input; that data must stay stable for Thold after FF2's edge. Meanwhile FF1 — clocked by the same edge N — has already launched new data toward FF2. If the new data takes the fastest path through the logic and arrives too soon, it corrupts the capture in progress:
{ "signal": [
{ "name": "CLK @ FF1 (launch)", "wave": "0.10.......", "node": ".a" },
{ "name": "CLK @ FF2 (capture)", "wave": "0...10.....", "node": "....b" },
{ "name": "FF2.D", "wave": "3.....4....", "data": ["D0 — being captured", "D1 — new"], "node": "......c" }
], "edge": ["a->b +skew delays capture edge", "b->c stable for at least Thold"], "head": { "text": "Hold - same edge, new data racing old capture" }, "config": { "hscale": 1.5 } }
Hold Constraint
Tc2q,min + Tcomb,min ≥ Thold + Tskew
Every delay term is a min — best case, fast corner, shortest path. The check asks: on the fastest day, does the new data still arrive late enough? Notice what's absent: Tclk appears nowhere. Both flops act on the same edge, so the clock period cancels out of the race entirely. That single algebraic fact is why no frequency change can ever fix a hold violation.
And notice what skew does here: positive skew pushes FF2's capture edge later, which extends the danger window the new data must survive. The same +Tskew that relaxed setup now sits on the requirement side of the hold inequality. Skew gives to one check exactly what it takes from the other:
flowchart LR
subgraph POSITIVE["Positive Skew"]
PS_SETUP["Setup: ✓ Helps"]
PS_HOLD["Hold: ✗ Hurts"]
end
subgraph NEGATIVE["Negative Skew"]
NS_SETUP["Setup: ✗ Hurts"]
NS_HOLD["Hold: ✓ Helps"]
end
style PS_SETUP fill:#d1fae5,stroke:#10b981
style PS_HOLD fill:#fee2e2,stroke:#ef4444
style NS_SETUP fill:#fee2e2,stroke:#ef4444
style NS_HOLD fill:#d1fae5,stroke:#10b981
Worked Example: How Fixing Setup Creates a Hold Violation
Take one register-to-register path and — this time — characterize it honestly, with both the longest and shortest routes through the logic cloud:
- Tc2q,max = 0.3 ns, Tc2q,min = 0.15 ns
- Tcomb,max = 2.0 ns (longest path), Tcomb,min = 0.4 ns (shortest path)
- Tsetup = 0.2 ns, Thold = 0.1 ns
Round 1 — zero skew.
// Setup (max path):
Tclk_min = 0.3 + 2.0 + 0.2 - 0 = 2.5 ns // Fmax = 400 MHz
// Hold (min path):
Hold_margin = 0.15 + 0.4 - 0 - 0.1 = +0.45 ns // comfortable
Round 2 — add +0.5 ns skew to chase frequency. The clock tree is tuned so FF2's clock arrives 0.5 ns late, buying the setup race half a nanosecond:
// Setup (max path):
Tclk_min = 0.3 + 2.0 + 0.2 - 0.5 = 2.0 ns // Fmax = 500 MHz — 25% faster!
// Hold (min path):
Hold_margin = 0.15 + 0.4 - 0.5 - 0.1 = -0.05 ns // VIOLATION
The path now closes setup at 500 MHz and fails hold at every frequency. This is not a contrived corner — it is the standard failure mode of skew-based setup fixing, and it's why place-and-route tools treat hold fixing as a dedicated step after clock tree synthesis. The repair is in the data path, not the clock: insert delay (a buffer, ~0.1 ns) on the short path only:
// After adding 0.1 ns delay buffer on the min path:
Hold_margin = 0.15 + 0.5 - 0.5 - 0.1 = +0.05 ns // met
// Setup unaffected: buffer sits on the short path, max path unchanged
flowchart LR
FF1["FF1"] --> BUF1["Delay
Buffer"] --> COMB["Short path
through cloud"] --> FF2["FF2"]
style FF1 fill:#dbeafe,stroke:#3b82f6
style FF2 fill:#dbeafe,stroke:#3b82f6
style BUF1 fill:#fef3c7,stroke:#f59e0b
style COMB fill:#f3f4f6,stroke:#9ca3af
Slack Summary
| Check | Slack Formula | Delays Used | Requirement |
|---|---|---|---|
| Setup | Tclk + Tskew − Tc2q,max − Tcomb,max − Tsetup | Max (slow corner) | ≥ 0 |
| Hold | Tc2q,min + Tcomb,min − Tskew − Thold | Min (fast corner) | ≥ 0 |
Fixing Violations
Setup Violations
- Reduce the long path: faster cells, logic restructuring, fewer levels
- Pipeline the path (add a register stage)
- Useful skew — deliberately delay the capture clock (see Advanced section)
- Last resort: lower the frequency — setup is the one check that lets you
Hold Violations
- Add delay on the short data path (buffers, weaker cells)
- Reduce the skew that created the problem
- Changing frequency does nothing — Tclk is not in the equation
Common Mistakes
- One Tcomb for both checks. The original mistake this rewrite exists to fix: analyze hold with the setup path's delay and hold margin looks enormous, forever. Hold lives on the shortest path at the fastest corner.
- Assuming the skew sign convention. "Positive skew helps setup" is only true under capture-minus-launch. Verify before applying to a tool report.
- Trying to waive a hold violation "because we'll run slower." Same-edge race; frequency-independent; the waiver is a respin with extra steps.
- Fixing hold before setup closure. Setup fixes (resizing, restructuring, useful skew) change both path delays and skew — hold buffers inserted too early get invalidated and re-done. Signoff flows fix hold last for a reason.
- Treating clock uncertainty as one number forever. Pre-CTS, uncertainty is a guess covering skew+jitter+margin; post-CTS, real propagated skew replaces part of it. Carrying the pre-CTS number through signoff double-counts pessimism.
Interview Corner
Q: Why can't a hold violation be fixed by slowing the clock?A: Because the hold check is a same-edge race — launch and capture happen on the same clock edge, so the clock period never enters the inequality: Tc2q,min + Tcomb,min ≥ Thold + Tskew. Changing Tclk changes nothing in that expression. The fix is adding delay to the short data path or reducing the skew.
Q: Which is worse in silicon — a setup or a hold violation — and why?A: Hold. A setup-limited part still works at reduced frequency, so it can ship in a slower bin. A hold violation fails at every frequency and can only be fixed by changing the physical design — a metal respin at best. This is also why hold checks get extra derating pessimism at signoff: the cost asymmetry is enormous.
Q: A path has +0.4 ns skew. What happens to its setup and hold checks?A: Under the capture-minus-launch convention, setup gains 0.4 ns of slack (the capture edge arrives later, extending the next-edge race) and hold loses 0.4 ns (the same-edge danger window extends by the same amount). Skew is zero-sum between the two checks on a given path.
Q: Which timing corners do setup and hold sign off at, and why aren't they the same?A: Setup is checked where paths are slowest (slow process, low voltage, worst temperature) because the question is "can the slowest data still make it?" Hold is checked where paths are fastest (fast process, high voltage) because the question is "can the fastest data arrive too early?" Modern signoff checks both at multiple corners with OCV derates, since launch and capture paths can sit at different points of the on-chip variation spread simultaneously.
Beyond the Basics: Advanced → Expert
The equations above are the whole story for one path with known numbers. Real timing closure is about where those numbers come from and how much you can trust them. Each rung below is one layer deeper into signoff reality.
Level 1 — Useful skew: turning the enemy into a tool
If skew shifts slack between setup and hold, you can spend it deliberately. Suppose a path fails setup by 0.3 ns while the next stage (FF2→FF3) has 0.8 ns of setup slack to spare. Delay FF2's clock by 0.3 ns: the failing FF1→FF2 path gains 0.3 ns of setup slack; the FF2→FF3 path gives up 0.3 of its 0.8 — everyone passes. That is useful skew: sculpting the clock tree so slack flows from rich paths to poor ones. The bill arrives at FF2's hold check, which just got 0.3 ns worse — useful-skew flows always re-run hold fixing behind them. In SDC/CTS this appears as set_clock_latency targets or CCOpt skew groups rather than a hand-placed buffer.
Level 2 — Decomposing clock uncertainty: why jitter mostly spares hold
The pre-CTS blanket set_clock_uncertainty lumps three physically different things: spatial skew (a fixed offset between branches), cycle-to-cycle jitter (edge N+1 wobbling relative to edge N), and margin. The decomposition matters because they hit the two checks differently. Setup spans two edges, so jitter — the wobble between edges — attacks it directly. Hold compares one edge against itself at two locations; the common-mode wobble largely cancels, and only the differential part (skew, plus a little uncorrelated tree noise) remains. This is why signoff scripts typically apply a smaller uncertainty to hold than to setup, and why "just add more uncertainty" as a safety blanket punishes setup closure without buying meaningful hold safety.
Level 3 — OCV and derating: launch and capture live on different chips
On-chip variation means two adjacent paths on the same die run at slightly different speeds. STA models this pessimistically: for a hold check, derate the launch path fast and the capture clock path slow simultaneously —
# Classic flat OCV derates (signoff hold check)
set_timing_derate -early 0.95 ;# data/launch path could be 5% faster
set_timing_derate -late 1.05 ;# capture clock path could be 5% slower
report_timing -delay_type min -derate
Flat derates are brutally pessimistic on deep clock trees, which led to AOCV (depth- and distance-aware derates) and then POCV/SOCV (statistical, per-cell sigma). One expert wrinkle worth knowing by name: when launch and capture clocks share part of the tree, derating the shared segment both fast and slow at once is physically impossible pessimism — CRPR (clock reconvergence pessimism removal) is the signoff feature that credits it back. If a hold report looks impossibly bad, the first question is whether CRPR is on.
Level 4 — GBA vs PBA: when the tool is lying to you (conservatively)
Default STA is graph-based (GBA): at every node the tool keeps the worst arrival across all paths through it, so a reported path can be a Frankenstein that no real signal ever traverses — worst slew from one path grafted onto worst arrival from another. Path-based analysis (PBA, report_timing -pba_mode path) re-times the specific path with its own slews and typically recovers real margin. The signoff trade: GBA is fast and safely pessimistic, PBA is expensive and accurate. Standard practice is GBA everywhere, then PBA on the violating tail before anyone starts ECOing cells — a surprising fraction of "violations" evaporate under PBA.
Level 5 — Bending the single-cycle rule: multicycle paths and latch borrowing
Everything above assumed data must cross in one cycle. Two escape hatches, both sharp-edged. set_multicycle_path 2 -setup tells the tool a path legitimately has two cycles — but the default hold edge moves with it, so the canonical idiom is paired: set_multicycle_path 2 -setup plus set_multicycle_path 1 -hold, returning the hold check to the launch edge. Forgetting the hold half is the most common SDC bug in real constraint decks, and it manifests as thousands of phantom hold buffers. The second hatch is level-sensitive latches: a latch is transparent for half the cycle, so late-arriving data can borrow time from the next phase without any constraint at all — the basis of high-performance time-borrowing pipelines, at the price of far murkier STA reports (borrowed time shows up as negative slack that isn't a violation).
Key Takeaways
- Setup and hold are different races on different paths: setup = next-edge race on the max path (slow corner); hold = same-edge race on the min path (fast corner). Never reuse one Tcomb across both.
- Skew is zero-sum: positive skew (capture later) buys setup slack and sells hold margin, one-for-one. Fixing setup with skew is how hold violations are manufactured.
- Tclk is absent from the hold equation — hold failures are frequency-independent, which is why they mean respins while setup failures mean speed bins.
- Beyond the basics, closure is a pessimism-management game: decomposed uncertainty, OCV derates with CRPR, PBA on the violating tail, and multicycle exceptions with their hold halves intact.
STA Commands (Synopsys PrimeTime)
# Setup analysis — max delays, slow corner
report_timing -delay_type max -path_type full
# Hold analysis — min delays, fast corner
report_timing -delay_type min -path_type full
# Skew visibility on the clock network
report_clock_timing -type skew
# Multicycle exception — note the mandatory hold half
set_multicycle_path 2 -setup -from [get_pins FF1/CK] -to [get_pins FF2/D]
set_multicycle_path 1 -hold -from [get_pins FF1/CK] -to [get_pins FF2/D]
Continue the series with Part 3: Reset Recovery & Removal Time — the same launch/capture reasoning applied to asynchronous reset deassertion. The 2-FF reset synchronizer it recommends is the one built in D Flip-Flop with Async Reset.
Comments (0)
Leave a Comment