Verilog RAM Models: Single-Port, Dual-Port, and a Task-Based Testbench

Memory models are the first place a Verilog beginner meets three ideas at once: a bidirectional bus, an inferred latch, and the difference between "simulates" and "synthesizes". This post merges four short 2015 write-ups into one guide: a single-port asynchronous RAM, the task-based testbench that exercises it, a dual-port RAM with separate read and write clocks, and a 4:1 multiplexer built from a Verilog task. Every block of code below was re-run for this rewrite with Icarus Verilog 13 and linted with Verilator 5, and the one thing the original posts got wrong is called out and fixed in Section 3.

Key takeaway The asynchronous models in Sections 1 and 4 are simulation models. They are fine for learning how a memory behaves and for a quick testbench, but they infer latches and drive tristate buses that no FPGA block RAM can implement. Section 3 gives the synchronous version you should actually synthesize.

1. Single-port asynchronous RAM

A single-port RAM has one address and one data path, so at any moment it is either reading or writing. The 2015 model uses a bidirectional data port and two enables: wr_en for writes and o_en (output enable) for reads. Only one should be high at a time.

PortDirectionDescription
datainoutBidirectional data bus, [WIDTH-1:0]
addrinputAddress, [ADDR-1:0]
resetinputAsynchronous reset, active high, clears every location
wr_eninputWrite enable
o_eninputOutput enable (read)

// Module: sp_ram.v
// Parameterized single-port asynchronous RAM (simulation model)

module sp_ram #(
  parameter WIDTH = 8,
  parameter DEPTH = 16,
  parameter ADDR  = 4
)(
  inout  [WIDTH-1:0] data,
  input  [ADDR-1:0]  addr,
  input              reset,
  input              wr_en,
  input              o_en
);

  reg [WIDTH-1:0] mem [DEPTH-1:0];
  integer i;

  // Read: drive the bus only when o_en=1 and wr_en=0, otherwise release it
  assign data = (o_en && !wr_en) ? mem[addr] : {WIDTH{1'bz}};

  // Write: level-sensitive, no clock
  always @(reset, data, addr, wr_en, o_en) begin
    if (reset) begin
      for (i = 0; i < DEPTH; i = i + 1)
        mem[i] = 0;
    end
    else if (wr_en && !o_en) begin
      mem[addr] = data;
    end
  end

endmodule

The truth table is the whole contract. Reading and writing at once is undefined, which is exactly why the testbench in the next section never does it.

wr_eno_enOperation
00Idle, bus released (high impedance)
01Read mem[addr] onto the bus
10Write the bus into mem[addr]
11Invalid, never drive both

Notice the {WIDTH{1'bz}} replication. A single 1'bz would be zero-extended to the bus width by the ternary operator, which is not what you want when releasing a bus. Replicating z across every bit is the correct idiom.

2. A task-based testbench

The testbench wraps each bus operation in a task, so the test sequence reads like a script instead of a wall of signal assignments. This is the habit that scales: a directed test written as a list of write and read calls is easy to review and easy to extend.


module sp_ram_tb;

  parameter WIDTH = 8;
  parameter ADDR  = 4;
  parameter DEPTH = 16;

  wire [WIDTH-1:0] data;
  reg  [WIDTH-1:0] temp;
  reg              reset, wr_en, o_en;
  reg  [ADDR-1:0]  addr;
  integer i;

  sp_ram inst(data, addr, reset, wr_en, o_en);

  // The testbench side of the bidirectional bus: drive only during writes
  assign data = (wr_en && !o_en) ? temp : 'bz;

  task init_t;
  begin
    {addr, reset, wr_en, o_en} = 0;
  end
  endtask

  task delay;
  begin
    #10;
  end
  endtask

  task reset_t;
  begin
    reset = 1;
    delay;
    reset = 0;
  end
  endtask

  task read_t;
    input [ADDR-1:0] rd_addr;
  begin
    wr_en = 0;
    o_en  = 1;
    addr  = rd_addr;
  end
  endtask

  task write_t;
    input [WIDTH-1:0] data_in;
    input [ADDR-1:0]  wr_addr;
  begin
    o_en  = 0;
    wr_en = 1;
    addr  = wr_addr;
    temp  = data_in;
  end
  endtask

  initial $monitor("t=%0t rst=%b wr=%b oe=%b addr=%0d data=%h",
                   $time, reset, wr_en, o_en, addr, data);

  initial begin
    init_t;
    reset_t;

    // Write 50+i to address i
    for (i = 0; i < DEPTH; i = i + 1) begin
      write_t(50 + i, i);
      delay;
    end

    // Read back in reverse order
    for (i = 0; i < DEPTH; i = i + 1) begin
      read_t(DEPTH - 1 - i);
      delay;
    end
    $finish;
  end

endmodule

The second assign data is the part beginners miss. Two things drive the data wire: the RAM (during reads) and the testbench (during writes). Each side must release the bus when it is not its turn, so the testbench drives temp only when wr_en && !o_en, the exact mirror of the RAM's read condition. Get either side wrong and you see x on the bus from contention.

Running it with Icarus gives the expected trace: sixteen writes of 0x32 through 0x41, then sixteen reads that return them in reverse.


t=0   rst=1 wr=0 oe=0 addr=0  data=zz
t=10  rst=0 wr=1 oe=0 addr=0  data=32
t=20  rst=0 wr=1 oe=0 addr=1  data=33
...
t=170 rst=0 wr=0 oe=1 addr=15 data=41
t=180 rst=0 wr=0 oe=1 addr=14 data=40
...
t=320 rst=0 wr=0 oe=1 addr=0  data=32

iverilog -Wall -o sp sp_ram.v sp_ram_tb.v && vvp sp

3. The fix: why the asynchronous model does not synthesize

The original posts presented sp_ram as RTL. It is not, and a lint run says so immediately:


%Warning-LATCH: sp_ram.v:24:3: Latch inferred for signal 'sp_ram.mem'
  (not all control paths of combinational always assign a value)
%Warning-UNOPTFLAT: sp_ram.v:17:19: Signal unoptimizable:
  Circular combinational logic: 'sp_ram.mem'

Both warnings point at the same root cause. The write block is level-sensitive and only assigns mem on some paths, so synthesis must hold the old value on the others, which is the definition of a latch. Worse, mem feeds data through the read assign, and data is in the sensitivity list of the block that writes mem, so the memory is combinationally its own input. Add the tristate inout, which exists on chip pads but not inside FPGA fabric, and you have a model that only makes sense in a simulator.

The synthesizable shape is a synchronous RAM: one clock, a write enable, separate write and read data, and a registered read. This is the template that Xilinx, Intel and Lattice tools recognize and map onto block RAM.


// Module: sp_ram_sync.v
// Synchronous single-port RAM, read-first, synthesizable

module sp_ram_sync #(
  parameter WIDTH = 8,
  parameter DEPTH = 16,
  parameter ADDR  = 4
)(
  input                  clk,
  input                  we,
  input  [ADDR-1:0]      addr,
  input  [WIDTH-1:0]     wdata,
  output reg [WIDTH-1:0] rdata
);

  reg [WIDTH-1:0] mem [0:DEPTH-1];

  always @(posedge clk) begin
    if (we)
      mem[addr] <= wdata;
    rdata <= mem[addr];   // read-first: a write cycle returns the old contents
  end

endmodule

Three things changed and each one matters:

  • No reset on the array. Block RAMs have no reset input. A for loop that clears every location on reset forces the tool to build the memory out of flip-flops instead. Initialize with an initial block and $readmemh if you need known contents, or write zeros with a small state machine after reset.
  • Registered read. rdata is valid one clock after addr, which is what real memories do. Testbenches must account for the latency, as the read_check task below does.
  • Read-during-write is defined. With the read assignment after the write, a cycle that writes and reads the same address returns the old data. Put the read before the write, or bypass wdata, and you get write-first behaviour. Pick one on purpose.

module sp_ram_sync_tb;
  parameter WIDTH = 8, DEPTH = 16, ADDR = 4;
  reg clk = 0, we = 0;
  reg [ADDR-1:0] addr = 0;
  reg [WIDTH-1:0] wdata = 0;
  wire [WIDTH-1:0] rdata;
  integer i, errors = 0;

  sp_ram_sync #(WIDTH, DEPTH, ADDR) dut(.clk(clk), .we(we), .addr(addr),
                                        .wdata(wdata), .rdata(rdata));
  always #5 clk = ~clk;

  task write(input [ADDR-1:0] a, input [WIDTH-1:0] d);
    begin @(negedge clk); we = 1; addr = a; wdata = d; @(negedge clk); we = 0; end
  endtask

  task read_check(input [ADDR-1:0] a, input [WIDTH-1:0] exp);
    begin
      @(negedge clk); addr = a;
      @(negedge clk);              // registered read: valid one clock later
      if (rdata !== exp) begin
        errors = errors + 1;
        $display("FAIL addr=%0d exp=%h got=%h", a, exp, rdata);
      end
    end
  endtask

  initial begin
    for (i = 0; i < DEPTH; i = i + 1) write(i, 8'd50 + i);
    for (i = DEPTH-1; i >= 0; i = i - 1) read_check(i, 8'd50 + i);
    if (errors == 0) $display("PASS: %0d locations verified", DEPTH);
    $finish;
  end
endmodule

This one is self-checking, which the 2015 testbench was not. It prints PASS: 16 locations verified under Icarus, and Verilator lints the RAM with no warnings.

4. Dual-port RAM with separate read and write clocks

A dual-port RAM has two independent address and data paths, so two agents can access it at once. The 2015 model goes one step further and clocks writes with wr_clk and reads with rd_clk, which is the shape of a clock-domain-crossing buffer.

PortDirectionDescription
wr_clk, rd_clkinputWrite and read clocks, may differ in frequency
resetinputAsynchronous reset, clears the array
data_0, data_1inoutBidirectional data per port
addr_0, addr_1inputAddress per port
wr_en_0, wr_en_1inputWrite enable per port
o_en_0, o_en_1inputOutput enable per port

// Module: dp_async_ram.v
// Dual-port RAM, writes on wr_clk, reads on rd_clk (simulation model)

module dp_async_ram #(
  parameter DEPTH = 16,
  parameter WIDTH = 8,
  parameter ADDR  = 4
)(
  input              wr_clk,
  input              rd_clk,
  input              reset,
  inout  [WIDTH-1:0] data_0, data_1,
  input  [ADDR-1:0]  addr_0, addr_1,
  input              wr_en_0, wr_en_1,
  input              o_en_0, o_en_1
);

  reg [WIDTH-1:0] data_0_reg, data_1_reg;
  integer i;

  reg [WIDTH-1:0] mem [DEPTH-1:0];

  // Writes: synchronous to wr_clk, asynchronous reset
  always @(posedge wr_clk or posedge reset) begin
    if (reset) begin
      for (i = 0; i < DEPTH; i = i + 1)
        mem[i] <= 0;
    end
    else begin
      if (wr_en_0 && !o_en_0)
        mem[addr_0] <= data_0;
      if (wr_en_1 && !o_en_1)
        mem[addr_1] <= data_1;
    end
  end

  // Reads: registered on rd_clk, driven onto the bus only when enabled
  assign data_0 = (o_en_0 && !wr_en_0) ? data_0_reg : {WIDTH{1'bz}};
  assign data_1 = (o_en_1 && !wr_en_1) ? data_1_reg : {WIDTH{1'bz}};

  always @(posedge rd_clk) begin
    if (o_en_0 && !wr_en_0)
      data_0_reg <= mem[addr_0];
    if (o_en_1 && !wr_en_1)
      data_1_reg <= mem[addr_1];
  end

endmodule
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#e0f2fe', 'primaryTextColor': '#0f172a', 'primaryBorderColor': '#0066cc', 'lineColor': '#475569', 'secondaryColor': '#f8fafc'}}}%%
flowchart LR
    subgraph Port0["Port 0"]
        D0[data_0]
        A0[addr_0]
        WE0[wr_en_0]
        OE0[o_en_0]
    end

    subgraph Port1["Port 1"]
        D1[data_1]
        A1[addr_1]
        WE1[wr_en_1]
        OE1[o_en_1]
    end

    MEM{{Memory Array}}

    D0 <--> MEM
    D1 <--> MEM
    A0 --> MEM
    A1 --> MEM

    style MEM fill:#d1fae5,stroke:#10b981,stroke-width:2px

Two hazards live in this design, and a verification engineer should be able to name both:

  • Write collision. If both ports write the same address in the same wr_clk cycle, the second if wins in simulation because nonblocking assignments to the same element resolve in source order. Silicon makes no such promise. Either constrain the stimulus so it never happens, or add an assertion that flags it: assert property (@(posedge wr_clk) !(wr_en_0 && wr_en_1 && addr_0 == addr_1));
  • Clock domain crossing. The array is written in the wr_clk domain and read in the rd_clk domain with no synchronization. A read that lands while the same location is being written can sample a half-updated word. The real answer is a FIFO with Gray-coded pointers, which is exactly what the asynchronous FIFO post builds on top of a dual-port memory like this one.

The reset loop and the inout ports carry the same synthesis caveats as Section 3. For hardware, use the synchronous template with two clocked ports and let the vendor's true-dual-port block RAM absorb the cross-domain access rules.

5. Bonus: a Verilog task inside RTL

Tasks are usually thought of as testbench tools, but a task with no timing controls, called from a combinational always block, is legal synthesizable Verilog. It is a way to name a block of logic that is reused several times inside one module. Here is the 2015 example, a 4:1 multiplexer, with one fix: the task's output argument was also named y, which shadowed the module port and drew a VARHIDDEN warning from Verilator.


module mux_t(
  input  [3:0] in,
  input  [1:0] sel,
  output reg   y
);

  task t_mux;
    input  [3:0] i;
    input  [1:0] s;
    output       o;
  begin
    case (s)
      2'b00: o = i[0];
      2'b01: o = i[1];
      2'b10: o = i[2];
      2'b11: o = i[3];
    endcase
  end
  endtask

  always @(*) begin
    t_mux(in, sel, y);
  end

endmodule

Rules for a task that must synthesize:

  • No #, @ or wait inside the task. Tasks with timing controls are simulation only.
  • Assign every output on every path, otherwise the latch warning from Section 3 comes back.
  • Prefer a function when there is a single return value. Functions cannot contain timing controls at all, so the synthesizability question never arises.

Key takeaways

  • A bidirectional bus needs both sides to release it. The RAM drives on read, the testbench drives on write, and the conditions must be exact mirrors.
  • Level-sensitive writes and reset loops over a memory array produce latches and flip-flop memories. Use the synchronous, registered-read template for anything that will be synthesized.
  • Self-checking testbenches beat waveform inspection. The 2015 testbench only printed values; the rewrite compares them.
  • Dual-port memories with two clocks are a clock-domain crossing. Constrain write collisions with an assertion and use a proper asynchronous FIFO for data transfer.
  • Tasks are fine in RTL when they are purely combinational, but a function is usually the clearer choice.

Verified with

Icarus Verilog 13.0 and Verilator 5.052 on macOS. Commands:


iverilog -Wall -o sp   sp_ram.v      sp_ram_tb.v      && vvp sp
iverilog -Wall -o sync sp_ram_sync.v sp_ram_sync_tb.v && vvp sync
verilator --lint-only -Wall sp_ram_sync.v
verilator --lint-only -Wall mux_t.v
Author
Mayur Kubavat
DV engineer working on SoC verification. Writes here about UVM, PCIe, SystemVerilog, and the everyday craft of getting designs to tape-out.

Comments (0)

Leave a Comment