Skip to content

FMA: why the addend alignment has two extra bits (#970), and how one could be removed - #1919

Closed
davidharrishmc wants to merge 1 commit into
openhwfoundation:mainfrom
davidharrishmc:dh/fma-970-save-bit
Closed

davidharrishmc wants to merge 1 commit into
openhwfoundation:mainfrom
davidharrishmc:dh/fma-970-save-bit

Conversation

@davidharrishmc

@davidharrishmc davidharrishmc commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Explains #970. Closed without merging; the RTL is unchanged.

PR #793 widened the FMA's aligned addend and adder by two bits, one above the product and one below it, to fix #578. Issue #970 asked why each bit is needed. This PR records the answer.

  • The bit above the product is required.
  • The bit below the product is required only for the smallest subnormal multiplicand.
  • That bit can be removed by remapping that one input value; this branch shows how. Synthesis shows the area saving is below synthesis noise, while the remap adds logic in front of the multiplier and the alignment count. So the RTL keeps FMALEN = 3Nf+6.

Datapath layout

The multiplier produces the exact product Pm = Xm·Ym in U(2.2Nf) (two integer bits, because the product is in [1,4)), with exponent Pe. The alignment shifter places the addend Zm relative to the product, and the adder (Am, Sm, FMALEN bits) holds, from the top:

Field Wally (kept) With the remap (this branch)
Addend Zm when it is not shifted Nf+1 Nf+1
Gap between the addend and the product 1 1
Exact product 2Nf+2 2Nf+2
Bits below the product 2 1
FMALEN 3Nf+6 3Nf+5

Below the adder, the alignment shifter keeps Nf more bits that only feed the sticky bit. Those bits are unchanged.

Why the bit above the product is needed (#578)

The product is killed, meaning it only contributes to the sticky bit, when the addend is so far above it that it cannot reach the result's guard bit. Because the product can be almost 4·2^Pe, this is safe only when Ze ≥ Pe+Nf+4. Before #793 the product was killed one binade earlier, at Ze ≥ Pe+Nf+3. Most of the time that works, because the result keeps the addend's exponent and its guard bit is above any product bit.

The exception is an effective subtraction with Zm = 1.0 exactly. Then Z−P falls just below 2^Ze, the result loses its leading bit, and its guard bit moves down to 2^(Ze−Nf−2) = 2^(Pe+1). A product in [2,4)·2^Pe reaches that bit. This is the #578 vector:

bcafffffe7ffffff × 3fdfffffffffffff + 3ff0000000000000 = 3fefffffffffffff   (RNE)

Killing the product gives 3ff0000000000000. With the kill moved to Ze ≥ Pe+Nf+4, an addend at Ze = Pe+Nf+3 must stay in the adder. Its lowest bit then sits one position above the product's top bit, and that gap bit becomes the result's last bit in the cancelling case. Removing it makes TestFloat fail on the vector above.

Why the bit below the product was needed

Wally does not normalize subnormal multiplicands. A subnormal X is unpacked as 0.f with biased exponent 1. If Xm has k leading zeros, the product's leading one is up to k positions below 2^Pe, but its lowest bit is still at 2^(Pe−2Nf). Suppose an effective subtraction then cancels one more bit. The result's last bit is at 2^(Pe−k−Nf−1), and its guard bit is at 2^(Pe−k−Nf−2). The addend bits at those positions must be added exactly rather than folded into the sticky bit. So the adder needs k−Nf+2 bits below the product. A product of normal multiplicands (k = 0) needs none.

Only the smallest subnormal, with fraction 1, has k = Nf, and it needs both bits. Every other subnormal has k ≤ Nf−1 and needs at most one. With one bit, the smallest subnormal goes wrong. The addend's last bit lands in the sticky bit, and when the remainder is exactly half an ulp, the datapath sees "less than half" and rounds the wrong way. One example, constructed directly from this analysis:

0000000000000001 × 4350000000000000 + 801ffffffffffffd = 0020000000000002   (RNE)

With one bit, this returns 0020000000000001. TestFloat finds the same class of error in f64 and in f128, each time with X equal to the smallest subnormal, for example 0000000000000001 c3d00001fffffffd 00148a95c0ea9ca0 gives 80aff5beb51f8aab instead of 80aff5beb51f8aac.

The #578 vectors do not need this bit; the bit above the product alone fixes them. So #793's second bit was correct, but for a different case than the one it was added for.

How the bit could be removed (this branch; not merged)

In fma.sv, a multiplicand equal to the smallest subnormal (significand 2^−Nf, biased exponent 1) is passed to the multiplier and the exponent logic as significand 2^−(Nf−1) with biased exponent 0. The value is the same, but the product now has at most Nf−1 leading zeros, so one bit below the product is enough. For each of X and Y:

XMinSub = ~Xm[NF] & ~|Xm[NF-1:1] & Xm[0];               // smallest subnormal
XmM     = {Xm[NF:2], Xm[1] | XMinSub, Xm[0] & ~XMinSub}; // move fraction bit 0 to bit 1
XeM     = {Xe[NE-1:1], Xe[0] & ~XMinSub};                // biased exponent 1 -> 0

The unpacker always gives a subnormal biased exponent bit 0 = 1, including subnormals of the smaller formats converted to the internal bias, so clearing that bit subtracts 1 in every format. In the smaller formats the fraction sits at the top of the Nf-bit field and bit 0 is 0, so the remap never fires there; it isn't needed because those formats have slack below their product. Z is not remapped.

The remaining edits follow from the narrower adder:

  • FMALEN becomes 3Nf+5.
  • The product's padding in the adder and LZA becomes one bit instead of two.
  • The killed-product placement becomes {Nf+3 zeros, Zm, 2Nf+1 zeros}.
  • The killed-addend threshold drops from ACnt > 3Nf+5 to ACnt > 3Nf+4, so the shifter still holds every addend bit until the addend can only affect the sticky bit.

The sum's exponent and the normalization shift amounts are measured from the top of the sum, so they don't change.

Tradeoffs

The saving is one bit of a datapath about 160 bits wide (161 for double), in each of these places:

  • the alignment shifter output;
  • the addend inversion;
  • both carry-propagate adders (positive and negative sum);
  • the leading-zero anticipator;
  • the E/M pipeline register for Sm;
  • the normalization shifter input, wherever FMALEN+2 sets NORMSHIFTSZ.

The cost is two (Nf−1)-input zero-detects and six gates. They sit in front of the multiplier's two low input bits and the low bit of the alignment count ACnt, which feeds the alignment shifter's controls. So the remap adds a few gate delays at the start of a path that matters.

In sky130 at 200 MHz, the change is smaller than synthesis noise:

Synthesis main this branch this branch without the remap (incorrect, for reference)
FMA area 174,701 µm² 172,775 µm² 173,456 µm²
FMA critical-path arrival 5.53 ns 5.59 ns 5.49 ns
FPU area 668,194 µm² 667,913 µm² —

The multiplier is identical RTL in all three FMA runs, yet its area varied by up to 2,250 µm². By hand the net saving is a few tens of cells, which these runs can't resolve.

Alternatives:

  • Share the detection with the unpacker. It already ORs the whole fraction for FracZero, so exporting a smallest-subnormal flag costs about two gates per operand. That makes the remap nearly free, but adds a port through unpackinput, unpack, fpu, fma, and testbench_fp.
  • Shift every subnormal multiplicand left by one (biased exponent 0, significand {f, 0}). This needs no zero-detect, but costs Nf two-input muxes per operand, more than the bit saves.
  • Normalize subnormal multiplicands completely. This would remove both bits below the product and most of the sticky-only shifter bits, but needs a leading-zero counter and a shifter in front of the multiplier.
  • Keep 3Nf+6 and only document the reasoning above.

Decision: keep 3Nf+6. The saving is too small to justify extra logic on a timing-relevant path, or new ports through five files.

Testing of the remap in this branch

  • Directed vectors: 82,500 f64 vectors that target the gap bit and the bits below the product, with expected results from an exact rational-arithmetic reference that agrees with TestFloat. They pass; removing any bit makes them fail.
  • TestFloat: add, sub, mul and fma pass on fd_rv64gc and fdqh_rv64gc, covering all rounding modes in half, single, double and quad.
  • Coverfloat: the 220 coverfloat D-FMA tests from Integrate Cover-Float Tests in ACT4 riscv/riscv-arch-test#1528 pass. They do not detect the removal of any of these bits.
  • Regression and lint: regression-wally and lint-wally --nightly pass, apart from failures unrelated to the FPU.

🤖 Generated with Claude Code

The second bit below the product (added in openhwfoundation#793) is needed only when a
multiplicand is the smallest subnormal, which gives the product NF leading
zeros so the guard bit falls two bits below the product.  Presenting that
value as 2^-(NF-1) with biased exponent 0 caps the product at NF-1 leading
zeros, so one bit below the product is enough and FMALEN drops from 3NF+6 to
3NF+5.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: David Harris <David_Harris@hmc.edu>
@davidharrishmc davidharrishmc changed the title FMA: remove one adder bit below the product FMA: why the addend alignment has two extra bits (#970), and how one could be removed Oct 3, 2026
@davidharrishmc

Copy link
Copy Markdown
Contributor Author

Closing without merging: the RTL stays at FMALEN = 3Nf+6. The description records why both extra bits from #793 exist and how the one below the product could be removed, for reference.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

TestFloat fails on fma tests

1 participant