FMA: why the addend alignment has two extra bits (#970), and how one could be removed - #1919
Closed
davidharrishmc wants to merge 1 commit into
Closed
davidharrishmc wants to merge 1 commit into
davidharrishmc wants to merge 1 commit into
Conversation
The second bit below the product (added in openhwfoundation#793) is needed only when a multiplicand is the smallest subnormal, which gives the product NF leading zeros so the guard bit falls two bits below the product. Presenting that value as 2^-(NF-1) with biased exponent 0 caps the product at NF-1 leading zeros, so one bit below the product is enough and FMALEN drops from 3NF+6 to 3NF+5. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: David Harris <David_Harris@hmc.edu>
Contributor
Author
|
Closing without merging: the RTL stays at FMALEN = 3Nf+6. The description records why both extra bits from #793 exist and how the one below the product could be removed, for reference. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Explains #970. Closed without merging; the RTL is unchanged.
PR #793 widened the FMA's aligned addend and adder by two bits, one above the product and one below it, to fix #578. Issue #970 asked why each bit is needed. This PR records the answer.
Datapath layout
The multiplier produces the exact product Pm = Xm·Ym in U(2.2Nf) (two integer bits, because the product is in [1,4)), with exponent Pe. The alignment shifter places the addend Zm relative to the product, and the adder (
Am,Sm,FMALENbits) holds, from the top:Below the adder, the alignment shifter keeps Nf more bits that only feed the sticky bit. Those bits are unchanged.
Why the bit above the product is needed (#578)
The product is killed, meaning it only contributes to the sticky bit, when the addend is so far above it that it cannot reach the result's guard bit. Because the product can be almost 4·2^Pe, this is safe only when Ze ≥ Pe+Nf+4. Before #793 the product was killed one binade earlier, at Ze ≥ Pe+Nf+3. Most of the time that works, because the result keeps the addend's exponent and its guard bit is above any product bit.
The exception is an effective subtraction with Zm = 1.0 exactly. Then Z−P falls just below 2^Ze, the result loses its leading bit, and its guard bit moves down to 2^(Ze−Nf−2) = 2^(Pe+1). A product in [2,4)·2^Pe reaches that bit. This is the #578 vector:
Killing the product gives 3ff0000000000000. With the kill moved to Ze ≥ Pe+Nf+4, an addend at Ze = Pe+Nf+3 must stay in the adder. Its lowest bit then sits one position above the product's top bit, and that gap bit becomes the result's last bit in the cancelling case. Removing it makes TestFloat fail on the vector above.
Why the bit below the product was needed
Wally does not normalize subnormal multiplicands. A subnormal X is unpacked as 0.f with biased exponent 1. If Xm has k leading zeros, the product's leading one is up to k positions below 2^Pe, but its lowest bit is still at 2^(Pe−2Nf). Suppose an effective subtraction then cancels one more bit. The result's last bit is at 2^(Pe−k−Nf−1), and its guard bit is at 2^(Pe−k−Nf−2). The addend bits at those positions must be added exactly rather than folded into the sticky bit. So the adder needs k−Nf+2 bits below the product. A product of normal multiplicands (k = 0) needs none.
Only the smallest subnormal, with fraction 1, has k = Nf, and it needs both bits. Every other subnormal has k ≤ Nf−1 and needs at most one. With one bit, the smallest subnormal goes wrong. The addend's last bit lands in the sticky bit, and when the remainder is exactly half an ulp, the datapath sees "less than half" and rounds the wrong way. One example, constructed directly from this analysis:
With one bit, this returns 0020000000000001. TestFloat finds the same class of error in f64 and in f128, each time with X equal to the smallest subnormal, for example
0000000000000001 c3d00001fffffffd 00148a95c0ea9ca0gives80aff5beb51f8aabinstead of80aff5beb51f8aac.The #578 vectors do not need this bit; the bit above the product alone fixes them. So #793's second bit was correct, but for a different case than the one it was added for.
How the bit could be removed (this branch; not merged)
In
fma.sv, a multiplicand equal to the smallest subnormal (significand 2^−Nf, biased exponent 1) is passed to the multiplier and the exponent logic as significand 2^−(Nf−1) with biased exponent 0. The value is the same, but the product now has at most Nf−1 leading zeros, so one bit below the product is enough. For each of X and Y:The unpacker always gives a subnormal biased exponent bit 0 = 1, including subnormals of the smaller formats converted to the internal bias, so clearing that bit subtracts 1 in every format. In the smaller formats the fraction sits at the top of the Nf-bit field and bit 0 is 0, so the remap never fires there; it isn't needed because those formats have slack below their product. Z is not remapped.
The remaining edits follow from the narrower adder:
FMALENbecomes 3Nf+5.{Nf+3 zeros, Zm, 2Nf+1 zeros}.The sum's exponent and the normalization shift amounts are measured from the top of the sum, so they don't change.
Tradeoffs
The saving is one bit of a datapath about 160 bits wide (161 for double), in each of these places:
Sm;FMALEN+2setsNORMSHIFTSZ.The cost is two (Nf−1)-input zero-detects and six gates. They sit in front of the multiplier's two low input bits and the low bit of the alignment count
ACnt, which feeds the alignment shifter's controls. So the remap adds a few gate delays at the start of a path that matters.In sky130 at 200 MHz, the change is smaller than synthesis noise:
The multiplier is identical RTL in all three FMA runs, yet its area varied by up to 2,250 µm². By hand the net saving is a few tens of cells, which these runs can't resolve.
Alternatives:
FracZero, so exporting a smallest-subnormal flag costs about two gates per operand. That makes the remap nearly free, but adds a port throughunpackinput,unpack,fpu,fma, andtestbench_fp.{f, 0}). This needs no zero-detect, but costs Nf two-input muxes per operand, more than the bit saves.Decision: keep 3Nf+6. The saving is too small to justify extra logic on a timing-relevant path, or new ports through five files.
Testing of the remap in this branch
regression-wallyandlint-wally --nightlypass, apart from failures unrelated to the FPU.🤖 Generated with Claude Code