Add float8_e4m3 #161

apivovarov · 2024-08-01T04:21:55Z

This PR adds f8E4M3 type.

f8E4M3 type follows IEEE 754 convention

f8E4M3 (IEEE 754)
- Exponent bias: 7
- Maximum stored exponent value: 14 (binary 1110)
- Maximum unbiased exponent value: 14 - 7 = 7
- Minimum stored exponent value: 1 (binary 0001)
- Minimum unbiased exponent value: 1 − 7 = −6
- Precision specifies the total number of bits used for the significand (mantisa), 
    including implicit leading integer bit = 3 + 1 = 4
- Follows IEEE 754 conventions for representation of special values
- Has Positive and Negative zero
- Has Positive and Negative infinity
- Has NaNs

Additional details:
- Max exp (unbiased): 7
- Min exp (unbiased): -6
- Infinities (+/-): S.1111.000
- Zeros (+/-): S.0000.000
- NaNs: S.1111.{001, 010, 011, 100, 101, 110, 111}
- Max normal number: S.1110.111 = +/-2^(7) x (1 + 0.875) = +/-240
- Min normal number: S.0001.000 = +/-2^(-6)
- Max subnormal number: S.0000.111 = +/-2^(-6) x 0.875 = +/-2^(-9) x 7
- Min subnormal number: S.0000.001 = +/-2^(-6) x 0.125 = +/-2^(-9)

Related LLVM PRs:

LLVM PR-97179 [APFloat] Add support for f8E4M3 IEEE 754 type (Merged)
LLVM PR-97118 [MLIR] Add f8E4M3 IEEE 754 type (Merged)

RFCs:

StableHLO PR-2486 [RFC] Add f8E4M3 and f8E3M4 types support

Related ml_dtypes PRs:

PR-57 Add float8_e4m3fnuz and float8_e5m2fnuz

C++ Testing

Tested as described in PR-123 Add CMakeLists for C++ tests

cmake --build build -- all test

100% tests passed, 0 tests failed out of 250

Total Test time (real) =  12.50 sec

apivovarov · 2024-08-05T19:32:56Z

Peter, Jake, whenever you have time, could you please review this pull request?
f8E4M3 support has already been merged to LLVM APFloat and MLIR.
I also need to add f8E4M3 to XLA. Internally XLA uses tsl which depends on ml_dtypes - tsl/platform/ml_dtypes.h

@hawkinsp @jakevdp

hawkinsp · 2024-08-07T19:49:37Z

@apivovarov I'm on vacation this week, but I'll take a look next week. Sorry for the delay...

apivovarov · 2024-08-20T05:40:56Z

@apivovarov I'm on vacation this week, but I'll take a look next week. Sorry for the delay...

Hi Peter, it seems this PR got overlooked. @hawkinsp

BTW, I've also attached a link to the StableHLO [RFC] Add f8E4M3 and f8E3M4 types support openxla/stablehlo#2486

Not sure what change should be merged first StableHLO or ml_dtypes.

hawkinsp

This looks good to me. Would you mind also adding an entry to the changelog?

ml_dtypes/_src/dtypes.cc

hawkinsp · 2024-08-20T18:28:41Z

The linter is also sad, please fix.

hawkinsp · 2024-08-20T18:41:31Z

And it doesn't matter whether the stablehlo change or the ml_dtypes change is merged first, really, they aren't directly coupled.

hawkinsp · 2024-08-21T13:53:53Z

Linter is still sad.

apivovarov · 2024-08-21T16:00:26Z

Linter is still sad.

Fixed. pre-commit run --all-files - all green now

### Summary This is a proposal to add `Float8E4M3` and `Float8E3M4` floating point types to StableHLO. Feedback welcome, see [RFC: Float8E4M3 and Float8E3M4](https://github.com/apivovarov/stablehlo/blob/rfc_f8E4M3_f8E3M4/rfcs/20240808-f8E4M3_f8E3M4.md) for more details. ### References and Links - LLVM [PR-97179](llvm/llvm-project#97179) [APFloat] Add support for f8E4M3 IEEE 754 type (Merged) - LLVM [PR-97118](llvm/llvm-project#97118) [MLIR] Add f8E4M3 IEEE 754 type (Merged) - LLVM [PR-99698](llvm/llvm-project#99698) [APFloat] Add support for f8E3M4 IEEE 754 type (Merged) - LLVM [PR-101230](llvm/llvm-project#101230) [MLIR] Add f8E3M4 IEEE 754 type (Merged) - [RFC: FP8 in StableHLO](https://github.com/openxla/stablehlo/blob/main/rfcs/20221031-fp8.md) - [RFC: Float8E4M3FNUZ and Float8E5M2FNUZ](https://github.com/openxla/stablehlo/blob/main/rfcs/20230321-fp8_fnuz.md) - StableHLO [PR-2482](#2482) Add f8E4M3 and f8E3M4 types support - [Amazon EC2 Trn1 Instances](https://aws.amazon.com/ec2/instance-types/trn1/) - ml_dtypes [PR-161](jax-ml/ml_dtypes#161) Add float8_e4m3 (Merged) - ml_dtypes [PR-171](jax-ml/ml_dtypes#171) Add float8_e3m4 (Merged) - XLA [PR-16585](openxla/xla#16585) Add support for float8_e4m3

This PR adds f8E4M3 and f8E3M4 types support. f8E4M3 and f8E3M4 types follow IEEE 754 convention. ```c f8E4M3 (IEEE 754) - Exponent bias: 7 - Maximum stored exponent value: 14 (binary 1110) - Maximum unbiased exponent value: 14 - 7 = 7 - Minimum stored exponent value: 1 (binary 0001) - Minimum unbiased exponent value: 1 − 7 = −6 - Precision specifies the total number of bits used for the significand (mantisa), including implicit leading integer bit = 3 + 1 = 4 - Follows IEEE 754 conventions for representation of special values - Has Positive and Negative zero - Has Positive and Negative infinity - Has NaNs Additional details: - Max exp (unbiased): 7 - Min exp (unbiased): -6 - Infinities (+/-): S.1111.000 - Zeros (+/-): S.0000.000 - NaNs: S.1111.{001, 010, 011, 100, 101, 110, 111} - Max normal number: S.1110.111 = +/-2^(7) x (1 + 0.875) = +/-240 - Min normal number: S.0001.000 = +/-2^(-6) - Max subnormal number: S.0000.111 = +/-2^(-6) x 0.875 = +/-2^(-9) x 7 - Min subnormal number: S.0000.001 = +/-2^(-6) x 0.125 = +/-2^(-9) ``` ```c f8E3M4 (IEEE 754) - Exponent bias: 3 - Maximum stored exponent value: 6 (binary 110) - Maximum unbiased exponent value: 6 - 3 = 3 - Minimum stored exponent value: 1 (binary 001) - Minimum unbiased exponent value: 1 − 3 = −2 - Precision specifies the total number of bits used for the significand (mantissa), including implicit leading integer bit = 4 + 1 = 5 - Follows IEEE 754 conventions for representation of special values - Has Positive and Negative zero - Has Positive and Negative infinity - Has NaNs Additional details: - Max exp (unbiased): 3 - Min exp (unbiased): -2 - Infinities (+/-): S.111.0000 - Zeros (+/-): S.000.0000 - NaNs: S.111.{0,1}⁴ except S.111.0000 - Max normal number: S.110.1111 = +/-2^(6-3) x (1 + 15/16) = +/-2^3 x 31 x 2^(-4) = +/-15.5 - Min normal number: S.001.0000 = +/-2^(1-3) x (1 + 0) = +/-2^(-2) - Max subnormal number: S.000.1111 = +/-2^(-2) x 15/16 = +/-2^(-2) x 15 x 2^(-4) = +/-15 x 2^(-6) - Min subnormal number: S.000.0001 = +/-2^(-2) x 1/16 = +/-2^(-2) x 2^(-4) = +/-2^(-6) ``` Related PRs: - LLVM [PR-97179](llvm/llvm-project#97179) [APFloat] Add support for f8E4M3 IEEE 754 type (Merged) - LLVM [PR-97118](llvm/llvm-project#97118) [MLIR] Add f8E4M3 IEEE 754 type (Merged) - LLVM [PR-99698](llvm/llvm-project#99698) [APFloat] Add support for f8E3M4 IEEE 754 type (Merged) - LLVM [PR-101230](llvm/llvm-project#101230) [MLIR] Add f8E3M4 IEEE 754 type (Merged) - StableHLO [PR-2486](#2486) [RFC] Add f8E4M3 and f8E3M4 types support - ml_dtypes [PR-161](jax-ml/ml_dtypes#161) Add float8_e4m3 (Merged) - ml_dtypes [PR-171](jax-ml/ml_dtypes#171) Add float8_e3m4 (Merged) - XLA [PR-16585](openxla/xla#16585) Add support for float8_e4m3

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes openxla/xla#17075 PiperOrigin-RevId: 674396944

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes #17075 PiperOrigin-RevId: 674396944

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes openxla/xla#17075 PiperOrigin-RevId: 674396944

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes #17075 PiperOrigin-RevId: 674396944

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes openxla/xla#17075 PiperOrigin-RevId: 674396944

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes #17075 PiperOrigin-RevId: 674396944

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes openxla/xla#17075 PiperOrigin-RevId: 675687080

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes #17075 PiperOrigin-RevId: 675687080

ml_dtypes Updates: Add float8_e4m3 and float8_e3m4 types support Fix float divmod with zero denominator Add int2 and uint2 types ml_dtypes/commits Related PRs ml_dtypes PR Add float8_e4m3 jax-ml/ml_dtypes#161 Add float8_e4m3 (Merged) XLA PR Add support for float8_e4m3 #16585 (In Review) This closes openxla/xla#17075 PiperOrigin-RevId: 675687080

jakevdp requested a review from hawkinsp August 1, 2024 16:49

jakevdp assigned hawkinsp Aug 1, 2024

apivovarov force-pushed the e4m3 branch from 315e379 to ab2fbdc Compare August 1, 2024 22:10

This was referenced Aug 2, 2024

[MLIR] Add f8E4M3 IEEE 754 type llvm/llvm-project#97118

Merged

[APFloat] Add support for f8E4M3 IEEE 754 type llvm/llvm-project#97179

Merged

This was referenced Aug 13, 2024

Remove dup third_party/py/ml_dtypes as it is not used. openxla/xla#15831

Closed

[RFC] Add f8E4M3 and f8E3M4 types support openxla/stablehlo#2486

Merged

hawkinsp reviewed Aug 20, 2024

View reviewed changes

ml_dtypes/_src/dtypes.cc Outdated Show resolved Hide resolved

apivovarov force-pushed the e4m3 branch from ab2fbdc to 27d81a6 Compare August 21, 2024 02:17

Add float8_e4m3

1833c0c

apivovarov force-pushed the e4m3 branch from 27d81a6 to 1833c0c Compare August 21, 2024 15:59

hawkinsp approved these changes Aug 21, 2024

View reviewed changes

hawkinsp added the pull ready label Aug 21, 2024

copybara-service bot merged commit 30f2497 into jax-ml:main Aug 21, 2024
11 of 12 checks passed

apivovarov mentioned this pull request Aug 22, 2024

Add float8_e3m4 #171

Merged

apivovarov deleted the e4m3 branch August 23, 2024 19:02

This was referenced Aug 23, 2024

np.floor_divide(?, ml_dtypes.???(0.0)) return NaN but np.float16 returns Inf. #170

Closed

Add support for float8_e4m3 and float8_e3m4 types openxla/xla#16585

Open

Add f8E4M3 and f8E3M4 types support openxla/stablehlo#2482

Merged

This was referenced Sep 11, 2024

[TSL] Bump ml_dtypes to version 0.5.0 openxla/xla#17075

Closed

Add float8_e4m3 and float8_e3m4 types support jax-ml/jax#23585

Open

copybara-service bot mentioned this pull request Sep 13, 2024

[TSL] Bump ml_dtypes. Add float8_e4m3, float8_e3m4 google/tsl#2693

Merged

copybara-service bot mentioned this pull request Sep 13, 2024

[TSL] Bump ml_dtypes. Add float8_e4m3, float8_e3m4 openxla/xla#17167

Merged

copybara-service bot mentioned this pull request Sep 13, 2024

[TSL] Bump ml_dtypes. Add float8_e4m3, float8_e3m4 tensorflow/tensorflow#75735

Merged

apivovarov mentioned this pull request Sep 16, 2024

[TSL] Bump ml_dtypes to 0.5.0 openxla/xla#17230

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Add float8_e4m3 #161

Add float8_e4m3 #161

apivovarov commented Aug 1, 2024 •

edited

Loading

apivovarov commented Aug 5, 2024

hawkinsp commented Aug 7, 2024

apivovarov commented Aug 20, 2024

hawkinsp left a comment

hawkinsp commented Aug 20, 2024

hawkinsp commented Aug 20, 2024

hawkinsp commented Aug 21, 2024

apivovarov commented Aug 21, 2024

Add float8_e4m3 #161

Add float8_e4m3 #161

Conversation

apivovarov commented Aug 1, 2024 • edited Loading

C++ Testing

apivovarov commented Aug 5, 2024

hawkinsp commented Aug 7, 2024

apivovarov commented Aug 20, 2024

hawkinsp left a comment

Choose a reason for hiding this comment

hawkinsp commented Aug 20, 2024

hawkinsp commented Aug 20, 2024

hawkinsp commented Aug 21, 2024

apivovarov commented Aug 21, 2024

apivovarov commented Aug 1, 2024 •

edited

Loading