[Common] Experimental CuTeDSL MXFP4 backend - #3223
Draft
janekb04 wants to merge 50 commits into
Draft
Conversation
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
…it__.py Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci
for more information, see https://pre-commit.ci
for more information, see https://pre-commit.ci
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
janekb04
force-pushed
the
cutedsl_nvfp4
branch
3 times, most recently
from
July 22, 2026 21:02
c69c284 to
8d04ce3
Compare
for more information, see https://pre-commit.ci
Route the optimized 1D NVFP4 quantize-transpose path through the new CuTe DSL backend, falling back to the existing CUDA kernel when the CuTe DSL backend declines to handle the case. Adds the FP4 dtype mapping to the TVM FFI bridge and an empty stub header for the backend. Signed-off-by: Jan Bielak <jbielak@nvidia.com>
Add the C++ and Python scaffolding for dispatching NVFP4 quantize-transpose to the CuTe DSL backend via TVM FFI, mirroring the general structure and file layout of the MXFP8 backend (quantize_mxfp8_cutedsl.cuh and quantize_mxfp8.py). Kernel instantiation parameters and validation are left as todos. Signed-off-by: Jan Bielak <jbielak@nvidia.com>
Populate NVFP4QuantizeConfig with the kernel instantiation parameters (stochastic rounding, fast math, row-scaled NVFP4 and transpose flags) on both the C++ and Python sides, and add the input/output validation mirroring the CUDA implementation in quantize_transpose_nvfp4.cuh. The config values are derived from the quantization config and output tensor and forwarded across the TVM FFI boundary, and the backend is guarded behind FP4_TYPE_SUPPORTED. Signed-off-by: Jan Bielak <jbielak@nvidia.com>
for more information, see https://pre-commit.ci
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This PR adds a CuTe DSL NVFP4 quantization backend. It is based on the work in #3137.
In this PR only a CuTe DSL version of
quantize_transpose_tuned_1Dis implemented. It handles the case of 1Dbfloat16to NVFP4 quantization, mirroring the capabilities of the original CUDA implementation inquantize_transpose_nvfp4_tuned_1D.cuh.To be rebased once #3137 merges.
Overview
Below is the overview of #3137, which introduces the CuTe DSL infrastructure that this PR is based on.
Kernels using CuTe DSL are implemented in
quantize_mxfp8.py.These kernels are implemented as
@cute.jitfunctions, which are wrapped inKernelclasses with a@cute.kernel__call__method.get_mxfp8_quantization_functionis registered from the Python side with TVM FFI. It returns a Python function callable from the C++ side that runs the CuTE DSL kernel for a particular configuration.The C++ side of the TVM FFI bridge is implemented in
quantize_mxfp8_cutedsl.cuh. Themxfp8_quantize_cutedslfunction has the same interface as the originalmxfp8::quantize. It tries to call the Python side implementation.The
MXFP8QuantConfigis what gets carried over between the sides. The kernel is instantiated once per every config. It is like the kernel's template parameters.quantize.cuhsimply tries to call the CuTe DSL version if it is available instead of the regular implementation.This PR mirrors the structure of #3137:
quantize_transpose_nvfp4.py(vs.quantize_mxfp8.py), registeringget_nvfp4_quantization_function(vs.get_mxfp8_quantization_function) with TVM FFI.quantize_transpose_nvfp4_cutedsl.cuh(vs.quantize_mxfp8_cutedsl.cuh), exposingnvfp4_quantize_transpose_cutedsl(vs.mxfp8_quantize_cutedsl) with the same interface asnvfp4::quantize_transpose.NVFP4QuantConfig(vs.MXFP8QuantConfig).quantize.cuhtriescutedsl_backend::nvfp4_quantize_transpose_cutedslfirst and falls back to the CUDAnvfp4::quantize_transposekernel if it declines (same fallback pattern as the MXFP8 path).Type of change
Changes
Support for the CuTe DSL backend is implemented in a top-down manner, one commit at a time:
quantize.cuh(done)tvm_ffi_bridge.h) and an empty stub header for the backend.quantize_transpose_nvfp4_cutedsl.cuh, mirroringquantize_mxfp8_cutedsl.cuh.quantize_transpose_nvfp4.py, mirroringquantize_mxfp8.py.NVFP4QuantizeConfigwith the kernel instantiation parameters (stochastic rounding, fast math, row-scaled NVFP4 and transpose flags) on both the C++ and Python sides, mirroringquantize_transpose_nvfp4.cuh:1441. The config values are derived from the quantization config and output tensor and forwarded across the TVM FFI boundary.FP4_TYPE_SUPPORTED.DLTensorWrapperand actually invoke the Python-side entrypoint, instead of just resolving it.cute.compile(..., options="--enable-tvm-ffi")and register it under the config-derived key viatvm_ffi.register_global_func, with graceful fallback to the CUDA kernel on unsupported configs or compilation failure.@cute.kerneldevice function and the host-side launch (tile/thread/stage sizing, grid/block computation, TMA atom setup) in__call__are still stubs.test_mxfp8_cutedsl_backend.py) once the kernel is functional.Checklist: