github NVIDIA/cudnn-frontend v1.0-pre-release-5
cudnn FE 1.0 pre-release-5

pre-release2 years ago

Pre-release-5 release notes:

[API change] Based on user feedback, we have removed distinction between the graph and plan objects. With the new API, plan remains embedded in the graph and all operations are performed on the graph object.

Previously,

    REQUIRE(graph.validate().is_good());
    REQUIRE(graph.build_operation_graph(handle).is_good());
    auto plans = graph.get_execution_plan_list({fe::HeurMode_t::A});
    REQUIRE(plans.check_support(handle).is_good());
    REQUIRE(graph.set_execution_plans(plans).is_good());

Now,

    REQUIRE(graph.validate().is_good());
    REQUIRE(graph.build_operation_graph(handle).is_good());
    REQUIRE(graph.create_execution_plans({fe::HeurMode_t::A}).is_good());
    REQUIRE(graph.check_support(handle).is_good());
    REQUIRE(graph.build_plans(handle).is_good());

Also, with this change the following new API have been introduced on the graph class.

error_t
build_plans(cudnnHandle_t const &handle,
            BuildPlanPolicy_t const policy     = BuildPlanPolicy_t::HEURISTICS_CHOICE,
            bool const do_multithreaded_builds = false);

Graph & deselect_workspace_greater_than(int64_t const workspace);

Graph & deselect_behavior_notes(std::vector<BehaviorNote_t> const &notes);

Graph & deselect_numeric_notes(std::vector<NumericalNote_t> const &notes);

int64_t get_workspace_size() const 


int64_t get_autotune_workspace_size() const;

error_t autotune(cudnnHandle_t handle,
             std::unordered_map<std::shared_ptr<Tensor_attributes>, void *> variants,
             void *workspace,
             void *user_impl = nullptr);

[API change] Removes the implicit validate call made in build_operation_graph. Now, the expectation is that the user explicitly calls validate on the graph before calling build_operation_graph. This helps the user distinguish errors between malformed graphs and error occuring due to lowering into cudnn.

[API change] Return error codes from the graph API have now been marked nodiscard.

[New API] Have added a new graph::key() -> int64_t as an API that returns a hash on the graph object. This can be used as key for graph caching. Eg. of this usage is shown in the samples.

[New API] Have added new python API create_handle, destroy_handle, set_stream, get_stream to allow custom handle and stream management on the graph object.

[New functionality] sdpa backward can now compute dbias if the fprop had a bias operation. This functionality was added in cudnn 8.9.6.

[Enhancement] There is a extension in behavior of CUDNN_FRONTEND_ATTN_DP_WORKSPACE_LIMIT. This is documented in docs/operation/Attention.md

[Enhancement] Have added better error checks to make sure all the tensors of the node have been created. This prevents unexpected segmentation faults seen earlier.

[Bug Fix] Fix issues in instancenorm, which had caused invalid memory access earlier.

[Enhancement] Have moved the v0.9 API samples to samples/legacy_samples folder for better organization.

Don't miss a new cudnn-frontend release

NewReleases is sending notifications on new releases.