The lifecycle of a model is implemented in the following steps:
Open the model file and read the shapes of the weights, but don’t load the weights yet.
Using the loaded shapes and optional metadata, instantiate a model struct with
Tensors, representing the shape and layout of each layer of the NN.Compile the model struct and it’s
forwardfunction into an accelerator specific executable. Theforwardfunction describes the mathematical operations corresponding to the model inference.Load the model weights from disk, onto the accelerator memory.
Load some user inputs, and copy them to the accelerator.
Call the executable using the weights and the user inputs.
Fetch the returned model output from accelerator into host memory, and finally present it to the user.
When all user inputs have been processed, free the executable resources, weights and inputs.