package pipeline
- Alphabetic
- Public
- All
The incremental task API lets tasks run more quickly when they are called more than once.
The incremental task API lets tasks run more quickly when they are called more than once. The idea is to do less work when tasks are called a second time, by skipping any work that has already been done. In other words, tasks only perform the “incremental” work that is necessary since they were last run.
To analyse which work needs to be done, a task’s work is broken up into a number of sub-operations, each of which can be run independently. Each operation takes input parameters and can read and write files. The incremental task API keeps a record of which operations have been run so that that those operations don’t need to be repeated in the future.
Here is how tasks interact with the API:
- Tasks call the API with a list of potential operations to perform and with a function to run operations.
- The API takes care of pruning the list of operations to find the incremental operations that need to be run.
- The API then calls the supplied function to run the pruned list of operations. This method returns a list of results, one for each operation.
- If an operation succeeds, the details are recorded so that the operation can be skipped in the future, if possible.
Behind the scenes, syncIncremental maintains a record of each operation that succeeds. It uses these records to work out which operations need to be run and which can be skipped.
Each operation is assumed to take some input parameters, optionally read and write some files, and either succeed or fail when it runs. An operation which fails will always be run again even if its parameters and input files remain the same. But an operation which succeeds will only need to be run again if its input parameters change or if the contents of any files it read from or wrote to have changed.