tf namespace
taskflow namespace
Classes
- class ChromeObserver
- observer interface based on Chrome tracing format
- class CriticalSection
- class to create a critical region of limited workers to run tasks
-
template<unsigned NT, unsigned VT>class cudaExecutionPolicy
- class to define execution policy for CUDA standard algorithms
- class cudaFlow
- class for building a CUDA task dependency graph
- class cudaFlowCapturer
- class for building a CUDA task dependency graph through stream capture
- class cudaLinearCapturing
- class to capture a linear CUDA graph using a sequential stream
- class cudaRoundRobinCapturing
- class to capture a CUDA graph using a round-robin algorithm
- class cudaScopedDevice
- RAII-styled device context switch.
- class cudaScopedPerThreadEvent
- class that provides RAII-styled guard of event acquisition
- class cudaScopedPerThreadStream
- class that provides RAII-styled guard of stream acquisition
- class cudaSequentialCapturing
- class to capture a CUDA graph using a sequential stream
- class cudaTask
- handle to a node of the internal CUDA graph
- class Executor
- execution interface for running a taskflow graph
- class FlowBuilder
- building methods of a task dependency graph
-
template<typename T>class Future
- class to access the result of task execution
- class ObserverInterface
- The interface class for creating an executor observer.
- class Semaphore
- class to create a semophore object for building a concurrency constraint
- class Subflow
- class to construct a subflow graph from the execution of a dynamic task
- class syclFlow
- class for building a SYCL task dependency graph
- class syclTask
- handle to a node of the internal CUDA graph
- class Task
- handle to a node in a task dependency graph
- class Taskflow
- main entry to create a task dependency graph
- class TaskView
- class to access task information from the observer interface
- class TFProfObserver
- observer interface based on the built-in taskflow profiler format
- class WorkerView
- class to create an immutable view of a worker in an executor
Enums
- enum class TaskType: int { PLACEHOLDER = 0, CUDAFLOW, SYCLFLOW, STATIC, DYNAMIC, CONDITION, MODULE, ASYNC, UNDEFINED }
- enumeration of all task types
- enum class ObserverType: int { TFPROF = 0, CHROME, UNDEFINED }
- enumeration of all observer types
- enum class cudaTaskType: int { EMPTY = 0, HOST, MEMSET, MEMCPY, KERNEL, SUBFLOW, CAPTURE, UNDEFINED }
- enumeration of all cudaTask types
Typedefs
-
using observer_stamp_t = std::
chrono:: time_point<std:: chrono:: steady_clock> - default time point type of observers
- using cudaPerThreadStreamPool = cudaPerThreadDeviceObjectPool<cudaStream_t, cudaStreamCreator, cudaStreamDeleter>
- alias of per-thread stream pool type
- using cudaPerThreadEventPool = cudaPerThreadDeviceObjectPool<cudaEvent_t, cudaEventCreator, cudaEventDeleter>
- alias of per-thread event pool type
- using cudaDefaultExecutionPolicy = cudaExecutionPolicy<512, 9>
- default execution policy
Functions
- auto to_string(TaskType type) -> const char*
- convert a task type to a human-readable string
-
auto operator<<(std::
ostream& os, const Task& task) -> std:: ostream& - overload of ostream inserter operator for cudaTask
- auto to_string(ObserverType type) -> const char*
- convert an observer type to a human-readable string
- auto cuda_get_num_devices() -> size_t
- queries the number of available devices
- auto cuda_get_device() -> int
- gets the current device associated with the caller thread
- void cuda_set_device(int id)
- switches to a given device context
- void cuda_get_device_property(int i, cudaDeviceProp& p)
- obtains the device property
- auto cuda_get_device_property(int i) -> cudaDeviceProp
- obtains the device property
-
void cuda_dump_device_property(std::
ostream& os, const cudaDeviceProp& p) - dumps the device property
- auto cuda_get_device_max_threads_per_block(int d) -> size_t
- queries the maximum threads per block on a device
- auto cuda_get_device_max_x_dim_per_block(int d) -> size_t
- queries the maximum x-dimension per block on a device
- auto cuda_get_device_max_y_dim_per_block(int d) -> size_t
- queries the maximum y-dimension per block on a device
- auto cuda_get_device_max_z_dim_per_block(int d) -> size_t
- queries the maximum z-dimension per block on a device
- auto cuda_get_device_max_x_dim_per_grid(int d) -> size_t
- queries the maximum x-dimension per grid on a device
- auto cuda_get_device_max_y_dim_per_grid(int d) -> size_t
- queries the maximum y-dimension per grid on a device
- auto cuda_get_device_max_z_dim_per_grid(int d) -> size_t
- queries the maximum z-dimension per grid on a device
- auto cuda_get_device_max_shm_per_block(int d) -> size_t
- queries the maximum shared memory size in bytes per block on a device
- auto cuda_get_device_warp_size(int d) -> size_t
- queries the warp size on a device
- auto cuda_get_device_compute_capability_major(int d) -> int
- queries the major number of compute capability of a device
- auto cuda_get_device_compute_capability_minor(int d) -> int
- queries the minor number of compute capability of a device
- auto cuda_get_device_unified_addressing(int d) -> bool
- queries if the device supports unified addressing
- auto cuda_get_driver_version() -> int
- queries the latest CUDA version (1000 * major + 10 * minor) supported by the driver
- auto cuda_get_runtime_version() -> int
- queries the CUDA Runtime version (1000 * major + 10 * minor)
- auto cuda_get_free_mem(int d) -> size_t
- queries the free memory (expensive call)
- auto cuda_get_total_mem(int d) -> size_t
- queries the total available memory (expensive call)
-
template<typename T>auto cuda_malloc_device(size_t N, int d) -> T*
- allocates memory on the given device for holding
Nelements of typeT -
template<typename T>auto cuda_malloc_device(size_t N) -> T*
- allocates memory on the current device associated with the caller
-
template<typename T>auto cuda_malloc_shared(size_t N) -> T*
- allocates shared memory for holding
Nelements of typeT -
template<typename T>void cuda_free(T* ptr, int d)
- frees memory on the GPU device
-
template<typename T>void cuda_free(T* ptr)
- frees memory on the GPU device
- void cuda_memcpy_async(cudaStream_t stream, void* dst, const void* src, size_t count)
- copies data between host and device asynchronously through a stream
- void cuda_memset_async(cudaStream_t stream, void* devPtr, int value, size_t count)
- initializes or sets GPU memory to the given value byte by byte
- auto cuda_per_thread_stream_pool() -> cudaPerThreadStreamPool&
- acquires the per-thread cuda stream pool
- auto cuda_per_thread_event_pool() -> cudaPerThreadEventPool&
- per-thread cuda event pool
- auto to_string(cudaTaskType type) -> const char* constexpr
- convert a cuda_task type to a human-readable string
-
auto operator<<(std::
ostream& os, const cudaTask& ct) -> std:: ostream& - overload of ostream inserter operator for cudaTask
-
template<typename P, typename C>void cuda_single_task_async(P&& p, C c)
- runs a callable asynchronously using one kernel thread
-
template<typename P, typename C>void cuda_single_task(P&& p, C c)
- runs a callable using one kernel thread
-
template<typename P, typename I, typename C>void cuda_for_each_async(P&& p, I first, I last, C c)
- performs asynchronous parallel iterations over a range of items
-
template<typename P, typename I, typename C>void cuda_for_each(P&& p, I first, I last, C c)
- performs parallel iterations over a range of items
-
template<typename P, typename I, typename C>void cuda_for_each_index_async(P&& p, I first, I last, I inc, C c)
- performs asynchronous parallel iterations over an index-based range of items
-
template<typename P, typename I, typename C>void cuda_for_each_index(P&& p, I first, I last, I inc, C c)
- performs parallel iterations over an index-based range of items
-
template<typename P, typename I, typename O, typename C>void cuda_transform_async(P&& p, I first, I last, O output, C op)
- performs asynchronous parallel transforms over a range of items
-
template<typename P, typename I, typename O, typename C>void cuda_transform(P&& p, I first, I last, O output, C op)
- performs parallel transforms over a range of items
-
template<typename P, typename I1, typename I2, typename O, typename C>void cuda_transform_async(P&& p, I1 first1, I1 last1, I2 first2, O output, C op)
- performs asynchronous parallel transforms over two ranges of items
-
template<typename P, typename I1, typename I2, typename O, typename C>void cuda_transform(P&& p, I1 first1, I1 last1, I2 first2, O output, C op)
- performs parallel transforms over two ranges of items
-
template<typename P, typename T>auto cuda_reduce_buffer_size(unsigned count) -> unsigned
- queries the buffer size in bytes needed to call reduce kernels
-
template<typename P, typename I, typename T, typename O>void cuda_reduce(P&& p, I first, I last, T* res, O op)
- performs parallel reduction over a range of items
-
template<typename P, typename I, typename T, typename O>void cuda_reduce_async(P&& p, I first, I last, T* res, O op, void* buf)
- performs asynchronous parallel reduction over a range of items
-
template<typename P, typename I, typename T, typename O>void cuda_uninitialized_reduce(P&& p, I first, I last, T* res, O op)
- performs parallel reduction over a range of items without an initial value
-
template<typename P, typename I, typename T, typename O>void cuda_uninitialized_reduce_async(P&& p, I first, I last, T* res, O op, void* buf)
- performs asynchronous parallel reduction over a range of items without an initial value
-
template<typename P, typename I, typename T, typename O, typename U>void cuda_transform_reduce(P&& p, I first, I last, T* res, O bop, U uop)
- performs parallel reduction over a range of transformed items without an initial value
-
template<typename P, typename I, typename T, typename O, typename U>void cuda_transform_reduce_async(P&& p, I first, I last, T* res, O bop, U uop, void* buf)
- performs asynchronous parallel reduction over a range of transformed items without an initial value
-
template<typename P, typename I, typename T, typename O, typename U>void cuda_transform_uninitialized_reduce(P&& p, I first, I last, T* res, O bop, U uop)
- performs parallel reduction over a range of transformed items with an initial value
-
template<typename P, typename I, typename T, typename O, typename U>void cuda_transform_uninitialized_reduce_async(P&& p, I first, I last, T* res, O bop, U uop, void* buf)
- performs asynchronous parallel reduction over a range of transformed items with an initial value
-
template<typename P, typename T>auto cuda_scan_buffer_size(unsigned count) -> unsigned
- queries the buffer size in bytes needed to call scan kernels
-
template<typename P, typename I, typename O, typename C>void cuda_inclusive_scan(P&& p, I first, I last, O output, C op)
- performs inclusive scan over a range of items
-
template<typename P, typename I, typename O, typename C>void cuda_inclusive_scan_async(P&& p, I first, I last, O output, C op, void* buf)
- performs asynchronous inclusive scan over a range of items
-
template<typename P, typename I, typename O, typename C, typename U>void cuda_transform_inclusive_scan(P&& p, I first, I last, O output, C bop, U uop)
- performs inclusive scan over a range of transformed items
-
template<typename P, typename I, typename O, typename C, typename U>void cuda_transform_inclusive_scan_async(P&& p, I first, I last, O output, C bop, U uop, void* buf)
- performs asynchronous inclusive scan over a range of transformed items
-
template<typename P, typename I, typename O, typename C>void cuda_exclusive_scan(P&& p, I first, I last, O output, C op)
- performs exclusive scan over a range of items
-
template<typename P, typename I, typename O, typename C>void cuda_exclusive_scan_async(P&& p, I first, I last, O output, C op, void* buf)
- performs asynchronous exclusive scan over a range of items
-
template<typename P, typename I, typename O, typename C, typename U>void cuda_transform_exclusive_scan(P&& p, I first, I last, O output, C bop, U uop)
- performs exclusive scan over a range of items
-
template<typename P, typename I, typename O, typename C, typename U>void cuda_transform_exclusive_scan_async(P&& p, I first, I last, O output, C bop, U uop, void* buf)
- performs asynchronous exclusive scan over a range of items
-
template<typename P>auto cuda_merge_buffer_size(unsigned a_count, unsigned b_count) -> unsigned
- queries the buffer size in bytes needed to call merge kernels
-
template<typename P, typename a_keys_it, typename a_vals_it, typename b_keys_it, typename b_vals_it, typename c_keys_it, typename c_vals_it, typename C>void cuda_merge(P&& p, a_keys_it a_keys_first, a_vals_it a_vals_first, a_keys_it a_keys_last, b_keys_it b_keys_first, b_vals_it b_vals_first, b_keys_it b_keys_last, c_keys_it c_keys_first, c_vals_it c_vals_first, C comp)
- performs key-value merge over a range of keys and values
-
template<typename P, typename a_keys_it, typename a_vals_it, typename b_keys_it, typename b_vals_it, typename c_keys_it, typename c_vals_it, typename C>void cuda_merge_async(P&& p, a_keys_it a_keys_first, a_vals_it a_vals_first, a_keys_it a_keys_last, b_keys_it b_keys_first, b_vals_it b_vals_first, b_keys_it b_keys_last, c_keys_it c_keys_first, c_vals_it c_vals_first, C comp, void* buf)
- performs asynchronous key-value merge over a range of keys and values
-
template<typename P, typename a_keys_it, typename b_keys_it, typename c_keys_it, typename C>void cuda_merge(P&& p, a_keys_it a_keys_first, a_keys_it a_keys_last, b_keys_it b_keys_first, b_keys_it b_keys_last, c_keys_it c_keys_first, C comp)
- performs key-only merge over a range of keys
-
template<typename P, typename a_keys_it, typename b_keys_it, typename c_keys_it, typename C>void cuda_merge_async(P&& p, a_keys_it a_keys_first, a_keys_it a_keys_last, b_keys_it b_keys_first, b_keys_it b_keys_last, c_keys_it c_keys_first, C comp, void* buf)
- performs asynchronous key-only merge over a range of keys
-
template<typename P, typename K, typename V = cudaEmpty>auto cuda_sort_buffer_size(unsigned count) -> unsigned
- queries the buffer size in bytes needed to call sort kernels for the given number of elements
-
template<typename P, typename K_it, typename V_it, typename C>void cuda_sort(P&& p, K_it k_first, K_it k_last, V_it v_first, C comp)
- performs key-value sort on a range of items
-
template<typename P, typename K_it, typename V_it, typename C>void cuda_sort_async(P&& p, K_it k_first, K_it k_last, V_it v_first, C comp, void* buf)
- performs asynchronous key-value sort on a range of items
-
template<typename P, typename K_it, typename C>void cuda_sort(P&& p, K_it k_first, K_it k_last, C comp)
- performs key-only sort on a range of items
-
template<typename P, typename K_it, typename C>void cuda_sort_async(P&& p, K_it k_first, K_it k_last, C comp, void* buf)
- performs asynchronous key-only sort on a range of items
-
auto operator<<(std::
ostream& os, const syclTask& ct) -> std:: ostream& - overload of ostream inserter operator for syclTask
- auto version() -> const char* constexpr
- queries the version information in a string format
major.minor.patch
Variables
-
std::
array<TaskType, 8> TASK_TYPES constexpr - array of all task types (used for iterating task types)
-
template<typename C>bool is_static_task_v constexpr
- determines if a callable is a static task
-
template<typename C>bool is_dynamic_task_v constexpr
- determines if a callable is a dynamic task
-
template<typename C>bool is_condition_task_v constexpr
- determines if a callable is a condition task
-
template<typename C>bool is_cudaflow_task_v constexpr
- determines if a callable is a cudaFlow task
-
template<typename C>bool is_syclflow_task_v constexpr
- determines if a callable is a syclFlow task
Enum documentation
enum class tf:: TaskType: int
enumeration of all task types
| Enumerators | |
|---|---|
| PLACEHOLDER |
placeholder task type |
| CUDAFLOW |
cudaFlow task type |
| SYCLFLOW |
syclFlow task type |
| STATIC |
static task type |
| DYNAMIC |
dynamic (subflow) task type |
| CONDITION |
condition task type |
| MODULE |
module task type |
| ASYNC |
asynchronous task type |
| UNDEFINED |
undefined task type (for internal use only) |
enum class tf:: cudaTaskType: int
enumeration of all cudaTask types
| Enumerators | |
|---|---|
| EMPTY |
empty task type |
| HOST |
host task type |
| MEMSET |
memory set task type |
| MEMCPY |
memory copy task type |
| KERNEL |
memory copy task type |
| SUBFLOW |
subflow (child graph) task type |
| CAPTURE |
capture task type |
| UNDEFINED |
undefined task type |
Function documentation
template<typename T>
T* tf:: cuda_malloc_device(size_t N,
int d)
allocates memory on the given device for holding N elements of type T
The function calls cudaMalloc to allocate N*sizeof(T) bytes of memory on the given device d and returns a pointer to the starting address of the device memory.
template<typename T>
T* tf:: cuda_malloc_device(size_t N)
allocates memory on the current device associated with the caller
The function calls malloc_device from the current device associated with the caller.
template<typename T>
T* tf:: cuda_malloc_shared(size_t N)
allocates shared memory for holding N elements of type T
The function calls cudaMallocManaged to allocate N*sizeof(T) bytes of memory and returns a pointer to the starting address of the shared memory.
template<typename T>
void tf:: cuda_free(T* ptr,
int d)
frees memory on the GPU device
| Template parameters | |
|---|---|
| T | pointer type |
| Parameters | |
| ptr | device pointer to memory to free |
| d | device context identifier |
This methods call cudaFree to free the memory space pointed to by ptr using the given device context.
template<typename T>
void tf:: cuda_free(T* ptr)
frees memory on the GPU device
| Template parameters | |
|---|---|
| T | pointer type |
| Parameters | |
| ptr | device pointer to memory to free |
This methods call cudaFree to free the memory space pointed to by ptr using the current device context of the caller.
void tf:: cuda_memcpy_async(cudaStream_t stream,
void* dst,
const void* src,
size_t count)
copies data between host and device asynchronously through a stream
| Parameters | |
|---|---|
| stream | stream identifier |
| dst | destination memory address |
| src | source memory address |
| count | size in bytes to copy |
The method calls cudaMemcpyAsync with the given stream using cudaMemcpyDefault to infer the memory space of the source and the destination pointers. The memory areas may not overlap.
void tf:: cuda_memset_async(cudaStream_t stream,
void* devPtr,
int value,
size_t count)
initializes or sets GPU memory to the given value byte by byte
| Parameters | |
|---|---|
| stream | stream identifier |
| devPtr | pointer to GPU mempry |
| value | value to set for each byte of the specified memory |
| count | size in bytes to set |
The method calls cudaMemsetAsync with the given stream to fill the first count bytes of the memory area pointed to by devPtr with the constant byte value value.
template<typename P, typename C>
void tf:: cuda_single_task_async(P&& p,
C c)
runs a callable asynchronously using one kernel thread
| Template parameters | |
|---|---|
| P | execution policy type |
| C | closure type |
| Parameters | |
| p | execution policy |
| c | closure to run by one kernel thread |
template<typename P, typename C>
void tf:: cuda_single_task(P&& p,
C c)
runs a callable using one kernel thread
| Template parameters | |
|---|---|
| P | execution policy type |
| C | closure type |
| Parameters | |
| p | execution policy |
| c | closure to run by one kernel thread |
template<typename P, typename I, typename C>
void tf:: cuda_for_each_async(P&& p,
I first,
I last,
C c)
performs asynchronous parallel iterations over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| C | unary operator type |
| Parameters | |
| p | execution policy object |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| c | unary operator to apply to each dereferenced iterator |
Please refer to Parallel Iterations for details.
template<typename P, typename I, typename C>
void tf:: cuda_for_each(P&& p,
I first,
I last,
C c)
performs parallel iterations over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| C | unary operator type |
| Parameters | |
| p | execution policy object |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| c | unary operator to apply to each dereferenced iterator |
Please refer to Parallel Iterations for details.
template<typename P, typename I, typename C>
void tf:: cuda_for_each_index_async(P&& p,
I first,
I last,
I inc,
C c)
performs asynchronous parallel iterations over an index-based range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input index type |
| C | unary operator type |
| Parameters | |
| p | execution policy object |
| first | index to the beginning of the range |
| last | index to the end of the range |
| inc | step size between successive iterations |
| c | unary operator to apply to each index |
Please refer to Parallel Iterations for details.
template<typename P, typename I, typename C>
void tf:: cuda_for_each_index(P&& p,
I first,
I last,
I inc,
C c)
performs parallel iterations over an index-based range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input index type |
| C | unary operator type |
| Parameters | |
| p | execution policy object |
| first | index to the beginning of the range |
| last | index to the end of the range |
| inc | step size between successive iterations |
| c | unary operator to apply to each index |
Please refer to Parallel Iterations for details.
template<typename P, typename I, typename O, typename C>
void tf:: cuda_transform_async(P&& p,
I first,
I last,
O output,
C op)
performs asynchronous parallel transforms over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| O | output iterator type |
| C | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| output | iterator to the beginning of the output range |
| op | unary operator to apply to transform each item |
This method is equivalent to the parallel execution of the following loop on a GPU:
while (first != last) { *output++ = op(*first++); }
Please refer to Parallel Transforms for details.
template<typename P, typename I, typename O, typename C>
void tf:: cuda_transform(P&& p,
I first,
I last,
O output,
C op)
performs parallel transforms over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| O | output iterator type |
| C | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| output | iterator to the beginning of the output range |
| op | unary operator to apply to transform each item |
This method is equivalent to the parallel execution of the following loop on a GPU:
while (first != last) { *output++ = op(*first++); }
Please refer to Parallel Transforms for details.
template<typename P, typename I1, typename I2, typename O, typename C>
void tf:: cuda_transform_async(P&& p,
I1 first1,
I1 last1,
I2 first2,
O output,
C op)
performs asynchronous parallel transforms over two ranges of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I1 | first input iterator type |
| I2 | second input iterator type |
| O | output iterator type |
| C | binary operator type |
| Parameters | |
| p | execution policy |
| first1 | iterator to the beginning of the first range |
| last1 | iterator to the end of the first range |
| first2 | iterator to the beginning of the second range |
| output | iterator to the beginning of the output range |
| op | binary operator to apply to transform each pair of items |
This method is equivalent to the parallel execution of the following loop on a GPU:
while (first1 != last1) { *output++ = op(*first1++, *first2++); }
Please refer to Parallel Transforms for details.
template<typename P, typename I1, typename I2, typename O, typename C>
void tf:: cuda_transform(P&& p,
I1 first1,
I1 last1,
I2 first2,
O output,
C op)
performs parallel transforms over two ranges of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I1 | first input iterator type |
| I2 | second input iterator type |
| O | output iterator type |
| C | binary operator type |
| Parameters | |
| p | execution policy |
| first1 | iterator to the beginning of the first range |
| last1 | iterator to the end of the first range |
| first2 | iterator to the beginning of the second range |
| output | iterator to the beginning of the output range |
| op | binary operator to apply to transform each pair of items |
This method is equivalent to the parallel execution of the following loop on a GPU:
while (first1 != last1) { *output++ = op(*first1++, *first2++); }
Please refer to Parallel Transforms for details.
template<typename P, typename T>
unsigned tf:: cuda_reduce_buffer_size(unsigned count)
queries the buffer size in bytes needed to call reduce kernels
| Template parameters | |
|---|---|
| P | execution policy type |
| T | value type |
| Parameters | |
| count | number of elements to reduce |
The function is used to allocate a buffer for calling asynchronous reduce. Please refer to Parallel Reduction for details.
template<typename P, typename I, typename T, typename O>
void tf:: cuda_reduce(P&& p,
I first,
I last,
T* res,
O op)
performs parallel reduction over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| T | value type |
| O | binary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| res | pointer to the result |
| op | binary operator to apply to reduce elements |
Please refer to Parallel Reduction for details.
template<typename P, typename I, typename T, typename O>
void tf:: cuda_reduce_async(P&& p,
I first,
I last,
T* res,
O op,
void* buf)
performs asynchronous parallel reduction over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| T | value type |
| O | binary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| res | pointer to the result |
| op | binary operator to apply to reduce elements |
| buf | pointer to the temporary buffer |
Please refer to Parallel Reduction for details.
template<typename P, typename I, typename T, typename O>
void tf:: cuda_uninitialized_reduce(P&& p,
I first,
I last,
T* res,
O op)
performs parallel reduction over a range of items without an initial value
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| T | value type |
| O | binary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| res | pointer to the result |
| op | binary operator to apply to reduce elements |
Similar to tf::res to reduce. Please refer to Parallel Reduction for more details.
template<typename P, typename I, typename T, typename O>
void tf:: cuda_uninitialized_reduce_async(P&& p,
I first,
I last,
T* res,
O op,
void* buf)
performs asynchronous parallel reduction over a range of items without an initial value
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| T | value type |
| O | binary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| res | pointer to the result |
| op | binary operator to apply to reduce elements |
| buf | pointer to the temporary buffer |
Asynchronous version of tf::
template<typename P, typename I, typename T, typename O, typename U>
void tf:: cuda_transform_reduce(P&& p,
I first,
I last,
T* res,
O bop,
U uop)
performs parallel reduction over a range of transformed items without an initial value
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| T | value type |
| O | binary operator type |
| U | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| res | pointer to the result |
| bop | binary operator to apply to reduce elements |
| uop | unary operator to apply to transform elements |
Transforms each element in the range using the unary operator uop and then reduce these transformed elements to res using the binary operator bop. Please refer to Parallel Reduction for more details.
template<typename P, typename I, typename T, typename O, typename U>
void tf:: cuda_transform_reduce_async(P&& p,
I first,
I last,
T* res,
O bop,
U uop,
void* buf)
performs asynchronous parallel reduction over a range of transformed items without an initial value
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| T | value type |
| O | binary operator type |
| U | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| res | pointer to the result |
| bop | binary operator to apply to reduce elements |
| uop | unary operator to apply to transform elements |
| buf | pointer to the temporary buffer |
Asynchronous version of tf::
template<typename P, typename I, typename T, typename O, typename U>
void tf:: cuda_transform_uninitialized_reduce(P&& p,
I first,
I last,
T* res,
O bop,
U uop)
performs parallel reduction over a range of transformed items with an initial value
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| T | value type |
| O | binary operator type |
| U | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| res | pointer to the result |
| bop | binary operator to apply to reduce elements |
| uop | unary operator to apply to transform elements |
Similar to tf::
template<typename P, typename I, typename T, typename O, typename U>
void tf:: cuda_transform_uninitialized_reduce_async(P&& p,
I first,
I last,
T* res,
O bop,
U uop,
void* buf)
performs asynchronous parallel reduction over a range of transformed items with an initial value
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator type |
| T | value type |
| O | binary operator type |
| U | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the range |
| last | iterator to the end of the range |
| res | pointer to the result |
| bop | binary operator to apply to reduce elements |
| uop | unary operator to apply to transform elements |
| buf | pointer to the temporary buffer |
Asynchronous version of tf::
template<typename P, typename T>
unsigned tf:: cuda_scan_buffer_size(unsigned count)
queries the buffer size in bytes needed to call scan kernels
| Template parameters | |
|---|---|
| P | execution policy type |
| T | value type |
| Parameters | |
| count | number of elements to scan |
The function is used to allocate a buffer for calling asynchronous scan. Please refer to Parallel Scan for details.
template<typename P, typename I, typename O, typename C>
void tf:: cuda_inclusive_scan(P&& p,
I first,
I last,
O output,
C op)
performs inclusive scan over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator |
| O | output iterator |
| C | binary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the input range |
| last | iterator to the end of the input range |
| output | iterator to the beginning of the output |
| op | binary operator to apply to scan |
Please refer to Parallel Scan for details.
template<typename P, typename I, typename O, typename C>
void tf:: cuda_inclusive_scan_async(P&& p,
I first,
I last,
O output,
C op,
void* buf)
performs asynchronous inclusive scan over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator |
| O | output iterator |
| C | binary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the input range |
| last | iterator to the end of the input range |
| output | iterator to the beginning of the output range |
| op | binary operator to apply to scan |
| buf | pointer to the temporary buffer |
Please refer to Parallel Scan for details.
template<typename P, typename I, typename O, typename C, typename U>
void tf:: cuda_transform_inclusive_scan(P&& p,
I first,
I last,
O output,
C bop,
U uop)
performs inclusive scan over a range of transformed items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator |
| O | output iterator |
| C | binary operator type |
| U | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the input range |
| last | iterator to the end of the input range |
| output | iterator to the beginning of the output range |
| bop | binary operator to apply to scan |
| uop | unary operator to apply to transform each item before scan |
Please refer to Parallel Scan for details.
template<typename P, typename I, typename O, typename C, typename U>
void tf:: cuda_transform_inclusive_scan_async(P&& p,
I first,
I last,
O output,
C bop,
U uop,
void* buf)
performs asynchronous inclusive scan over a range of transformed items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator |
| O | output iterator |
| C | binary operator type |
| U | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the input range |
| last | iterator to the end of the input range |
| output | iterator to the beginning of the output range |
| bop | binary operator to apply to scan |
| uop | unary operator to apply to transform each item before scan |
| buf | pointer to the temporary buffer |
Please refer to Parallel Scan for details.
template<typename P, typename I, typename O, typename C>
void tf:: cuda_exclusive_scan(P&& p,
I first,
I last,
O output,
C op)
performs exclusive scan over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator |
| O | output iterator |
| C | binary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the input range |
| last | iterator to the end of the input range |
| output | iterator to the beginning of the output range |
| op | binary operator to apply to scan |
Please refer to Parallel Scan for details.
template<typename P, typename I, typename O, typename C>
void tf:: cuda_exclusive_scan_async(P&& p,
I first,
I last,
O output,
C op,
void* buf)
performs asynchronous exclusive scan over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator |
| O | output iterator |
| C | binary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the input range |
| last | iterator to the end of the input range |
| output | iterator to the beginning of the output range |
| op | binary operator to apply to scan |
| buf | pointer to the temporary buffer |
Please refer to Parallel Scan for details.
template<typename P, typename I, typename O, typename C, typename U>
void tf:: cuda_transform_exclusive_scan(P&& p,
I first,
I last,
O output,
C bop,
U uop)
performs exclusive scan over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator |
| O | output iterator |
| C | binary operator type |
| U | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the input range |
| last | iterator to the end of the input range |
| output | iterator to the beginning of the output range |
| bop | binary operator to apply to scan |
| uop | unary operator to apply to transform each item before scan |
Please refer to Parallel Scan for details.
template<typename P, typename I, typename O, typename C, typename U>
void tf:: cuda_transform_exclusive_scan_async(P&& p,
I first,
I last,
O output,
C bop,
U uop,
void* buf)
performs asynchronous exclusive scan over a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| I | input iterator |
| O | output iterator |
| C | binary operator type |
| U | unary operator type |
| Parameters | |
| p | execution policy |
| first | iterator to the beginning of the input range |
| last | iterator to the end of the input range |
| output | iterator to the beginning of the output range |
| bop | binary operator to apply to scan |
| uop | unary operator to apply to transform each item before scan |
| buf | pointer to the temporary buffer |
Please refer to Parallel Scan for details.
template<typename P>
unsigned tf:: cuda_merge_buffer_size(unsigned a_count,
unsigned b_count)
queries the buffer size in bytes needed to call merge kernels
| Template parameters | |
|---|---|
| P | execution polity type |
| Parameters | |
| a_count | number of elements in the first input array |
| b_count | number of elements in the second input array |
The function is used to allocate a buffer for calling asynchronous merge. Please refer to Parallel Merge for details.
template<typename P, typename a_keys_it, typename a_vals_it, typename b_keys_it, typename b_vals_it, typename c_keys_it, typename c_vals_it, typename C>
void tf:: cuda_merge(P&& p,
a_keys_it a_keys_first,
a_vals_it a_vals_first,
a_keys_it a_keys_last,
b_keys_it b_keys_first,
b_vals_it b_vals_first,
b_keys_it b_keys_last,
c_keys_it c_keys_first,
c_vals_it c_vals_first,
C comp)
performs key-value merge over a range of keys and values
| Template parameters | |
|---|---|
| P | execution policy type |
| a_keys_it | first key iterator type |
| a_vals_it | first value iterator type |
| b_keys_it | second key iterator type |
| b_vals_it | second value iterator type |
| c_keys_it | output key iterator type |
| c_vals_it | output value iterator type |
| C | comparator type |
| Parameters | |
| p | execution policy |
| a_keys_first | iterator to the beginning of the first key range |
| a_vals_first | iterator to the beginning of the first value range |
| a_keys_last | iterator to the end of the first key range |
| b_keys_first | iterator to the beginning of the second key range |
| b_vals_first | iterator to the beginning of the second value range |
| b_keys_last | iterator to the end of the second key range |
| c_keys_first | iterator to the beginning of the output key range |
| c_vals_first | iterator to the beginning of the output value range |
| comp | comparator |
Please refer to Parallel Merge for details.
template<typename P, typename a_keys_it, typename a_vals_it, typename b_keys_it, typename b_vals_it, typename c_keys_it, typename c_vals_it, typename C>
void tf:: cuda_merge_async(P&& p,
a_keys_it a_keys_first,
a_vals_it a_vals_first,
a_keys_it a_keys_last,
b_keys_it b_keys_first,
b_vals_it b_vals_first,
b_keys_it b_keys_last,
c_keys_it c_keys_first,
c_vals_it c_vals_first,
C comp,
void* buf)
performs asynchronous key-value merge over a range of keys and values
| Template parameters | |
|---|---|
| P | execution policy type |
| a_keys_it | first key iterator type |
| a_vals_it | first value iterator type |
| b_keys_it | second key iterator type |
| b_vals_it | second value iterator type |
| c_keys_it | output key iterator type |
| c_vals_it | output value iterator type |
| C | comparator type |
| Parameters | |
| p | execution policy |
| a_keys_first | iterator to the beginning of the first key range |
| a_vals_first | iterator to the beginning of the first value range |
| a_keys_last | iterator to the end of the first key range |
| b_keys_first | iterator to the beginning of the second key range |
| b_vals_first | iterator to the beginning of the second value range |
| b_keys_last | iterator to the end of the second key range |
| c_keys_first | iterator to the beginning of the output key range |
| c_vals_first | iterator to the beginning of the output value range |
| comp | comparator |
| buf | pointer to the temporary buffer |
Please refer to Parallel Merge for details.
template<typename P, typename a_keys_it, typename b_keys_it, typename c_keys_it, typename C>
void tf:: cuda_merge(P&& p,
a_keys_it a_keys_first,
a_keys_it a_keys_last,
b_keys_it b_keys_first,
b_keys_it b_keys_last,
c_keys_it c_keys_first,
C comp)
performs key-only merge over a range of keys
| Template parameters | |
|---|---|
| P | execution policy type |
| a_keys_it | first key iterator type |
| b_keys_it | second key iterator type |
| c_keys_it | output key iterator type |
| C | comparator type |
| Parameters | |
| p | execution policy |
| a_keys_first | iterator to the beginning of the first key range |
| a_keys_last | iterator to the end of the first key range |
| b_keys_first | iterator to the beginning of the second key range |
| b_keys_last | iterator to the end of the second key range |
| c_keys_first | iterator to the beginning of the output key range |
| comp | comparator |
Please refer to Parallel Merge for details.
template<typename P, typename a_keys_it, typename b_keys_it, typename c_keys_it, typename C>
void tf:: cuda_merge_async(P&& p,
a_keys_it a_keys_first,
a_keys_it a_keys_last,
b_keys_it b_keys_first,
b_keys_it b_keys_last,
c_keys_it c_keys_first,
C comp,
void* buf)
performs asynchronous key-only merge over a range of keys
| Template parameters | |
|---|---|
| P | execution policy type |
| a_keys_it | first key iterator type |
| b_keys_it | second key iterator type |
| c_keys_it | output key iterator type |
| C | comparator type |
| Parameters | |
| p | execution policy |
| a_keys_first | iterator to the beginning of the first key range |
| a_keys_last | iterator to the end of the first key range |
| b_keys_first | iterator to the beginning of the second key range |
| b_keys_last | iterator to the end of the second key range |
| c_keys_first | iterator to the beginning of the output key range |
| comp | comparator |
| buf | pointer to the temporary buffer |
Please refer to Parallel Merge for details.
template<typename P, typename K, typename V = cudaEmpty>
unsigned tf:: cuda_sort_buffer_size(unsigned count)
queries the buffer size in bytes needed to call sort kernels for the given number of elements
| Template parameters | |
|---|---|
| P | execution policy type |
| K | key type |
| V | value type (default tf::cudaEmpty) |
| Parameters | |
| count | number of keys/values to sort |
The function is used to allocate a buffer for calling asynchronous sort. Please refer to Parallel Sort for details.
template<typename P, typename K_it, typename V_it, typename C>
void tf:: cuda_sort(P&& p,
K_it k_first,
K_it k_last,
V_it v_first,
C comp)
performs key-value sort on a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| K_it | key iterator type |
| V_it | value iterator type |
| C | comparator type |
| Parameters | |
| p | execution policy |
| k_first | iterator to the beginning of the key range |
| k_last | iterator to the end of the key range |
| v_first | iterator to the beginning of the value range |
| comp | binary comparator |
Please refer to Parallel Sort for details.
template<typename P, typename K_it, typename V_it, typename C>
void tf:: cuda_sort_async(P&& p,
K_it k_first,
K_it k_last,
V_it v_first,
C comp,
void* buf)
performs asynchronous key-value sort on a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| K_it | key iterator type |
| V_it | value iterator type |
| C | comparator type |
| Parameters | |
| p | execution policy |
| k_first | iterator to the beginning of the key range |
| k_last | iterator to the end of the key range |
| v_first | iterator to the beginning of the value range |
| comp | binary comparator |
| buf | pointer to the temporary buffer |
Please refer to Parallel Sort for details.
template<typename P, typename K_it, typename C>
void tf:: cuda_sort(P&& p,
K_it k_first,
K_it k_last,
C comp)
performs key-only sort on a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| K_it | key iterator type |
| C | comparator type |
| Parameters | |
| p | execution policy |
| k_first | iterator to the beginning of the key range |
| k_last | iterator to the end of the key range |
| comp | binary comparator |
Please refer to Parallel Sort for details.
template<typename P, typename K_it, typename C>
void tf:: cuda_sort_async(P&& p,
K_it k_first,
K_it k_last,
C comp,
void* buf)
performs asynchronous key-only sort on a range of items
| Template parameters | |
|---|---|
| P | execution policy type |
| K_it | key iterator type |
| C | comparator type |
| Parameters | |
| p | execution policy |
| k_first | iterator to the beginning of the key range |
| k_last | iterator to the end of the key range |
| comp | binary comparator |
| buf | pointer to the temporary buffer |
Please refer to Parallel Sort for details.
Variable documentation
template<typename C>
bool tf:: is_static_task_v constexpr
determines if a callable is a static task
A static task is a callable object constructible from std::function<void()>.
template<typename C>
bool tf:: is_dynamic_task_v constexpr
determines if a callable is a dynamic task
A dynamic task is a callable object constructible from std::function<void(Subflow&)>.
template<typename C>
bool tf:: is_condition_task_v constexpr
determines if a callable is a condition task
A condition task is a callable object constructible from std::function<int()>.
template<typename C>
bool tf:: is_cudaflow_task_v constexpr
determines if a callable is a cudaFlow task
A cudaFlow task is a callable object constructible from std::function<void(tf::cudaFlow&)> or std::function<void(tf::cudaFlowCapturer&)>.
template<typename C>
bool tf:: is_syclflow_task_v constexpr
determines if a callable is a syclFlow task
A syclFlow task is a callable object constructible from std::function<void(tf::syclFlow&)>.