gpuDevice command very slow
显示 更早的评论
I am running CUDA kernels using the parallel computing toolbox and r2012a. Recently upgraded to a 600 series (Kepler) gpu. To setup the CUDA kernel we extract the maximum threads per block using: gpu_han=gpuDevice(1); k = parallel.gpu.CUDAKernel('gpu_tfm_linear_arb.ptx', gpu_tfm_linear_arb.cu'); k.ThreadBlockSize = gpu_han.MaxThreadsPerBlock;
This is now executing very slowly (order 2mins). If I specify the threadblocksize manually to the max of the card (1024 in this case), it executes in 0.1 s.
This used to run quickly with a 400 series card. Any help gratefully received
采纳的回答
更多回答(2 个)
Andrei Pokrovsky
2016-9-15
编辑:Andrei Pokrovsky
2016-9-15
3 个投票
Try setting these env vars:
export CUDA_CACHE_MAXSIZE=2147483647
export CUDA_CACHE_DISABLE=0
This cured the problem on my GTX1080.
https://devblogs.nvidia.com/parallelforall/cuda-pro-tip-understand-fat-binaries-jit-caching/
Anthony
2013-6-17
0 个投票
2 个评论
Edric Ellis
2013-6-18
The cache is not stored where the program lives, this page from NVIDIA has all the gory details, including this:
- on Windows, %APPDATA%\NVIDIA\ComputeCache,
- on MacOS, $HOME/Library/Application\ Support/NVIDIA/ComputeCache,
- on Linux, ~/.nv/ComputeCache
Anthony
2013-7-12
类别
在 帮助中心 和 File Exchange 中查找有关 GPU Computing 的更多信息
Community Treasure Hunt
Find the treasures in MATLAB Central and discover how the community can help you!
Start Hunting!