Why is vectorization faster than the parallel computing？

Question

0 个投票

I am trying to speed my code, which processing some big geometry surface. From scanning on the website, I found 2 way to optimize my code. One is vectorization, the other is using parallel computing toolbox. My computer has 16 cores, but with relative low speed, 2.0 ghz. As my experience, I found vectorization is always faster than parallel computing. I just wonder the reason about it. Does Matlab build-in function could do the vectorization on many "virtual small processors", which is much more than the computer cores? Like GPU or something else? I want to know a little about the machinism under it.

Thank you very much for any hints and explainations.

2 个评论
显示无隐藏无

Jan 2019-9-4

The more specific a question is, the easier is an answer. If you post some code, it is possible to explain, what happens. Asking aboput the general mechanism demands for an exhaustive answer, which will most likely do not match your point.

Adam 2019-9-4

Vectorisation takes advantage of the fact that the same (usually fairly simple, at a component level) operation is being performed on many elements of the array, which it can process very efficiently at the low-level.

Parallelisation has data copying overheads and many other considerations, as discussed by Jan in his answer below.

Where vectorisation is possibly it is usually preferable (faster) to parallelisation.

请先登录，再进行评论。

请先登录，再回答此问题。

Follow Question

Answer 1

Jan 2019-9-4

1 个投票

It depends on the problem. Parallelization is not trivial. If you use e.g. 16 cores and write the results in neighboring elements of a UINT8 vector, you get a collision in the cache-line. As result the total computing time can exceed the time of a serial code, because the threads are waiting for eachother. Such collisions can occur in other resources also, e.g. if the memory is exhausted and expensive disk caching is used, or if data are requested through a network.

Many Matlab functions are mutli-threaded, e.g. sum: For large inputs Matlab computes the sum in several parts using different threads. For a 1e5 x 1e5 matrix all cores are used (most likely). Computing this by parallelization in a parfor loop is less efficient, because there is some overhead for starting the threads. The multi-threaded functions are written such, that resource collisions are avoided (at least in most cases. In some cases, e.g. logical indexing or cell2mat there is some potential for improvements).

So before you start to parallelize a function, check if it uses many cores already in the sequential version. Then starting mutliple threads of a parpool on a single machine will not improve the efficiency.

1 个评论
显示 -1更早的评论隐藏 -1更早的评论

Sterling Baird 2020-10-26

How does one check if many cores are already used in the sequential version?

For example:

mldivide ('\')
pdist (Statistics & Machine Learning Toolbox)
dsearchn
fitrgp (without hyperoptimization)

请先登录，再进行评论。

Why is vectorization faster than the parallel computing？

2 个评论
显示无隐藏无

采纳的回答

1 个评论
显示 -1更早的评论隐藏 -1更早的评论

更多回答（0 个）

类别

产品

版本

标签

Community Treasure Hunt

Why is vectorization faster than the parallel computing？

2 个评论 显示 无 隐藏 无

采纳的回答

1 个评论 显示 -1更早的评论 隐藏 -1更早的评论

更多回答（0 个）

类别

产品

版本

标签

另请参阅

Community Treasure Hunt

2 个评论
显示无隐藏无

1 个评论
显示 -1更早的评论隐藏 -1更早的评论