repeat large number of vector elements
1 次查看(过去 30 天)
显示 更早的评论
Hi, I have a vector storing only unique values: v = (0, 1, 2), and another vector about frequency of the unique values: c = (3, 2, 4). Now I want to create a new vector with repeated values as follows:
v2 = (0, 0, 0, 1, 1, 2, 2, 2, 2)
Since both vectors v and c are very large, how to efficiently derive v2 in matlab. Many thanks.
1 个评论
Jan
2013-2-20
Please specify "very large" explicitly: Some users tell 1000 elements large, others need billions for this term.
回答(5 个)
Azzi Abdelmalek
2013-2-20
编辑:Azzi Abdelmalek
2013-2-21
v=[0 1 2];
c=[3,2,4];
out=cell2mat(arrayfun(@(x) repmat(v(x),1,c(x)),1:numel(v),'un',0))
EDIT
i1=0;
out=zeros(1,sum(c));
for k=1:numel(v)
i0=i1+1;
i1=i0+c(k)-1;
out(i0:i1)=v(k)*ones(1,c(k));
end
5 个评论
Youssef Khmou
2013-2-20
编辑:Youssef Khmou
2013-2-20
for i=1:length(c)
V{i}=(i-1)*ones(c(i),1);
end
f=cell2mat(V(:))'
0 个评论
Jan
2013-2-21
编辑:Jan
2013-2-21
Some timings (R2009a/64/Win7/Core2Duo):
x = rand(1, 10000);
n = randi([0,10], 1, 10000);
tic; for k=1:100; r=runlength(x,n); end; toc
% With runlength() is a function with the following contents:
% Azzi's ARRAYFUN approach:
r = cell2mat(arrayfun(@(k) repmat(x(k), 1, n(k)), 1:numel(x), 'un', 0))
Elapsed time is 48.976947 seconds.
% Azzi's approach from the comment section: [EDITED start]
Elapsed time is 2.674559 seconds.
% Azzi's approach from the comment section with slight modification:
i1 = 0;
result = zeros(1, sum(c));
for k = 1:numel(v)
i0 = i1+1;
i1 = i0 + c(k) - 1;
result(i0:i1) = v(k); % Without "*ones(1,c(k))" !!
end
Elapsed time is 1.656987 seconds. [EDITED end]
% Youssef' approach:
V = cell(1, length(n));
for k = 1:length(n)
V{k} = x(k) * ones(n(k), 1); % With x(k) instead of (i-1)
end
r = cell2mat(V(:))'
Elapsed time is 1.422626 seconds
% The CUMSUM method from my former answer:
Elapsed time is 0.083854 seconds.
% An equivalent MEX implementation:
Elapsed time is 0.025775 seconds.
Conclusions:
- ARRAYFUN and the anonymous function is not efficient compared to a simple loop. Therefore I suggest to avoid this combination strictly.
- Collecting the single parts in a cell is a good idea. Then CELL2MAT can be replaced by the faster FEX: Cell2Vec : 0.97 sec
- Creating the large temporary vectors cumsum(x) and index seem to be not as bad as I was afraid. The C-MEX is not a dramatic enhancement anymore.
2 个评论
Jan
2013-2-21
编辑:Jan
2013-2-21
Yes, Azzi, I haven't seen it. It is much faster, especially when the multiplication with ONES is omitted. A good improvement compared to ARRAYFUN!
I do not understand why this is slower than collecting the partial vectors in a cell array at first. It looks much more efficient.
irina mihai
2015-12-6
Jan, Azzi Thanks for your help ! I used the cumsum function and the v(index) idea which i didn't know. I have a very large database (over 10^6 lines) and I was looking for something like this. Brilliant !
另请参阅
类别
在 Help Center 和 File Exchange 中查找有关 Loops and Conditional Statements 的更多信息
Community Treasure Hunt
Find the treasures in MATLAB Central and discover how the community can help you!
Start Hunting!