Remove duplicate rows in CSV file

hello dear mathworkers,
I have a dataset consist of approximatlly 4 millions records, and i want to remove the duplicated rows or records, can any one help me with the way, i am using matlab 2018a . thanks in advance

7 个评论

Upload a sample file.
You can use unique with 'rows' option
dear madhan ravi this is a sample of the whol data
@Mohammad Alsajri: You mentioned that the rows are duplicated, but in you sample sheet, I tried searching for duplicated data, but I could not find duplicated data..
For example, if I filtered the data for value '215' (at column E), still I dont see the duplicated value at other columns like column 'W' and 'X'..
Capture.JPG
So let us know.. what is mean by duplicate entries and how to consider it..
@Shameer Parmar duplicate means the entire row match another row for all columns , this is just sample it is 4 million records so of cours there is a duplicated rows
Mohammed: Alex's solution should have solved your problem.
thanks for help guys

请先登录,再进行评论。

 采纳的回答

Since all is numeric data, you can use:
data = xlsread('kdd.xlsx');
datanew = unique(data,'rows');

2 个评论

This is not working, because non of data is similar.. I dont find duplicate entries in this sheet provided by Mohammad Alsajri..
using your command, the 'data' and 'datanew' both are getting exact same..
This code works!
I guess the excel provided by Mohammad is just a small portion of the dataset (4 million of rows).

请先登录,再进行评论。

更多回答(0 个)

类别

在 帮助中心 和 File Exchange 中查找有关 Data Import and Analysis 的更多信息

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by