How to make textscan robust against non-matching lines?

Question

Joan Vazquez 2021-4-8

0
链接

此问题的直接链接

https://ww2.mathworks.cn/matlabcentral/answers/796102-how-to-make-textscan-robust-against-non-matching-lines

评论： Stephen23 2021-4-9

data.txt

I have files with lines that I want to parse, preferably with textscan. In between those lines, there may be lines to be skipped (unpredictable format and abundance, but definetely new lines). What is the best way to deal with it?

E.g. for the data in attachment, this will stop outputiing #HELLOMATHWORKS messages after line 4.

fid = fopen('data.txt');
out = textscan(fid,'#HELLOMATHWORKS,%[^,],%n');
fclose(fid);

This is a MWE out of a large code base.

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

请先登录，再进行评论。

请先登录，再回答此问题。

Answer 1

Stephen23 2021-4-8

0
链接

此回答的直接链接

https://ww2.mathworks.cn/matlabcentral/answers/796102-how-to-make-textscan-robust-against-non-matching-lines#answer_670412

编辑：Stephen23 2021-4-9

在 MATLAB Online 中打开

data.txt

str = fileread('data.txt');
tkn = regexp(str,'#HELLOMATHWORKS,([^,]+),(\S+)','tokens');
tkn = vertcat(tkn{:})
tkn = 6×2 cell array
    {'COM1'}    {'2146'}
    {'COM1'}    {'2147'}
    {'COM1'}    {'2148'}
    {'COM1'}    {'2149'}
    {'COM1'}    {'2150'}
    {'COM1'}    {'2151'}
vec = str2double(tkn(:,2))
vec = 6×1
        2146
        2147
        2148
        2149
        2150
        2151

2 个评论
显示无隐藏无

Joan Vazquez 2021-4-8

编辑：Joan Vazquez 2021-4-8

This does not produce the same output as my code:

tmp =

1×2 cell array

{6×1 cell} {6×1 double}

(Actually my messages have many more fields, this was just a MWE with 2... I have many similar functions using texscan to parse messages and I wanted to avoid refactoring them)

It is a good idea to work directly with regular expressions, but it seems that the formatSpec input parameter of textscan is not just any regular expression, it is more limited...

Anyway, It's OK for the moment, I'll accept the answer, thanks

Stephen23 2021-4-9

@Joan Vazquez: I presume that the text #HELLOMATHWORKS is not what is actually in your file. If the actual text contains some unique character that does not exist anywhere else in the file, you might be able to leverage the LineEnding/EndOfLine option to achieve the goal of reading the file data using textscan.

请先登录，再进行评论。

Answer 2

Joan Vazquez 2021-4-8

1
链接

此回答的直接链接

https://ww2.mathworks.cn/matlabcentral/answers/796102-how-to-make-textscan-robust-against-non-matching-lines#answer_670177

在 MATLAB Online 中打开

This works, but it does not seem the best solution...Ideally, I would tell textscan "skip everything until a new line starts with #HELLOMATHWORKS"

filetext = fileread('data.txt');
expr = '[^\n]*#HELLOMATHWORKS[^\n]*';
% Find and return all lines that contain the text '#HELLOMATHWORKS'.
matches = regexp(filetext,expr,'match');
% Make it a 1xN char to feed textscan
goodlines = sprintf('%s\n', matches{:});
tmp = textscan(goodlines,'#HELLOMATHWORKS,%[^,],%n');

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

请先登录，再进行评论。

How to make textscan robust against non-matching lines?

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

采纳的回答

2 个评论
显示无隐藏无

更多回答（1 个）

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

另请参阅

类别

标签

产品

版本

Community Treasure Hunt

How to make textscan robust against non-matching lines?

0 个评论 显示 -2更早的评论隐藏 -2更早的评论

采纳的回答

2 个评论 显示 无隐藏 无

更多回答（1 个）

0 个评论 显示 -2更早的评论隐藏 -2更早的评论

另请参阅

类别

标签

产品

版本

Community Treasure Hunt

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

2 个评论
显示无隐藏无

0 个评论
显示 -2更早的评论隐藏 -2更早的评论