How to search a string with multiple rows for text?

Question

0 Stimmen

Hello, After running seq=getgenpept('NP_036795'); . I want to search seq.Features for some text value 'Protein' . I have been unable to find the correct function to search a string with multiple rows.

Running: k=strfind(seq.Features,'Protein') results with "Error using strfind. Input strings must have one row."

Any thoughts? Best, Joe

3 Kommentare
1 älteren Kommentar anzeigen 1 älteren Kommentar ausblenden

per isakson am 27 Mär. 2015

Bearbeitet: per isakson am 27 Mär. 2015

In MATLAB Online öffnen

Excerpt from doc of getgenpept

Features: [40x64 char]

strfind cannot handle multi-row character arrays.

What does this array of characters look like? &nbsp BTW: it's allowed to use for-loops.

Luuk van Oosten am 28 Mär. 2015

Looks like the pic below.

What kind of info are you trying to extract from 'Protein'?

Melden Sie sich an, um zu kommentieren.

Melden Sie sich an, um diese Frage zu beantworten.

Follow Question

Answer 1

per isakson am 28 Mär. 2015

Bearbeitet: per isakson am 29 Mär. 2015

In MATLAB Online öffnen

0 Stimmen

I guess this block of characters is easier to read on screen than to read and parse automatically. "find the correct function" I don't think there is the function; a small program is needed. Anyhow, the script below creates a structure, sas, which is a start

    %%Create test data. (The OCR-program missed most of the underscore.)
    buf = { 'source   1..116                                                '
            '         /organism="Rattus norvegicus"                         '
            '         /dbxref="taxon: 10116^                                '
            '         /chromosome=^10^                                      '
            '         /map="10824"                                          '
            'Protein  1..116                                                '
            '         /product="vesicle-associated membrane protein 2^      '
            '         /note="VAMP-2; synaptobrevin-2; Synaptobrevin 2       '
            '         (vesicle-associated membrane protein VAMP-2);         '
            '         Vesicle-associated membrane protein (synaptobrevin 2)"'
            '         /calculated mol wt=12560                              '
            'Region   28..101                                               '
            '         /region name="Synaptobrevin"                          '
            '         /note="Synaptobrevin; pfam00957"                      '
            '         /dbxref="CDD:250253"                                  '
            'Site     95..114                                               '
            '         /site type="transmembrane region"                     '
            '         /inference="non-experimental evidence, no additional  '
            '         details recorded"                                     '
            '         /note="propagated from UniProt./Swiss-Prot (P63045.2).'
            'CDS      1..116                                                '
            '         /gene="Vamp2^                                         '
            '         /gene synonym="RATVAMPB; RATVAMPIR; SYS; Syb2^        '
            '         /coded by="NM 012663.2:83..433"                       '
            '         /dbxref="GeneID:24803^                                '
            '         /dbxref="RGD:3949"                                    '};
    str_array = char( buf );
    %%read and parse
    for rr = 1 : size( str_array, 1 )
        % search rows starting with a word and followed by digits, two ".", digits
        buf = regexp( str_array(rr,:), '^(\w+)\s+(\d+\.{2}\d+)', 'tokens' );
        if not( isempty( buf ) )
            field_name = buf{1}{1};
            sas.(field_name) = buf{1}(2); 
        else
            sas.(field_name) = cat( 1, sas.(field_name)         ...
                                ,   strtrim( str_array(rr,:) )  );
        end
    end

The structure, sas, has one field for each sub-group

    >> sas
    sas = 
         source: {5x1 cell}
        Protein: {6x1 cell}
         Region: {4x1 cell}
           Site: {4x1 cell}
            CDS: {6x1 cell}
    >> sas.Protein
    ans = 
        '1..116'
        '/product="vesicle-associated membrane protein 2^'
        '/note="VAMP-2; synaptobrevin-2; Synaptobrevin 2'
        '(vesicle-associated membrane protein VAMP-2);'
        'Vesicle-associated membrane protein (synaptobrevin 2)"'
        '/calculated mol wt=12560'
    >> char( sas.Protein )
    ans =
    1..116                                                
    /product="vesicle-associated membrane protein 2^      
    /note="VAMP-2; synaptobrevin-2; Synaptobrevin 2       
    (vesicle-associated membrane protein VAMP-2);         
    Vesicle-associated membrane protein (synaptobrevin 2)"
    /calculated mol wt=12560                              
    >>

Next step is to parse the sub-blocks.

0 Kommentare
-2 ältere Kommentare anzeigen -2 ältere Kommentare ausblenden

Melden Sie sich an, um zu kommentieren.

How to search a string with multiple rows for text?

3 Kommentare
1 älteren Kommentar anzeigen 1 älteren Kommentar ausblenden

Antworten (1)

0 Kommentare
-2 ältere Kommentare anzeigen -2 ältere Kommentare ausblenden

Kategorien

Produkte

Tags

Community Treasure Hunt

How to search a string with multiple rows for text?

3 Kommentare 1 älteren Kommentar anzeigen 1 älteren Kommentar ausblenden

Antworten (1)

0 Kommentare -2 ältere Kommentare anzeigen -2 ältere Kommentare ausblenden

Kategorien

Produkte

Tags

Siehe auch

Community Treasure Hunt

3 Kommentare
1 älteren Kommentar anzeigen 1 älteren Kommentar ausblenden

0 Kommentare
-2 ältere Kommentare anzeigen -2 ältere Kommentare ausblenden