How do I get the least used character in the text file ?

a= textread('GreatExpectations.txt','%c');
[m]=length(a);
most_used_letter=char(mode(0+a))

1 Kommentar

Jan
Jan am 26 Apr. 2015
Is in 'aba' the 'b' or the 'c' the least used character?

Melden Sie sich an, um zu kommentieren.

 Akzeptierte Antwort

Mohammad Abouali
Mohammad Abouali am 25 Apr. 2015
Bearbeitet: Mohammad Abouali am 26 Apr. 2015
%Sample text; equivalent to your a
txt='this is a sample text.';
uniqueChars=unique(txt(isletter(txt)));
charCount=arrayfun(@(c) sum(txt==c), uniqueChars);
% Now printing results
fprintf('Char count\n');
fprintf(' %c %d\n',[uniqueChars;charCount])
fprintf('\nThe least Used Characters are:\n')
fprintf('%c\n',uniqueChars(charCount==min(charCount)))
fprintf('\nThe most Used Characters are:\n')
fprintf('%c\n',uniqueChars(charCount==max(charCount)))
Once you run it you well get this:
Char count
a 2
e 2
h 1
i 2
l 1
m 1
p 1
s 3
t 3
x 1
The least Used Characters are:
h
l
m
p
x
The most Used Characters are:
s
t

2 Kommentare

Tony
Tony am 26 Apr. 2015
Allah Kheleek , thank you!!
You are welcome.
In larger texts, when you are sure that all characters are used you can remove uniqueChars=unique(txt(isletter(txt))); and then modify the next line form
charCount=arrayfun(@(c) sum(txt==c), uniqueChars);
to
charCount=arrayfun(@(c) sum(txt==c), char([65:90 97:122]));
You have to be sure that all characters are used though. Otherwise, those characters that are not used would be returned as the least used characters.
Also if upper/lower case are not important; then once you are done reading the text file do this:
txt=lower(txt);
and in this case if you decided to use the second form you need char(97:122) instead of char([65:90 97:122]);
If this has answered your question, I appreciate it if you accept the answers.

Melden Sie sich an, um zu kommentieren.

Weitere Antworten (1)

per isakson
per isakson am 26 Apr. 2015
Bearbeitet: per isakson am 26 Apr. 2015
Here is a code based on histc
ffs = 'GreatExpectations.txt';
fid = fopen( ffs );
buf = fscanf( fid, '%1c' );
fclose( fid );
letters only
ascii = double( upper( buf ) );
ascii( ascii<double('A') | ascii>double('Z') ) = [];
ascii_ranges = ( double('A') : double('Z') );
[ ascii_counts, ~ ] = histc( ascii, ascii_ranges );
bar( ascii_ranges, ascii_counts, 'histc' )
min_val = min( ascii_counts );
ix_min = find( ascii_counts == min_val );
fprintf( 'The letters, %s, each appears %d time(s)\n' ...
, char('A'+ix_min-1), min_val )
it prints
The letters, JQZ, each appears 1 time(s)
&nbsp
Regarding char('A'+ix_min-1) see 'ABC'-'A' Considered Harmful

Tags

Gefragt:

am 25 Apr. 2015

Kommentiert:

Jan
am 26 Apr. 2015

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by