<p>how could I test a string against only valid characters like letters a-z?...</p> <pre class="prettyprint"><code>string name; cout << "Enter your name" cin >> name; string letters = "qwertyuiopasdfghjklzxcvbnm"; string::iterator it; for(it = name.begin(); it = name.end(); it++) { size_t found = letters.find(it); } </code></pre>

<p>First, using <code>std::cin >> name</code> will fail if the user enters <code>John Smith</code> because <code>>></code> splits input on whitespace characters. You should use <code>std::getline()</code> to get the name:</p> <pre class="prettyprint"><code>std::getline(std::cin, name); </code></pre> <h3>Here we go…</h3> <p>There are a number of ways to check that a string contains only alphabetic characters. The simplest is probably <code>s.find_first_not_of(t)</code>, which returns the index of the first character in <code>s</code> that is not in <code>t</code>:</p> <pre class="prettyprint"><code>bool contains_non_alpha = name.find_first_not_of("abcdefghijklmnopqrstuvwxyz") != std::string::npos; </code></pre> <p>That rapidly becomes cumbersome, however. To also match uppercase alphabetic characters, you’d have to add 26 more characters to that string! Instead, you may want to use a combination of <code>find_if</code> from the <code><algorithm></code> header and <code>std::isalpha</code> from <code><cctype></code>:</p> <pre class="prettyprint"><code>#include <algorithm> #include <cctype> struct non_alpha { bool operator()(char c) { return !std::isalpha(c); } }; bool contains_non_alpha = std::find_if(name.begin(), name.end(), non_alpha()) != name.end(); </code></pre> <p><code>find_if</code> searches a range for a value that matches a predicate, in this case a functor <code>non_alpha</code> that returns whether its argument is a non-alphabetic character. If <code>find_if(name.begin(), name.end(), ...)</code> returns <code>name.end()</code>, then no match was found.</p> <h3>But there’s more!</h3> <p>To do this as a one-liner, you can use the adaptors from the <code><functional></code> header:</p> <pre class="prettyprint"><code>#include <algorithm> #include <cctype> #include <functional> bool contains_non_alpha = std::find_if(name.begin(), name.end(), std::not1(std::ptr_fun((int(*)(int))std::isalpha))) != name.end(); </code></pre> <p>The <code>std::not1</code> produces a function object that returns the logical inverse of its input; by supplying a pointer to a function with <code>std::ptr_fun(...)</code>, we can tell <code>std::not1</code> to produce the logical inverse of <code>std::isalpha</code>. The cast <code>(int(*)(int))</code> is there to select the overload of <code>std::isalpha</code> which takes an <code>int</code> (treated as a character) and returns an <code>int</code> (treated as a Boolean).</p> <p>Or, if you can use a C++11 compiler, using a lambda cleans this up a lot:</p> <pre class="prettyprint"><code>#include <cctype> bool contains_non_alpha = std::find_if(name.begin(), name.end(), [](char c) { return !std::isalpha(c); }) != name.end(); </code></pre> <p><code>[](char c) -> bool { ... }</code> denotes a function that accepts a character and returns a <code>bool</code>. In our case we can omit the <code>-> bool</code> return type because the function body consists of only a <code>return</code> statement. This works just the same as the previous examples, except that the function object can be specified much more succinctly.</p> <h3>And (almost) finally…</h3> <p>In C++11 you can also use a regular expression to perform the match:</p> <pre class="prettyprint"><code>#include <regex> bool contains_non_alpha = !std::regex_match(name, std::regex("^[A-Za-z]+$")); </code></pre> <h3>But of course…</h3> <p>None of these solutions addresses the issue of locale or character encoding! For a locale-independent version of <code>isalpha()</code>, you’d need to use the C++ header <code><locale></code>:</p> <pre class="prettyprint"><code>#include <locale> bool isalpha(char c) { std::locale locale; // Default locale. return std::use_facet<std::ctype<char> >(locale).is(std::ctype<char>::alpha, c); } </code></pre> <p>Ideally we would use <code>char32_t</code>, but <code>ctype</code> doesn’t seem to be able to classify it, so we’re stuck with <code>char</code>. Lucky for us we can dance around the issue of locale entirely, because you’re probably only interested in English letters. There’s a handy header-only library called UTF8-CPP that will let us do what we need to do in a more encoding-safe way. First we define our version of <code>isalpha()</code> that uses UTF-32 code points:</p> <pre class="prettyprint"><code>bool isalpha(uint32_t c) { return (c >= 0x0041 && c <= 0x005A) || (c >= 0x0061 && c <= 0x007A); } </code></pre> <p>Then we can use the <code>utf8::iterator</code> adaptor to adapt the <code>basic_string::iterator</code> from octets into UTF-32 code points:</p> <pre class="prettyprint"><code>#include <utf8.h> bool contains_non_alpha = std::find_if(utf8::iterator(name.begin(), name.begin(), name.end()), utf8::iterator(name.end(), name.begin(), name.end()), [](uint32_t c) { return !isalpha(c); }) != name.end(); </code></pre> <p>For slightly better performance at the cost of safety, you can use <code>utf8::unchecked::iterator</code>:</p> <pre class="prettyprint"><code>#include <utf8.h> bool contains_non_alpha = std::find_if(utf8::unchecked::iterator(name.begin()), utf8::unchecked::iterator(name.end()), [](uint32_t c) { return !isalpha(c); }) != name.end(); </code></pre> <p>This will fail on some invalid input.</p> <p>Using UTF8-CPP in this way assumes that the host encoding is UTF-8, or a compatible encoding such as ASCII. In theory this is still an imperfect solution, but in practice it will work on the vast majority of platforms.</p> <p>I hope this answer is finally complete!</p>

how to test a string for letters only

Tags:

how could I test a string against only valid characters like letters a-z?...

string name;

cout << "Enter your name"
cin >> name;

string letters = "qwertyuiopasdfghjklzxcvbnm";

string::iterator it;

for(it = name.begin(); it = name.end(); it++)
{
  size_t found = letters.find(it);
}

919

asked Sep 30 '11 22:09

miatech

1 Answers

First, using std::cin >> name will fail if the user enters John Smith because >> splits input on whitespace characters. You should use std::getline() to get the name:

std::getline(std::cin, name);

Here we go…

There are a number of ways to check that a string contains only alphabetic characters. The simplest is probably s.find_first_not_of(t), which returns the index of the first character in s that is not in t:

bool contains_non_alpha
    = name.find_first_not_of("abcdefghijklmnopqrstuvwxyz") != std::string::npos;

That rapidly becomes cumbersome, however. To also match uppercase alphabetic characters, you’d have to add 26 more characters to that string! Instead, you may want to use a combination of find_if from the <algorithm> header and std::isalpha from <cctype>:

#include <algorithm>
#include <cctype>

struct non_alpha {
    bool operator()(char c) {
        return !std::isalpha(c);
    }
};

bool contains_non_alpha
    = std::find_if(name.begin(), name.end(), non_alpha()) != name.end();

find_if searches a range for a value that matches a predicate, in this case a functor non_alpha that returns whether its argument is a non-alphabetic character. If find_if(name.begin(), name.end(), ...) returns name.end(), then no match was found.

But there’s more!

To do this as a one-liner, you can use the adaptors from the <functional> header:

#include <algorithm>
#include <cctype>
#include <functional>

bool contains_non_alpha
    = std::find_if(name.begin(), name.end(),
                   std::not1(std::ptr_fun((int(*)(int))std::isalpha))) != name.end();

The std::not1 produces a function object that returns the logical inverse of its input; by supplying a pointer to a function with std::ptr_fun(...), we can tell std::not1 to produce the logical inverse of std::isalpha. The cast (int(*)(int)) is there to select the overload of std::isalpha which takes an int (treated as a character) and returns an int (treated as a Boolean).

Or, if you can use a C++11 compiler, using a lambda cleans this up a lot:

#include <cctype>

bool contains_non_alpha
    = std::find_if(name.begin(), name.end(),
                   [](char c) { return !std::isalpha(c); }) != name.end();

[](char c) -> bool { ... } denotes a function that accepts a character and returns a bool. In our case we can omit the -> bool return type because the function body consists of only a return statement. This works just the same as the previous examples, except that the function object can be specified much more succinctly.

And (almost) finally…

In C++11 you can also use a regular expression to perform the match:

#include <regex>

bool contains_non_alpha
    = !std::regex_match(name, std::regex("^[A-Za-z]+$"));

But of course…

None of these solutions addresses the issue of locale or character encoding! For a locale-independent version of isalpha(), you’d need to use the C++ header <locale>:

#include <locale>

bool isalpha(char c) {
    std::locale locale; // Default locale.
    return std::use_facet<std::ctype<char> >(locale).is(std::ctype<char>::alpha, c);
}

Ideally we would use char32_t, but ctype doesn’t seem to be able to classify it, so we’re stuck with char. Lucky for us we can dance around the issue of locale entirely, because you’re probably only interested in English letters. There’s a handy header-only library called UTF8-CPP that will let us do what we need to do in a more encoding-safe way. First we define our version of isalpha() that uses UTF-32 code points:

bool isalpha(uint32_t c) {
    return (c >= 0x0041 && c <= 0x005A)
        || (c >= 0x0061 && c <= 0x007A);
}

Then we can use the utf8::iterator adaptor to adapt the basic_string::iterator from octets into UTF-32 code points:

#include <utf8.h>

bool contains_non_alpha
    = std::find_if(utf8::iterator(name.begin(), name.begin(), name.end()),
                   utf8::iterator(name.end(), name.begin(), name.end()),
                   [](uint32_t c) { return !isalpha(c); }) != name.end();

For slightly better performance at the cost of safety, you can use utf8::unchecked::iterator:

#include <utf8.h>

bool contains_non_alpha
    = std::find_if(utf8::unchecked::iterator(name.begin()),
                   utf8::unchecked::iterator(name.end()),
                   [](uint32_t c) { return !isalpha(c); }) != name.end();

This will fail on some invalid input.

Using UTF8-CPP in this way assumes that the host encoding is UTF-8, or a compatible encoding such as ASCII. In theory this is still an imperfect solution, but in practice it will work on the vast majority of platforms.

I hope this answer is finally complete!

answered Oct 02 '22 07:10

Jon Purdy

Related questions
                            
                                Eclipse, adb, and ddms not detecting Android Emulator
                            
                                Rounding up to the nearest 0.05 in JavaScript
                            
                                Catch Segmentation fault in c++
                            
                                What is the left-child, right-sibling representation of a tree? Why would you use it?
                            
                                Attempt to insert non-property value Objective C
                            
                                Sort array on key value
                            
                                T-SQL Get File Extension Name from a Column
                            
                                jQuery load first 3 elements, click "load more" to display next 5 elements
                            
                                Time requests in NodeJS/Express
                            
                                How to set the color of the place holder text for a UITextField while preserving its existing properties?
                            
                                getActionBar().setDisplayHomeAsUpEnabled(true); throws NullPointerException on new activity creation (Google - Basic Tutorial)
                            
                                Convert 12 hour into 24 hour times

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With