How are non-capturing groups, i.e., <code>(?:)</code>, used in regular expressions and what are they good for?

Let me try to explain this with an example. Consider the following text: <pre class="prettyprint lang-none prettyprint-override"><code>http://stackoverflow.com/ https://stackoverflow.com/questions/tagged/regex </code></pre> Now, if I apply the regex below over it... <pre class="prettyprint lang-regex prettyprint-override"><code>(https?|ftp)://([^/\r\n]+)(/[^\r\n]*)? </code></pre> ... I would get the following result: <pre class="prettyprint"><code>Match "http://stackoverflow.com/" Group 1: "http" Group 2: "stackoverflow.com" Group 3: "/" Match "https://stackoverflow.com/questions/tagged/regex" Group 1: "https" Group 2: "stackoverflow.com" Group 3: "/questions/tagged/regex" </code></pre> But I don't care about the protocol -- I just want the host and path of the URL. So, I change the regex to include the non-capturing group <code>(?:)</code>. <pre class="prettyprint lang-regex prettyprint-override"><code>(?:https?|ftp)://([^/\r\n]+)(/[^\r\n]*)? </code></pre> Now, my result looks like this: <pre class="prettyprint"><code>Match "http://stackoverflow.com/" Group 1: "stackoverflow.com" Group 2: "/" Match "https://stackoverflow.com/questions/tagged/regex" Group 1: "stackoverflow.com" Group 2: "/questions/tagged/regex" </code></pre> See? The first group has not been captured. The parser uses it to match the text, but ignores it later, in the final result. <hr> <h3>EDIT:</h3> As requested, let me try to explain groups too. Well, groups serve many purposes. They can help you to extract exact information from a bigger match (which can also be named), they let you rematch a previous matched group, and can be used for substitutions. Let's try some examples, shall we? Imagine you have some kind of XML or HTML (be aware that regex may not be the best tool for the job, but it is nice as an example). You want to parse the tags, so you could do something like this (I have added spaces to make it easier to understand): <pre class="prettyprint lang-none prettyprint-override"><code> \<(?<TAG>.+?)\> [^<]*? \</\k<TAG>\> or \<(.+?)\> [^<]*? \</\1\> </code></pre> The first regex has a named group (TAG), while the second one uses a common group. Both regexes do the same thing: they use the value from the first group (the name of the tag) to match the closing tag. The difference is that the first one uses the name to match the value, and the second one uses the group index (which starts at 1). Let's try some substitutions now. Consider the following text: <pre class="prettyprint lang-none prettyprint-override"><code>Lorem ipsum dolor sit amet consectetuer feugiat fames malesuada pretium egestas. </code></pre> Now, let's use this dumb regex over it: <pre class="prettyprint lang-regex prettyprint-override"><code>\b(\S)(\S)(\S)(\S*)\b </code></pre> This regex matches words with at least 3 characters, and uses groups to separate the first three letters. The result is this: <pre class="prettyprint"><code>Match "Lorem" Group 1: "L" Group 2: "o" Group 3: "r" Group 4: "em" Match "ipsum" Group 1: "i" Group 2: "p" Group 3: "s" Group 4: "um" ... Match "consectetuer" Group 1: "c" Group 2: "o" Group 3: "n" Group 4: "sectetuer" ... </code></pre> So, if we apply the substitution string: <pre class="prettyprint lang-none prettyprint-override"><code>$1_$3$2_$4 </code></pre> ... over it, we are trying to use the first group, add an underscore, use the third group, then the second group, add another underscore, and then the fourth group. The resulting string would be like the one below. <pre class="prettyprint lang-none prettyprint-override"><code>L_ro_em i_sp_um d_lo_or s_ti_ a_em_t c_no_sectetuer f_ue_giat f_ma_es m_la_esuada p_er_tium e_eg_stas. </code></pre> You can use named groups for substitutions too, using <code>${name}</code>. To play around with regexes, I recommend http://regex101.com/, which offers a good amount of details on how the regex works; it also offers a few regex engines to choose from.

What is a non-capturing group in regular expressions?

1 Answers

Let me try to explain this with an example.

Consider the following text:

http://stackoverflow.com/ https://stackoverflow.com/questions/tagged/regex

Now, if I apply the regex below over it...

(https?|ftp)://([^/\r\n]+)(/[^\r\n]*)?

... I would get the following result:

Match "http://stackoverflow.com/"      Group 1: "http"      Group 2: "stackoverflow.com"      Group 3: "/"  Match "https://stackoverflow.com/questions/tagged/regex"      Group 1: "https"      Group 2: "stackoverflow.com"      Group 3: "/questions/tagged/regex"

But I don't care about the protocol -- I just want the host and path of the URL. So, I change the regex to include the non-capturing group (?:).

(?:https?|ftp)://([^/\r\n]+)(/[^\r\n]*)?

Now, my result looks like this:

Match "http://stackoverflow.com/"      Group 1: "stackoverflow.com"      Group 2: "/"  Match "https://stackoverflow.com/questions/tagged/regex"      Group 1: "stackoverflow.com"      Group 2: "/questions/tagged/regex"

See? The first group has not been captured. The parser uses it to match the text, but ignores it later, in the final result.

EDIT:

As requested, let me try to explain groups too.

Well, groups serve many purposes. They can help you to extract exact information from a bigger match (which can also be named), they let you rematch a previous matched group, and can be used for substitutions. Let's try some examples, shall we?

Imagine you have some kind of XML or HTML (be aware that regex may not be the best tool for the job, but it is nice as an example). You want to parse the tags, so you could do something like this (I have added spaces to make it easier to understand):

   \<(?<TAG>.+?)\> [^<]*? \</\k<TAG>\> or    \<(.+?)\> [^<]*? \</\1\>

The first regex has a named group (TAG), while the second one uses a common group. Both regexes do the same thing: they use the value from the first group (the name of the tag) to match the closing tag. The difference is that the first one uses the name to match the value, and the second one uses the group index (which starts at 1).

Let's try some substitutions now. Consider the following text:

Lorem ipsum dolor sit amet consectetuer feugiat fames malesuada pretium egestas.

Now, let's use this dumb regex over it:

\b(\S)(\S)(\S)(\S*)\b

This regex matches words with at least 3 characters, and uses groups to separate the first three letters. The result is this:

Match "Lorem"      Group 1: "L"      Group 2: "o"      Group 3: "r"      Group 4: "em" Match "ipsum"      Group 1: "i"      Group 2: "p"      Group 3: "s"      Group 4: "um" ...  Match "consectetuer"      Group 1: "c"      Group 2: "o"      Group 3: "n"      Group 4: "sectetuer" ...

So, if we apply the substitution string:

$1_$3$2_$4

... over it, we are trying to use the first group, add an underscore, use the third group, then the second group, add another underscore, and then the fourth group. The resulting string would be like the one below.

L_ro_em i_sp_um d_lo_or s_ti_ a_em_t c_no_sectetuer f_ue_giat f_ma_es m_la_esuada p_er_tium e_eg_stas.

You can use named groups for substitutions too, using ${name}.

To play around with regexes, I recommend http://regex101.com/, which offers a good amount of details on how the regex works; it also offers a few regex engines to choose from.

125

answered Oct 15 '22 14:10

Ricardo Nolde

Related questions
                            
                                Regex to replace multiple spaces with a single space
                            
                                Regex for password must contain at least eight characters, at least one number and both lower and uppercase letters and special characters
                            
                                How to match "anything up until this sequence of characters" in a regular expression?
                            
                                Regular expression to stop at first match
                            
                                How to use Regular Expressions (Regex) in Microsoft Excel both in-cell and loops
                            
                                How do I remove all non alphanumeric characters from a string except dash?
                            
                                Regular Expression for alphanumeric and underscores
                            
                                How to count string occurrence in string?
                            
                                Find and kill a process in one line using bash and regex
                            
                                How to negate specific word in regex? [duplicate]
                            
                                Regular Expressions: Is there an AND operator?
                            
                                What is the best regular expression to check if a string is a valid URL?
                            
                                Check whether a string matches a regex in JS
                            
                                How to validate phone numbers using regex
                            
                                Is there a regular expression to detect a valid regular expression?
                            
                                Negative matching using grep (match lines that do not contain foo)
                            
                                \d less efficient than [0-9]
                            
                                How do you access the matched groups in a JavaScript regular expression?
                            
                                How do you use a variable in a regular expression?
                            
                                RegEx match open tags except XHTML self-contained tags

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

What is a non-capturing group in regular expressions?

Tags:

regex

capturing-group

regex-group

never_had_a_name

People also ask

1 Answers

EDIT:

Ricardo Nolde

Recent Activity

Donate For Us