I have a string of arbitrary characters, some of which are digits. I would like to break the strings into fields consisting of digits and non-digits. For example, if my string has the value 'abc34d-f9', I would like to get an array
['abc','34','d-f','9']
I'm nearly there, using look-behind and look-ahead expressions:
s.split(/( (?<=\D)(?=\d) | (?<=\d)(?=\D) )/x)
This splits on transitions between boundaries digit->nondigit and vice versa. However, I also get empty elements, i.e. this would return
['abc','','34','','d-f','','9']
Of course it is trivial to filter out the nullstrings from the array. I just wonder: Why do I get them, and how can I do it better?
Use string.scan function to return an array of matched strings.
> 'abc34d-f9'.scan(/\D+|\d+/)
=> ["abc", "34", "d-f", "9"]
\D+ matches one or more non-digit characters where \d+ matches one or more digit characters.
Your regex also works fine if you remove the capturing group. Because capturing group would also return the delimiter(boundary on which the input string was splitted) to the final output.
> 'abc34d-f9'.split(/(?<=\D)(?=\d)|(?<=\d)(?=\D)/)
=> ["abc", "34", "d-f", "9"]
> 'abc34d-f9'.split(/ (?<=\D)(?=\d) | (?<=\d)(?=\D) /x)
=> ["abc", "34", "d-f", "9"]
If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!
Donate Us With