XPath axis, get all following nodes until

Tags:

I have the following example of HTML:

<!-- lots of html -->
<h2>Foo bar</h2>
<p>lorem</p>
<p>ipsum</p>
<p>etc</p>

<h2>Bar baz</h2>
<p>dum dum dum</p>
<p>poopfiddles</p>
<!-- lots more html ... -->

I'm looking to extract all paragraphs following the 'Foo bar' header, until I reach the 'Bar baz' header (the text for the 'Bar baz' header is unknown, so unfortunately I can't use the answer provided by bougyman). Now I can of course using something like //h2[text()='Foo bar']/following::p but that of course will grab all paragraphs following this header. So I have the option to traverse the nodeset and push paragraphs into an Array until the text matches that of the next following header, but let's be honest, that's never as cool as being able to do it in XPath.

Is there a way to do this that I'm missing?

845

asked Jan 22 '11 11:01

Lee Jarvis

3 Answers

Use:

(//h2[. = 'Foo bar'])[1]/following-sibling::p
   [1 = count(preceding-sibling::h2[1] | (//h2[. = 'Foo bar'])[1])]

In case it is guaranteed that every h2 has a distinct value, this may be simplified to:

//h2[. = 'Foo bar']/following-sibling::p
   [1 = count(preceding-sibling::h2[1] | ../h2[. = 'Foo bar'])]

This means: Select all p elements that are following siblings of the h2 (first or only one in the document) whose string value is 'Foo bar' and also the first preceding sibling h2 for all these p elements is exactly the h2(first or only one in the document) whose string value is'Foo bar'`.

Here we use a method of finding whether two nodes are identical:

count($n1 | $n2) = 1

is true() exactly when the nodes $n1 and $n2 are the same node.

This expression can be generalized:

$x/following-sibling::p
       [1 = count(preceding-sibling::node()[name() = name($x)][1] | $x)]

selects all "immediate following siblings" of any node specified by $x.

200

answered Oct 12 '22 04:10

Dimitre Novatchev

In XPath 2.0 (I know this doesn't help you...) the simplest solution is probably

h2[. = 'Foo bar']/following-sibling::* except h2[. = 'Bar baz']/(.|following-sibling::* )

But like other solutions presented, this is likely (in the absence of an optimizer that recognizes the pattern) to be linear in the number of elements beyond the second h2, whereas you would really like a solution whose performance depends only on the number of elements selected. I've always felt it would be nice to have an until operator:

h2[. = 'Foo bar']/(following-sibling::* until . = 'Bar baz')

In its absence an XSLT or XQuery solution using recursion is likely to perform better when the number of nodes to be selected is small compared with the number of following siblings.

answered Oct 12 '22 04:10

Michael Kay

This XPATH 1.0 statement selects all of the <p> that are siblings that follow an <h2> who's string value is equal to "Foo bar", that are also followed by an <h2> sibling element who's first preceding sibling <h2> has a string value of "Foo bar".

//p[preceding-sibling::h2[.='Foo bar']]
 [following-sibling::h2[
  preceding-sibling::h2[1][.='Foo bar']]]

answered Oct 12 '22 02:10

Mads Hansen

Related questions
                            
                                Cucumber/Capybara vs Selenium? [closed]
                            
                                Escape single quote in XPath with Nokogiri?
                            
                                Disable Sprockets asset caching in development
                            
                                Bug when switching to Rails 4.1 - `method_missing': undefined method `whitelist_attributes=' for ActiveRecord::Base:Class (NoMethodError)
                            
                                Ruby On Rails with Windows Vista - Best Setup? [closed]
                            
                                (Ruby) Is there a function to easily find the first number in a string?
                            
                                How do you pipe output from a Ruby script to 'head' without getting a broken pipe error
                            
                                Rails times oddness : "x days from now"
                            
                                Web application admin generators [closed]
                            
                                Converting a unique seed string into a random, yet deterministic, float value in Ruby
                            
                                Ruby request to https - "in `read_nonblock': Connection reset by peer (Errno::ECONNRESET)"
                            
                                Ruby on Rails, two models in one form
                            
                                Is there a solution to bypass 'can't add a new key into hash during iteration (RuntimeError)'?
                            
                                How do I call Windows DLL functions from Ruby?
                            
                                Ruby on rails: Devise, want to add invite code?
                            
                                Ruby, How to add a param to an URL that you don't know if it has any other param already
                            
                                Is there a way to see the raw SQL that a Sequel expression will generate?
                            
                                Navigate to Ruby function definition in VS Code
                            
                                about ruby range?
                            
                                Parsing simple XML with Nokogiri

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

XPath axis, get all following nodes until

Tags:

ruby

xpath

nokogiri