Scraping with Scrapy and Selenium

Tags:

I have a scrapy spider which crawls a site that reloads content via javascript on the page. In order to move to the next page to scrape, I have been using Selenium to click on the month link at the top of the site.

The problem is that, even though my code moves through each link as expected, the spider just scrapes the first month (Sept) data for the number of months and returns this duplicate data.

How can I get around this?

from selenium import webdriver

class GigsInScotlandMain(InitSpider):
        name = 'gigsinscotlandmain'
        allowed_domains = ["gigsinscotland.com"]
        start_urls = ["http://www.gigsinscotland.com"]


    def __init__(self):
        InitSpider.__init__(self)
        self.br = webdriver.Firefox()

    def parse(self, response):
        hxs = HtmlXPathSelector(response)
        self.br.get(response.url)
        time.sleep(2.5)
        # Get the string for each month on the page.
        months = hxs.select("//ul[@id='gigsMonths']/li/a/text()").extract()

        for month in months:
            link = self.br.find_element_by_link_text(month)
            link.click()
            time.sleep(5)

            # Get all the divs containing info to be scraped.
            listitems = hxs.select("//div[@class='listItem']")
            for listitem in listitems:
                item = GigsInScotlandMainItem()
                item['artist'] = listitem.select("div[contains(@class, 'artistBlock')]/div[@class='artistdiv']/span[@class='artistname']/a/text()").extract()
                #
                # Get other data ...
                #
                yield item

503

asked Sep 16 '13 19:09

puffin

1 Answers

The problem is that you are reusing HtmlXPathSelector that was defined for the initial response. Redefine it from selenium browser source_code:

...
for month in months:
    link = self.br.find_element_by_link_text(month)
    link.click()
    time.sleep(5)

    hxs = HtmlXPathSelector(self.br.page_source)

    # Get all the divs containing info to be scraped.
    listitems = hxs.select("//div[@class='listItem']")
...

145

answered Sep 25 '22 14:09

alecxe

Related questions
                            
                                psycopg2.ProgrammingError: syntax error at or near "\"
                            
                                Exception signature
                            
                                Apply function to a MultiIndex dataframe with pandas/python
                            
                                Reordering columns/rows of a pivot_table?
                            
                                How to find backlinks in a website with python [closed]
                            
                                How can I specify the version of Python that Perl's Inline::Python module is using?
                            
                                Resizing a 3D image (and resampling)
                            
                                python xmpp simple client error
                            
                                How do I start with Django ORM to easily switch to SQLAlchemy?
                            
                                Determine arguments where two numpy arrays intersect in Python
                            
                                Python - Test if Raw-Input has no entry
                            
                                How to execute awk command by python code
                            
                                Efficient way to find the index of the max upper triangular entry in a numpy array?
                            
                                How do I set min and max resize sizes for my main window?
                            
                                How to programmatically post to my own wall on Facebook?
                            
                                Consecutive requests with python Requests.Session() not working
                            
                                OpenCV find coloured in circle and position value Python
                            
                                Scipy quickly initialise sparse matrix from list of coordinates
                            
                                Is there a way to load the constant values stored in a header file via ctypes?
                            
                                Python filename from inode?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Scraping with Scrapy and Selenium

Tags:

python

selenium

scrapy

puffin

People also ask

1 Answers

alecxe

Recent Activity

Donate For Us