Extracting all Nouns from a text file using nltk

Tags:

python

nltk

Is there a more efficient way of doing this? My code reads a text file and extracts all Nouns.

import nltk

File = open(fileName) #open file
lines = File.read() #read all lines
sentences = nltk.sent_tokenize(lines) #tokenize sentences
nouns = [] #empty to array to hold all nouns

for sentence in sentences:
     for word,pos in nltk.pos_tag(nltk.word_tokenize(str(sentence))):
         if (pos == 'NN' or pos == 'NNP' or pos == 'NNS' or pos == 'NNPS'):
             nouns.append(word)

How do I reduce the time complexity of this code? Is there a way to avoid using the nested for loops?

Thanks in advance!

838

asked Nov 07 '15 20:11

Rakesh Adhikesavan

2 Answers

If you are open to options other than NLTK, check out TextBlob. It extracts all nouns and noun phrases easily:

>>> from textblob import TextBlob
>>> txt = """Natural language processing (NLP) is a field of computer science, artificial intelligence, and computational linguistics concerned with the inter
actions between computers and human (natural) languages."""
>>> blob = TextBlob(txt)
>>> print(blob.noun_phrases)
[u'natural language processing', 'nlp', u'computer science', u'artificial intelligence', u'computational linguistics']

131

answered Oct 17 '22 11:10

Aziz Alto

import nltk

lines = 'lines is some string of words'
# function to test if something is a noun
is_noun = lambda pos: pos[:2] == 'NN'
# do the nlp stuff
tokenized = nltk.word_tokenize(lines)
nouns = [word for (word, pos) in nltk.pos_tag(tokenized) if is_noun(pos)] 

print nouns
>>> ['lines', 'string', 'words']

Useful tip: it is often the case that list comprehensions are a faster method of building a list than adding elements to a list with the .insert() or append() method, within a 'for' loop.

answered Oct 17 '22 12:10

Boa

Related questions
                            
                                Angular UI Bootstrap Modal: [$injector:unpr] Unknown provider: $uibModalInstanceProvider
                            
                                Styling ionic 2 toast
                            
                                'this' is undefined in a Mongoose pre save hook [duplicate]
                            
                                How to recover deleted iPython Notebooks
                            
                                Could not write JSON: Infinite recursion (StackOverflowError); nested exception spring boot
                            
                                Collection View Compositional Layout with estimated height not working
                            
                                Parsing Binary Data in C?
                            
                                Is returning early from a function more elegant than an if statement?
                            
                                How do I empty Drupal Cache (without Devel)
                            
                                Using hibernate/hql to truncate a table?
                            
                                What's the point of valid CSS/HTML?
                            
                                What is the last event to fire when loading a new WPF/C# window?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With