Is there a bi gram or tri gram feature in Spacy?

Tags:

The below code breaks the sentence into individual tokens and the output is as below

 "cloud"  "computing"  "is" "benefiting"  " major"  "manufacturing"  "companies"


import en_core_web_sm
nlp = en_core_web_sm.load()

doc = nlp("Cloud computing is benefiting major manufacturing companies")
for token in doc:
    print(token.text)

What I would ideally want is, to read 'cloud computing' together as it is technically one word.

Basically I am looking for a bi gram. Is there any feature in Spacy that allows Bi gram or Tri grams ?

425

asked Dec 03 '18 16:12

4 Answers

Spacy allows the detection of noun chunks. So to parse your noun phrases as single entities do this:

Detect the noun chunks https://spacy.io/usage/linguistic-features#noun-chunks
Merge the noun chunks
Do dependency parsing again, it would parse "cloud computing" as single entity now.

>>> import spacy
>>> nlp = spacy.load('en')
>>> doc = nlp("Cloud computing is benefiting major manufacturing companies")
>>> list(doc.noun_chunks)
[Cloud computing, major manufacturing companies]
>>> for noun_phrase in list(doc.noun_chunks):
...     noun_phrase.merge(noun_phrase.root.tag_, noun_phrase.root.lemma_, noun_phrase.root.ent_type_)
... 
Cloud computing
major manufacturing companies
>>> [(token.text,token.pos_) for token in doc]
[('Cloud computing', 'NOUN'), ('is', 'VERB'), ('benefiting', 'VERB'), ('major manufacturing companies', 'NOUN')]

146

answered Oct 11 '22 20:10

['all what i want is that you give me back my code because i worked a lot on it. Just give me back my code', 'how are you? i am just showing you an example of how to make bigrams on spacy. We love bigrams on spacy', 'i love to repeat phrases to make bigrams because i love make bigrams']

then step 2 and 3:

doc = ' '.join(textList)
spacy_doc = parser(doc)
print(spacy_doc)

and will print this:

all what i want is that you give me back my code because i worked a lot on it. Just give me back my code how are you? i am just showing you an example of how to make bigrams on spacy. We love bigrams on spacy i love to repeat phrases to make bigrams because i love make bigrams

Finally step 4 (Zuzana's answer)

ngrams = list(textacy.extract.ngrams(spacy_doc, 2, min_freq=2))
print(ngrams)

will print this:

[make bigrams, make bigrams, make bigrams]

answered Oct 11 '22 18:10

iair linker

I had a similar problem (bigrams, trigrams, like your "cloud computing"). I made a simple list of the n-grams, word_3gram, word_2grams etc., with the gram as basic unit (cloud_computing).

Assume I have the sentence "I like cloud computing because it's cheap". The sentence_2gram is: "I_like", "like_cloud", "cloud_computing", "computing_because" ... Comparing that your bigram list only "cloud_computing" is recognized as a valid bigram; all other bigrams in the sentence are artificial. To recover all other words you just take the first part of the other words,

"I_like".split("_")[0] -> I; 
"like_cloud".split("_")[0] -> like
"cloud_computing" -> in bigram list, keep it. 
  skip next bi-gram "computing_because" ("computing" is already used)
"because_it's".split("_")[0]" -> "because" etc.

To also capture the last word in the sentence ("cheap") I added the token "EOL". I implemented this in python, and the speed was OK (500k words in 3min), i5 processor with 8G. Anyway, you have to do it only once. I find this more intuitive than the official (spacy-style) chunk approach. It also works for non-spacy frameworks.

I do this before the official tokenization/lemmatization, as you would get "cloud compute" as possible bigram. But I'm not certain if this is the best/right approach.

answered Oct 11 '22 20:10

user9165100

Related questions
                            
                                Collapsing rows in a Pandas dataframe
                            
                                pyinstaller add folder with images in exe file
                            
                                How to use tqdm with map for Dataframes
                            
                                Group by max or min in a numpy array
                            
                                Write function result to stdin
                            
                                Default to and select first item in Tkinter listbox
                            
                                Type error Iter - Python3
                            
                                How to while loop until the end of a file in Python without checking for empty line?
                            
                                How do I upload a Universal Python Wheel for Python 2 and 3?
                            
                                Standard solution for supporting Python 2 and Python 3
                            
                                Getting ValueError: The indices for endog and exog are not aligned
                            
                                ('Unexpected credentials type', None, 'Expected', 'service_account') with oauth2client (Python)
                            
                                python - cannot make corr work
                            
                                How to stop daemon thread?
                            
                                python : cannot import name JIRA
                            
                                How to solve the circular import error in django?
                            
                                Why is my image_path undefined when using export_graphviz? - Python 3
                            
                                lxml / BeautifulSoup parser warning
                            
                                How can I display all numbers in range 0-N that are "super numbers"
                            
                                how to apply yapf to every python file under a directory?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Is there a bi gram or tri gram feature in Spacy?

Tags:

python-3.x

tokenize

nlp

n-gram

spacy

venkatttaknev

People also ask

4 Answers

DhruvPathak

Suzana

iair linker

user9165100

Recent Activity

Donate For Us