R- Changing encoding of column in dataframe?

Question

I am trying to change the encoding of a column in a dataframe.

stri_enc_mark(data_updated$text)
#   [1] "UTF-8" "ASCII" "ASCII" "UTF-8" "ASCII" "ASCII" "UTF-8" "UTF-8" "UTF-8"
#  [10] "ASCII" "ASCII" "UTF-8" "ASCII" "UTF-8" "ASCII" "UTF-8" "ASCII" "UTF-8"
#  [19] "ASCII" "UTF-8" "ASCII" "UTF-8" "ASCII" "UTF-8" "UTF-8" "ASCII" "ASCII"
#  [28] "ASCII" "ASCII" "UTF-8" "ASCII" "ASCII" "ASCII" "UTF-8" "UTF-8" "ASCII"

When I try to convert it, it does not throw an error, but still has no effect on the vector:

d <- enc2utf8(data_updated$text)
stri_enc_mark(d)
#   [1] "UTF-8" "ASCII" "ASCII" "UTF-8" "ASCII" "ASCII" "UTF-8" "UTF-8" "UTF-8"
#  [10] "ASCII" "ASCII" "UTF-8" "ASCII" "UTF-8" "ASCII" "UTF-8" "ASCII" "UTF-8"
#  [19] "ASCII" "UTF-8" "ASCII" "UTF-8" "ASCII" "UTF-8" "UTF-8" "ASCII" "ASCII"
#  [28] "ASCII" "ASCII" "UTF-8" "ASCII" "ASCII" "ASCII" "UTF-8" "UTF-8" "ASCII"

Any suggestions?

I am on Windows 7, 32bit. Adding data snippet.

> Encoding(data_updated$text[1:35])
 [1] "UTF-8"   "unknown" "unknown" "UTF-8"   "unknown" "unknown" "UTF-8"  
 [8] "UTF-8"   "UTF-8"   "unknown" "unknown" "UTF-8"   "unknown" "UTF-8"  
[15] "unknown" "UTF-8"   "unknown" "UTF-8"   "unknown" "UTF-8"   "unknown"
[22] "UTF-8"   "unknown" "UTF-8"   "UTF-8"   "unknown" "unknown" "unknown"
[29] "unknown" "UTF-8"   "unknown" "unknown" "unknown" "UTF-8"   "UTF-8"

Data looks like this.

> data_updated$text[1:35]
 [1] "RT @satpalpandey: Majlis started in Sirsa Ashram.
Inform others too.
Live @ http://t.co/zGXWATGajX
IVR Airtel 55252
Reliance 56300403

#MSG…"
 [2] "Deal Talks for Here Mapping Service Expose Reliance on Location Data, via @nytimes #mapping #dilemma  http://t.co/wGdiS5OlRq"                      
 [3] "http://t.co/UZIyX1Rk7W The popping linksexploaded!! http://t.co/KpNntm1dH7 :) http://t.co/oku91uVxZ8"                                              
 [4] "RT @davidsunaria90: Wtch LIVE Mjlis Now
 http://t.co/GXNhe3eY7Y
IVR Airtel: 55252
Reliance: 56300403
Youtube Link : http://t.co/YewOVcz8bb
…" 
 [5] "Reliance Jio Infocomm: Indian carrier raises $750 million loan for 4G rollout  http://t.co/B2aWlkmwXz"                                             
 [6] "RT @SurjeetInsan: Majlis started in Sirsa Ashram.
Live @ http://t.co/PR6W5tzZes
IVR Airtel 55252
Reliance 56300403

#MSGPlsSaveTheEarth"      
 [7] "\"Deal Talks for Here Mapping Service Expose Reliance on Location Data\" by MARK SCOTT and MIKE ISAAC via NYT Techno… http://t.co/kyxTYIxks5"      
 [8] "RT @satpalpandey: Majlis started in Sirsa Ashram.
Inform others too.
Live @ http://t.co/zGXWATGajX
IVR Airtel 55252
Reliance 56300403

#MSG…"
 [9] "RT @jaameinsan: Watch LIVE Majlis Now
 http://t.co/nPQegnLXPa
IVR Airtel: 55252
Reliance: 56300403
Youtube Link : http://t.co/txXMtw3zFP
#M…" 
[10] "\"Deal Talks for Here Mapping Service Expose Reliance on Location Data\" by MARK SCOTT and MIKE ISAAC via NYT Technology"

These are tweets, and I think the "http://" links are dictating encoding here, given that they have expressions like "wGdiS5OlRq". For analysis I had removed these tags using regular expressions. But to store raw data in a DB i need these tweets. MongoDB does not have problem, but a RDBMS throws issues.

J.Delannoy · Accepted Answer

In case someone is still stuck : I used Encoding().

  for (col in colnames(mydataframe)){
  Encoding(mydataframe[[col]]) <- "UTF-8"}

NEO · Answer

It appears that we can use the conv() function to convert the encoding after we convert the vector into Factor and then back to character vector. It is a bit strange to be honest.

R- Changing encoding of column in dataframe?

Tags:

dataframe

r

encoding

NEO

2 Answers

J.Delannoy

NEO

Recent Activity

Donate For Us

R- Changing encoding of column in dataframe?

Tags:

dataframe

r

encoding

NEO

2 Answers

J.Delannoy

NEO

Related questions

Recent Activity

Donate For Us