This command works with HiveQL: <pre class="prettyprint"><code>insert overwrite directory '/data/home.csv' select * from testtable; </code></pre> But with Spark SQL I'm getting an error with an <code>org.apache.spark.sql.hive.HiveQl</code> stack trace: <pre class="prettyprint"><code>java.lang.RuntimeException: Unsupported language features in query: insert overwrite directory '/data/home.csv' select * from testtable </code></pre> Please guide me to write export to CSV feature in Spark SQL.

Since Spark <code>2.X</code> <code>spark-csv</code> is integrated as native datasource. Therefore, the necessary statement simplifies to (windows) <pre class="prettyprint"><code>df.write .option("header", "true") .csv("file:///C:/out.csv") </code></pre> or UNIX <pre class="prettyprint"><code>df.write .option("header", "true") .csv("/var/out.csv") </code></pre> <hr> Notice: as the comments say, it is creating the directory by that name with the partitions in it, not a standard CSV file. This, however, is most likely what you want since otherwise your either crashing your driver (out of RAM) or you could be working with a non distributed environment.

How to export data from Spark SQL to CSV

Tags:

export-to-csv

apache-spark

apache-spark-sql

hadoop

hiveql

This command works with HiveQL:

Click to copy

insert overwrite directory '/data/home.csv' select * from testtable;

But with Spark SQL I'm getting an error with an org.apache.spark.sql.hive.HiveQl stack trace:

Click to copy

java.lang.RuntimeException: Unsupported language features in query:     insert overwrite directory '/data/home.csv' select * from testtable

Please guide me to write export to CSV feature in Spark SQL.

808

asked Aug 11 '15 09:08

shashankS

2 Answers

You can use below statement to write the contents of dataframe in CSV format df.write.csv("/data/home/csv")

If you need to write the whole dataframe into a single CSV file, then use df.coalesce(1).write.csv("/data/home/sample.csv")

For spark 1.x, you can use spark-csv to write the results into CSV files

Below scala snippet would help

Click to copy

import org.apache.spark.sql.hive.HiveContext // sc - existing spark context val sqlContext = new HiveContext(sc) val df = sqlContext.sql("SELECT * FROM testtable") df.write.format("com.databricks.spark.csv").save("/data/home/csv")

To write the contents into a single file

Click to copy

import org.apache.spark.sql.hive.HiveContext // sc - existing spark context val sqlContext = new HiveContext(sc) val df = sqlContext.sql("SELECT * FROM testtable") df.coalesce(1).write.format("com.databricks.spark.csv").save("/data/home/sample.csv")

122

answered Oct 07 '22 15:10

sag

Since Spark 2.X spark-csv is integrated as native datasource. Therefore, the necessary statement simplifies to (windows)

Click to copy

df.write   .option("header", "true")   .csv("file:///C:/out.csv")

or UNIX

Click to copy

df.write   .option("header", "true")   .csv("/var/out.csv")

Notice: as the comments say, it is creating the directory by that name with the partitions in it, not a standard CSV file. This, however, is most likely what you want since otherwise your either crashing your driver (out of RAM) or you could be working with a non distributed environment.

answered Oct 07 '22 15:10

Boern

Related questions
                            
                                Schema evolution in parquet format
                            
                                How to write 'map only' hadoop jobs?
                            
                                COLLECT_SET() in Hive, keep duplicates?
                            
                                Default Namenode port of HDFS is 50070.But I have come across at some places 8020 or 9000 [closed]
                            
                                java.net.URISyntaxException when starting HIVE
                            
                                What is a container in YARN?
                            
                                What are SUCCESS and part-r-00000 files in hadoop
                            
                                Explode the Array of Struct in Hive
                            
                                Hadoop java.io.IOException: Mkdirs failed to create /some/path
                            
                                Hive installation issues: Hive metastore database is not initialized
                            
                                hdfs dfs -put with overwrite?
                            
                                Hive query output to file
                            
                                http://localhost:50070 does not work HADOOP
                            
                                Is it better to use the mapred or the mapreduce package to create a Hadoop Job?
                            
                                hadoop.mapred vs hadoop.mapreduce?
                            
                                What is RDD in spark
                            
                                Hadoop/Hive : Loading data from .csv on a local machine
                            
                                java.lang.RuntimeException: Unable to instantiate org.apache.hadoop.hive.ql.metadata.SessionHiveMetaStoreClient
                            
                                data block size in HDFS, why 64MB?
                            
                                Spark iterate HDFS directory

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

How to export data from Spark SQL to CSV

Tags:

export-to-csv

apache-spark

apache-spark-sql

hadoop

hiveql

shashankS

People also ask

2 Answers

sag

Boern

Recent Activity

Donate For Us