I am using Pyspark with Python 2.7. I have a date column in string (with ms) and would like to convert to timestamp This is what I have tried so far <pre class="prettyprint lang-py prettyprint-override"><code>df = df.withColumn('end_time', from_unixtime(unix_timestamp(df.end_time, '%Y-%M-%d %H:%m:%S.%f')) ) </code></pre> <code>printSchema()</code> shows <code>end_time: string (nullable = true)</code> when I expended timestamp as the type of variable

Try using <code>from_utc_timestamp</code>: <pre class="prettyprint lang-py prettyprint-override"><code>from pyspark.sql.functions import from_utc_timestamp df = df.withColumn('end_time', from_utc_timestamp(df.end_time, 'PST')) </code></pre> You'd need to specify a timezone for the function, in this case I chose <code>PST</code> If this does not work please give us an example of a few rows showing <code>df.end_time</code>

Create a sample dataframe with Time-stamp formatted as string: <pre class="prettyprint lang-py prettyprint-override"><code>import pyspark.sql.functions as F df = spark.createDataFrame([('22-Jul-2018 04:21:18.792 UTC', ),('23-Jul-2018 04:21:25.888 UTC',)], ['TIME']) df.show(2,False) df.printSchema() </code></pre> Output: <pre class="prettyprint"><code>+----------------------------+ |TIME | +----------------------------+ |22-Jul-2018 04:21:18.792 UTC| |23-Jul-2018 04:21:25.888 UTC| +----------------------------+ root |-- TIME: string (nullable = true) </code></pre> Converting string time-format (including milliseconds ) to unix_timestamp(double). Since unix_timestamp() function excludes milliseconds we need to add it using another simple hack to include milliseconds. Extracting milliseconds from string using substring method (start_position = -7, length_of_substring=3) and Adding milliseconds seperately to unix_timestamp. (Cast to substring to float for adding) <pre class="prettyprint lang-py prettyprint-override"><code>df1 = df.withColumn("unix_timestamp",F.unix_timestamp(df.TIME,'dd-MMM-yyyy HH:mm:ss.SSS z') + F.substring(df.TIME,-7,3).cast('float')/1000) </code></pre> Converting unix_timestamp(double) to timestamp datatype in Spark. <pre class="prettyprint lang-py prettyprint-override"><code>df2 = df1.withColumn("TimestampType",F.to_timestamp(df1["unix_timestamp"])) df2.show(n=2,truncate=False) </code></pre> This will give you following output <pre class="prettyprint"><code>+----------------------------+----------------+-----------------------+ |TIME |unix_timestamp |TimestampType | +----------------------------+----------------+-----------------------+ |22-Jul-2018 04:21:18.792 UTC|1.532233278792E9|2018-07-22 04:21:18.792| |23-Jul-2018 04:21:25.888 UTC|1.532319685888E9|2018-07-23 04:21:25.888| +----------------------------+----------------+-----------------------+ </code></pre> Checking the Schema: <pre class="prettyprint"><code>df2.printSchema() root |-- TIME: string (nullable = true) |-- unix_timestamp: double (nullable = true) |-- TimestampType: timestamp (nullable = true) </code></pre>

Pyspark from_unixtime (unix_timestamp) does not convert to timestamp

Tags:

date

pyspark

I am using Pyspark with Python 2.7. I have a date column in string (with ms) and would like to convert to timestamp

This is what I have tried so far

df = df.withColumn('end_time', from_unixtime(unix_timestamp(df.end_time, '%Y-%M-%d %H:%m:%S.%f')) )

printSchema() shows end_time: string (nullable = true)

when I expended timestamp as the type of variable

580

asked Jan 24 '19 01:01

qqplot

2 Answers

Try using from_utc_timestamp:

from pyspark.sql.functions import from_utc_timestamp

df = df.withColumn('end_time', from_utc_timestamp(df.end_time, 'PST'))

You'd need to specify a timezone for the function, in this case I chose PST

If this does not work please give us an example of a few rows showing df.end_time

answered Sep 27 '22 22:09

Tanjin

Create a sample dataframe with Time-stamp formatted as string:

import pyspark.sql.functions as F
df = spark.createDataFrame([('22-Jul-2018 04:21:18.792 UTC', ),('23-Jul-2018 04:21:25.888 UTC',)], ['TIME'])
df.show(2,False)
df.printSchema()

Output:

+----------------------------+
|TIME                        |
+----------------------------+
|22-Jul-2018 04:21:18.792 UTC|
|23-Jul-2018 04:21:25.888 UTC|
+----------------------------+
root
|-- TIME: string (nullable = true)

Converting string time-format (including milliseconds ) to unix_timestamp(double). Since unix_timestamp() function excludes milliseconds we need to add it using another simple hack to include milliseconds. Extracting milliseconds from string using substring method (start_position = -7, length_of_substring=3) and Adding milliseconds seperately to unix_timestamp. (Cast to substring to float for adding)

df1 = df.withColumn("unix_timestamp",F.unix_timestamp(df.TIME,'dd-MMM-yyyy HH:mm:ss.SSS z') + F.substring(df.TIME,-7,3).cast('float')/1000)

Converting unix_timestamp(double) to timestamp datatype in Spark.

df2 = df1.withColumn("TimestampType",F.to_timestamp(df1["unix_timestamp"]))
df2.show(n=2,truncate=False)

This will give you following output

+----------------------------+----------------+-----------------------+
|TIME                        |unix_timestamp  |TimestampType          |
+----------------------------+----------------+-----------------------+
|22-Jul-2018 04:21:18.792 UTC|1.532233278792E9|2018-07-22 04:21:18.792|
|23-Jul-2018 04:21:25.888 UTC|1.532319685888E9|2018-07-23 04:21:25.888|
+----------------------------+----------------+-----------------------+

Checking the Schema:

df2.printSchema()


root
 |-- TIME: string (nullable = true)
 |-- unix_timestamp: double (nullable = true)
 |-- TimestampType: timestamp (nullable = true)

answered Sep 27 '22 21:09

Sangram Gaikwad

Related questions
                            
                                Best way to find date nearest to target in a list of dates?
                            
                                How do I use the C date and time functions on UNIX?
                            
                                Group by date range on weeks/months interval
                            
                                Date Validation In Rails 3
                            
                                Datejs - Problem with 12:00 pm
                            
                                Add days to current date from MySQL with PHP
                            
                                How can dates and random numbers be used for evil in Javascript?
                            
                                Format(SomeDate,"MM/dd") = "12-15" in VBA
                            
                                Calculate number of weeks , days and hours from milliseconds
                            
                                jQuery UI datepicker: add 6 months to another datepicker
                            
                                Parse a String to Date in Java
                            
                                Arithmetics on calendar dates in C or C++ (add N days to given date)
                            
                                Ruby: comparing dates of two Time objects
                            
                                Why does java.util.Date represent Year as "year-1900"?
                            
                                IE Input type Date not appearing as Date Picker [duplicate]
                            
                                Get time in milliseconds based on a given time zone (Local time zone)
                            
                                Get same weekend last year using moment js
                            
                                HttpClient HttpResponseMessage LastModified date of file
                            
                                Format a local date at another time zone
                            
                                Convert month's number to Month name

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With