Spark SQL Row_number() PartitionBy Sort Desc

Tags:

I've successfully create a row_number() partitionBy by in Spark using Window, but would like to sort this by descending, instead of the default ascending. Here is my working code:

from pyspark import HiveContext
from pyspark.sql.types import *
from pyspark.sql import Row, functions as F
from pyspark.sql.window import Window

data_cooccur.select("driver", "also_item", "unit_count", 
    F.rowNumber().over(Window.partitionBy("driver").orderBy("unit_count")).alias("rowNum")).show()

That gives me this result:

 +------+---------+----------+------+
 |driver|also_item|unit_count|rowNum|
 +------+---------+----------+------+
 |   s10|      s11|         1|     1|
 |   s10|      s13|         1|     2|
 |   s10|      s17|         1|     3|

And here I add the desc() to order descending:

data_cooccur.select("driver", "also_item", "unit_count", F.rowNumber().over(Window.partitionBy("driver").orderBy("unit_count").desc()).alias("rowNum")).show()

And get this error:

AttributeError: 'WindowSpec' object has no attribute 'desc'

What am I doing wrong here?

732

asked Feb 06 '16 22:02

jKraut

2 Answers

desc should be applied on a column not a window definition. You can use either a method on a column:

from pyspark.sql.functions import col, row_number from pyspark.sql.window import Window  F.row_number().over(     Window.partitionBy("driver").orderBy(col("unit_count").desc()) )

or a standalone function:

from pyspark.sql.functions import desc from pyspark.sql.window import Window  F.row_number().over(     Window.partitionBy("driver").orderBy(desc("unit_count")) )

answered Oct 02 '22 15:10

zero323

Or you can use the SQL code in Spark-SQL:

from pyspark.sql import SparkSession  spark = SparkSession\     .builder\     .master('local[*]')\     .appName('Test')\     .getOrCreate()  spark.sql("""     select driver         ,also_item         ,unit_count         ,ROW_NUMBER() OVER (PARTITION BY driver ORDER BY unit_count DESC) AS rowNum     from data_cooccur """).show()

answered Oct 02 '22 17:10

kennyut

Related questions
                            
                                Why use Tornado and Flask together?
                            
                                Python, SQLAlchemy pass parameters in connection.execute
                            
                                Pylint: overriding max-line-length in individual file
                            
                                Calculate Matrix Rank using scipy
                            
                                Access the sole element of a set
                            
                                TypeError: 'list' object cannot be interpreted as an integer
                            
                                How can I get the name of an object?
                            
                                Python timedelta issue with negative values
                            
                                Testing file uploads in Flask
                            
                                Remove list from list in Python [duplicate]
                            
                                Python for loop and iterator behavior
                            
                                Group by and find top n value_counts pandas
                            
                                The number of GET/POST parameters exceeded settings.DATA_UPLOAD_MAX_NUMBER_FIELDS
                            
                                Localized date strftime in Django view
                            
                                How to query directly the table created by Django for a ManyToMany relation?
                            
                                How to add group labels for bar charts in matplotlib
                            
                                How can I install lxml in docker
                            
                                Python packages hash not matching whilst installing using pip
                            
                                Python SQLite parameter substitution with wildcards in LIKE
                            
                                converting currency with $ to numbers in Python pandas

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Spark SQL Row_number() PartitionBy Sort Desc

Tags:

python

window-functions

apache-spark

apache-spark-sql

pyspark

jKraut

People also ask

2 Answers

zero323

kennyut

Recent Activity

Donate For Us