Spark DataFrame equivalent to Pandas Dataframe `.iloc()` method?

Name: Selecting Rows and Columns from a Pandas DataFrame using .loc and .iloc
Uploaded: 2022-09-16 16:46:23
Description: Spark DataFrame equivalent to Pandas Dataframe `.iloc()` method?Is there a way to reference Spark DataFrame columns by position using an

Question

Is there a way to reference Spark DataFrame columns by position using an integer?

Analogous Pandas DataFrame operation:

df.iloc[:0] # Give me all the rows at column position 0

Chadee Fouad · Accepted Answer

The equivalent of Python df.iloc is collect

PySpark examples:

X = df.collect()[0]['age']

or

X = df.collect()[0][1]  #row 0 col 1

zero323 · Answer

Not really, but you can try something like this:

Python:

df = sc.parallelize([(1, "foo", 2.0)]).toDF()
df.select(*df.columns[:1])  # I assume [:1] is what you really want
## DataFrame[_1: bigint]

or

df.select(df.columns[1:3])
## DataFrame[_2: string, _3: double]

Scala

val df = sc.parallelize(Seq((1, "foo", 2.0))).toDF()
df.select(df.columns.slice(0, 1).map(col(_)): _*)

Note:

Spark SQL doesn't support and it is unlikely to ever support row indexing so it is not possible to index across row dimension.

Spark DataFrame equivalent to Pandas Dataframe `.iloc()` method?

Tags:

pandas

dataframe

scala

apache-spark

apache-spark-sql

conner.xyz

Video Answer

2 Answers

Chadee Fouad

zero323

Recent Activity

Donate For Us

Spark DataFrame equivalent to Pandas Dataframe `.iloc()` method?

Tags:

pandas

dataframe

scala

apache-spark

apache-spark-sql

conner.xyz

Video Answer

2 Answers

Chadee Fouad

zero323

Related questions

Recent Activity

Donate For Us