Is it possible to save a pandas data frame directly to a parquet file? If not, what would be the suggested process?
The aim is to be able to send the parquet file to another team, which they can use scala code to read/open it. Thanks!
Pandas provides a beautiful Parquet interface. Pandas leverages the PyArrow library to write Parquet files, but you can also write Parquet files directly from PyArrow.
Pandas has a core function to_parquet()
. Just write the dataframe to parquet format like this:
df.to_parquet('myfile.parquet')
You still need to install a parquet library such as fastparquet
. If you have more than one parquet library installed, you also need to specify which engine you want pandas to use, otherwise it will take the first one to be installed (as in the documentation). For example:
df.to_parquet('myfile.parquet', engine='fastparquet')
If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!
Donate Us With