logoalt Hacker News

gozzoo • yesterday at 1:10 PM • 10 replies • view on HN

I'm not following the trends closely, but has Polars become a full replacement for Pandas? Are there use cases where one is better suited than the other?


Replies

desipenguin • yesterday at 1:47 PM

From recent Python Bytes podcast (https://pythonbytes.fm/episodes/show/496/a-lake-house-in-sea...)

> 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory

Python Vs Rust : In terms for speed - No comparison

(The above episode transcript has a link to blog post titled "Pandas should go extinct" )

esco2292 • yesterday at 1:40 PM

Polars is effectively a full replacement for Pandas for 99.9% of all cases. The only exception I'm really aware of is if you're working with geospatial data, as there isn't yet a "Geopolars" equivalent of the commonly used "Geopandas". However, Geopolars is still in active development and should eventually be production ready.

➕ show 1 reply
SukadarBukadar • yesterday at 9:14 PM

Pandas is better for slight in-place or per-row modifications, for loading from less conventional datasets, for transposition/more nuanced row-based aggregation/multi-axis manipulation, for performance when multiprocessing can be used, for interoperability with other libraries (e.g. plotting and statistics)... When you need to operate on huge datasets, use DuckDB, because its performance is even now on-par with Polars regarding speed, while handling huge or more complex joins is a huge win for DuckDB because it better offloads intermediate results to disk, while Polars just dies on me. I haven't tried such joins in Polars 2 though

lmeyerov • yesterday at 2:26 PM

We were able to do a full port of GFQL from pandas to polars, cypher graph queries on dataframes, including both our CPU + GPU modes, and hit massive speedups: https://www.graphistry.com/blog/cypher-on-polars-cpu-gpu-gra...

It's been impressive!

niltecedu • yesterday at 2:52 PM

Yes and no, its not replacing the reason why pandas was popular ie data scientists, but it a full replacement of its pipeline usage, And I would saw also beating out spark

392 • yesterday at 1:23 PM

my understanding is Polars is faster, scales better without using external solutions, better API, +Rust. Pandas wins if you want to use what the vast majority of folks are using and have used in the past. Probably has a more complete set of helpers / recipes for the little things you bump into when using it thoroughly, but in the age of LLMs, I think that's minor.

➕ show 3 replies
seemaze • yesterday at 3:45 PM

It has been for me. I greatly prefer the API, it fits my mental model much better. Give it a try!

minimaxir • yesterday at 3:29 PM

See "Pandas should go extinct": https://news.ycombinator.com/item?id=49668198

tl;dr yes

bmitc • yesterday at 2:59 PM

There are awkward things. For example, if you ingest a nanosecond resolution timestamp, there's no way to re-export that out of the Polars dataframe with nanosecond resolution.

gremlinunderway • yesterday at 7:51 PM

From experience I'll say one thing that Polars doesn't have great interfacing with is doing things like string concatenation, like taking multiple columns and combining them in with static string text in complex ways to create new columns.

Pandas has a really simple ability to just define a new column with

`df['col_a'] + "text" + df[col_b']` where "text" can be any string text inbetween your column values from col_a and col_b

If i remember correctly while you can do pl.col("col_a") + pl.("col_b") for plain concatenation, you can't mix in static text strings like you can with pandas and I haven't found really elegant ways to do that personally. Whereas I've found polars doesn't have as simple of a way to do that. You can choose to add one separator and make that anything you want, but only one separator and only inbetween the two values (so no suffixes or prefixes for example).

That being said, I hate everything to do with the pandas API (especially with its indexing system) and really prefer the more polars API for anyone coming from a SQL or database background. Pandas really shows its sort of academia background rather than a data engineering origin.

➕ show 1 reply