Tratopedia
Side N繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

Dataframe Libraries · Python · Performance Engineering · Benchmarking · Comparison BriefRecord to 23 Sep 2026

Python's two main dataframe libraries, as they stand in September 2026

Polars and pandas: Which Suits Which Job

pandas has been open source since 2009 and is still Python's most-downloaded table library. Polars, first committed in June 2020 and written in Rust on the Apache Arrow memory format, was built to be faster. Both changed in 2026. pandas 3.0 turned on a dedicated string type and Copy-on-Write by default. Polars 2.0, now at its second release candidate, makes its streaming engine the default and stops guaranteeing row order to do it. The speed gap is real and large on big analytical work. What is missing is a published measurement of the two as they now are.

  • 3.0.6current pandas, released 17 September 2026; the 3.0 line began on 21 January
  • 1.44.2current Polars, released 9 September 2026; 2.0 is at its second release candidate
  • 94×Polars' lead over pandas 2.2.3 at about 10 GB, in Polars' own benchmark of June 2025
  • 10×pandas' lead in PyPI downloads over the last month: 525.5 million to 51.6 million

Graded by how well each item stands up Sorted by standing. What was measured, what a company says about itself and what is only promised are never in the same row.

StandingWhat is recordedWhere
ConfirmedCurrent releases: pandas 3.0.6 (17 September 2026) and Polars 1.44.2 (9 September 2026). Polars 2.0.0rc1 came out on 2 September and rc2 on 20 September. pandas needs Python 3.11 or later, Polars 3.10 or laterThe Python Package Index
Confirmedpandas 3.0 turned on a dedicated string type, backed by PyArrow when it is installed; made Copy-on-Write the default, so chained assignment stops working and SettingWithCopyWarning is gone; added a first form of pd.col() expressions; and now parses date strings to microseconds by defaultpandas' release notes
ConfirmedPolars 2.0 runs every lazy query on the streaming engine by default. That engine does not guarantee row order for joins, group-bys and unpivots unless asked with maintain_order. It also refuses lossy type coercions and mismatched concatenations it used to let throughPolars' 2.0 announcement
ConfirmedPolars has no row index, holds data in the Apache Arrow columnar format rather than NumPy arrays, runs more operations in parallel, and can plan and optimise a lazy query before running itPolars' guide for pandas users
MeasuredAt scale factor 10, about 10 GB of CSV, Polars' streaming engine ran the whole query set in 3.89 seconds and pandas 2.2.3 in 365.71. At about 100 GB pandas was not run: it hit out-of-memory failuresPolars' own PDS-H benchmark, June 2025
MeasuredWith large dataframes Polars used about one-eighth of pandas' energy on synthetic tasks and 63% on TPC-H tasks. With small dataframes the TPC-H tasks showed no significant difference. Versions tested: pandas 2.1.1, Polars 0.19.7Nahrstedt and others, EASE 2024, peer-reviewed
Company's claim, not verifiable hereAdoption grew “from 250k to over 23M monthly users” after the 2023 seed round, and the core library is used in production in finance, life sciences and logisticsPolars; TechCrunch, reporting Polars
Company's claim, not verifiable here“In aggregate we expect the streaming engine to be easily 5x faster”: a forecast in the 2.0 announcement, not a measurementPolars' 2.0 announcement
Promised, not yet shipped“Proper out-of-core support for the streaming engine”, meaning dependable work on data larger than memory, plus a cost-based planner and join reordering. All are listed as still in flightPolars' 2.0 announcement
Promised, not yet shippedPolars 2.0 final, due “in the following weeks” after 2 September. It had not shipped by 23 SeptemberPolars' 2.0 announcement; the Python Package Index

2008 to September 2026 Every speed and energy measurement below comes before pandas 3.0. None compares the current versions of both.

  1. 2008pandas development begins at AQR Capital Management, under Wes McKinney
  2. 2009pandas is open-sourced by the end of the year; version 0.1 reaches PyPI on 25 December
  3. 2015pandas becomes a NumFOCUS-sponsored project
  4. 21 Sep 2017McKinney writes that his rule of thumb for pandas is 5 to 10 times as much RAM as the dataset
  5. 23 Jun 2020Ritchie Vink makes the first commit to Polars
  6. 3 Aug 2023Vink and Chiel Peters announce a Polars company, with a seed round of about $4 million led by Bain Capital Ventures; the library stays MIT-licensed
  7. 18 Jan 2024scikit-learn 1.4 lets its transformers output Polars dataframes
  8. 18 Jun 2024A Vrije Universiteit Amsterdam team presents its energy and performance study of pandas 2.1.1 and Polars 0.19.7 at EASE 2024
  9. 1 Jul 2024Polars 1.0.0
  10. 17 Sep 2024Polars' GPU engine, built with NVIDIA on RAPIDS cuDF, opens as a beta
  11. 1 Jun 2025Polars publishes PDS-H results for Polars 1.30.0 against pandas 2.2.3, DuckDB, Dask and PySpark
  12. 29 Sep 2025Polars raises an €18 million Series A led by Accel, about $21 million, with Bain Capital returning
  13. 21 Jan 2026pandas 3.0.0: string type and Copy-on-Write on by default, first pd.col() expressions, Python 3.11 minimum
  14. 2 Sep 2026Polars 2.0.0rc1, with the streaming engine as the default for lazy queries
  15. 9 Sep 2026Polars 1.44.2, the current stable release
  16. 17 Sep 2026pandas 3.0.6, the current stable release
  17. 20 Sep 2026Polars 2.0.0rc2. No final 2.0 yet

Much faster on big work, and measured on versions that have moved on The case for Polars, as its makers measured it, then the two things that qualify it.

Polars' case is architectural, and its own benchmark puts a number on it. pandas runs each step as soon as it is written, on one core, with no optimiser. Polars can take a whole query, optimise it before it runs, and spread it across every core. In Polars' PDS-H run of June 2025, on a 96-core machine, that was the difference between 365.71 seconds and 3.89 at about 10 GB. Polars says pandas was left out at about 100 GB because it hit out-of-memory failures. The same run shows two things a reader should keep. DuckDB, a third engine, was close behind Polars at 10 GB and ahead of it at 100 GB. And Polars' older in-memory engine slowed sharply at that size, which is why 2.0 makes streaming the default.

  • 3.89 sPolars streaming, about 10 GB, whole query set
  • 365.71 spandas 2.2.3, same queries, same machine
Engine, in Polars' PDS-H run of June 2025About 10 GB, secondsAbout 100 GB, seconds
Polars 1.30.0, streaming3.8923.94
DuckDB 1.3.05.8719.65
Polars 1.30.0, in-memory9.68152.27
Dask 2025.5.146.02548.52
PySpark 4.0.0120.11312.43
pandas 2.2.3365.71not run: out of memory

Two things qualify that. The first is size. The one peer-reviewed comparison, from a Vrije Universiteit Amsterdam team at EASE 2024, ran the same kind of TPC-H tasks at 6 million rows and found no significant difference in energy, the authors noting a slight advantage for pandas. At 60 million rows, and on their own synthetic tasks, Polars was clearly ahead. The second is age, and it is this article's reading, not a source's. The newest pandas in any measurement here is 2.2.3, from September 2024. The study used pandas 2.1.1 and Polars 0.19.7. pandas 3.0's Arrow-backed strings and Copy-on-Write came after all of them, and so did fourteen Polars minor lines, 1.31 to 1.44. The gap is very probably still large on big work. How large, for the current versions, nobody here has measured.

  • 6Mrows: no significant difference on TPC-H tasks in the peer-reviewed study
  • 2.2.3the newest pandas anyone here measured; the current release is 3.0.6

Meanwhile the two have moved toward each other, and the costs have moved with them. This too is this article's reading. pandas 3.0 took on Arrow-backed strings, copy-on-write semantics and the first expression syntax, all ideas Polars was built around. Polars 2.0 goes the other way on one guarantee pandas users take for granted: under the new default, a join or group-by no longer promises to keep rows in their original order. Polars calls the change necessary for streaming speed and tells users who rely on order to ask for it with maintain_order. Code moved from one to the other, or upgraded within Polars, can change its output without raising an error.

Memory, energy, GPUs, and using both From the library's creator, the study's authors, NVIDIA, and the projects around the two.

  • Memory

    A rule of thumb, and a roadmap

    Wes McKinney · Polars

    • pandas' creator wrote in 2017 that his rule of thumb was 5 to 10 times as much RAM as the dataset: a 10 GB dataset wants 64 to 128 GB
    • That rule predates pandas 3.0 by more than eight years. No comparable figure for Polars was found at a primary source
    • Polars lists “proper out-of-core support”, meaning dependable work beyond RAM, as still to come in 2.x
  • Energy

    One controlled study

    EASE 2024 · the authors, on Polars' blog

    • 28 tasks, 10 repetitions each in random order, 560 runs, on an isolated server
    • Large dataframes: Polars used about one-eighth of pandas' energy on synthetic tasks, and 63% on TPC-H tasks
    • Energy tracked execution time closely, and the authors credit Polars' use of every CPU core
  • GPU

    Up to 13×, at the top end

    Polars · NVIDIA

    • An optional engine on NVIDIA's cuDF, switched on with collect(engine="gpu"); queries it cannot run fall back to the CPU
    • “Up to 13x” is the best of the four best speedups NVIDIA charted, out of 22 PDS-H queries at scale factor 80, on an H100
    • Released as an open beta in September 2024; its status today was not established here
  • Together

    Less of an either-or

    scikit-learn · Narwhals

    • Since 1.4, in January 2024, scikit-learn's transformers can return Polars dataframes with set_output(transform="polars")
    • Narwhals, a compatibility layer with no dependencies, lets a library written against a subset of Polars' API accept pandas, Polars, cuDF, Modin or PyArrow
    • It also offers lazy-only support for Dask, DuckDB, Ibis, PySpark and SQLFrame
  • Adoption

    One order of magnitude apart

    PyPI Stats · Polars

    • Downloads in the last month: pandas 525,518,404, Polars 51,587,659, a ratio of 10.2
    • Downloads count installs, every automated build included, not people
    • In August 2023 Polars reported over 6 million downloads in total; by September 2025 it reported over 23 million monthly users

How much weight the headline benchmark bears depends on who ran it, so its terms matter. PDS-H is Polars' own derivative of TPC-H, and Polars states that its results are not comparable with published TPC-H results. Its rules forbid hand-pruning a table before a join and hand-reordering joins, and require each solution to run every query on a single engine: “No cherry picking.” The post also names where Polars lost: DuckDB at 100 GB, and query 21, a range join the streaming engine had not yet implemented. The EASE paper, written without the vendor, observed that the benchmarks before it were neither systematic nor peer-reviewed. On large data its results point the same way as Polars'.

Which one suits which job Two free libraries that increasingly work together. The question is the reader's workload, not which one wins.

  • pandas

    For existing code and step-by-step work

    • Suits: code already written against it, work that labels rows with an index or edits a table in place, and datasets small enough that the controlled study found no significant difference on its TPC-H tasks
    • Best at: reach. It is downloaded ten times as often as Polars, and it has been open source since 2009
    • Best at, since 3.0: more predictable copies and faster strings without leaving the API people know
    • Trade-off: eager and largely single-threaded, and memory-hungry on large data by its own creator's rule of thumb
  • Polars

    For large analytical pipelines

    • Suits: joins, group-bys and scans over tens of millions of rows or gigabytes, pipelines written as whole queries, and machines with many cores
    • Best at: speed and energy on big work. It was two orders of magnitude ahead of pandas 2.2.3 at about 10 GB in its own benchmark, and far more frugal on large data in the independent study
    • Best at, too: strictness. It raises on lossy casts and mismatched lengths instead of silently producing a different answer
    • Trade-off: no index and a new way of thinking in expressions; under 2.0, row order in joins and group-bys must be requested; working beyond RAM is still being finished

What is established

The current versions are pandas 3.0.6 and Polars 1.44.2, with Polars 2.0 at its second release candidate. pandas 3.0 made Arrow-backed strings and Copy-on-Write the default. Polars 2.0 makes streaming the default and drops guaranteed row order in joins and group-bys. On large analytical work, older versions of Polars were much faster than older versions of pandas, in Polars' own benchmark and in the peer-reviewed study. On the study's small TPC-H tasks the difference was not significant. And the two can be used together, through scikit-learn and Narwhals.

What to hold loosely

Any single speed-up figure: 94× is one vendor run at one size against pandas 2.2.3, and 13× is the best GPU query of 22. The 5× streaming gain and dependable work beyond RAM are forecasts. The user and production claims are the company's own. The 5–10× RAM figure is a 2017 rule of thumb. And the reading that the gap is still large for the current versions is this article's: no published measurement compares pandas 3.0 with Polars yet.

Sources: The pandas development team, “What's new in 3.0.0”, pandas documentation (21 January 2026); the pandas project, “About pandas”, pandas.pydata.org (read 23 September 2026); the Python Package Index, release histories of pandas and polars (read 23 September 2026); PyPI Stats, download statistics for pandas and polars (read 23 September 2026); Wes McKinney, “Apache Arrow and the ‘10 Things I Hate About pandas’”, wesmckinney.com (21 September 2017); Ritchie Vink and Chiel Peters, “Company announcement”, Polars (3 August 2023); Ritchie Vink, “Polars raises €18M Series A to build fast, ergonomic data processing at any scale”, Polars (29 September 2025); Anna Heim, “The startup behind open source tool Polars raises $21M from Accel”, TechCrunch (29 September 2025); Ritchie Vink, “Updated PDS-H benchmark results”, Polars (1 June 2025); Ritchie Vink, “Pre-release of Polars 2.0”, Polars (2 September 2026); Ritchie Vink, Robin van den Brink and Lawrence Mitchell, “GPU acceleration with Polars and NVIDIA RAPIDS”, Polars (17 September 2024); Jamil Semaan, “Polars GPU Engine Powered by NVIDIA cuDF Now Available in Open Beta”, NVIDIA Technical Blog (17 September 2024); Felix Nahrstedt, Mehdi Karmouche, Karolina Bargieł, Pouyeh Banijamali, Apoorva Nalini Pradeep Kumar and Ivano Malavolta, “An Empirical Study on the Energy Usage and Performance of Pandas and Polars Data Analysis Python Libraries”, EASE 2024, ACM (18 June 2024); Ivano Malavolta and Karolina Bargieł, “Benchmarking energy usage and performance of Polars and pandas”, Polars (22 August 2024); the Polars project, “Coming from Pandas”, Polars user guide (read 23 September 2026); the scikit-learn developers, “Release Highlights for scikit-learn 1.4”, scikit-learn.org (January 2024); the Narwhals developers, “Narwhals”, Narwhals documentation (read 23 September 2026).

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0001 current

The current version is also at latest/.