Polars 2.0
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Polars says version 2.0 makes its streaming engine the default for lazy queries and enables initial spill-to-disk support, alongside SQL and performance changes. Its published TPC-H and TPC-DS results favor Polars in most tested benchmarks, but those results come from the project’s own tests and have stated limitations.

The Polars project has released Polars 2.0, making its streaming engine the default for collecting lazy queries and enabling initial spill-to-disk support for workloads that exceed available memory. The major release also expands SQL support and adds a dedicated Map data type, changes that affect how data practitioners run queries and handle certain results.

Under the new default, calling collect on a LazyFrame uses the streaming engine. Polars says this can bring memory and performance improvements for many queries. The change can also affect row ordering: the release post says operations including joins, group-bys and unpivots do not guarantee observable row order by default. Users who require ordering can set maintain_order=True for supported operations.

Polars describes its out-of-core feature as an initial release. It is enabled by default and begins spilling data to disk at about 80% of RAM use, with a default disk budget of 64 GB. At launch, supported operations include sorting, window functions and many expressions; joins and group-bys are not yet supported for out-of-core execution, though the project says they are on its roadmap. The release post says the threshold may need tuning.

Version 2.0 also adds a native Map dtype corresponding to Arrow’s MapType, replacing the earlier representation as a list of key-value structs. The project highlights map operations such as retrieving a value by key, checking whether a key exists, and accessing keys or values. It also cites optimizer and engine work, including join reordering, improved common-subplan elimination and dynamic predicates or bloom filters, as part of its SQL and performance changes.

At a glance
announcementWhen: Announced in the Polars 2.0 release pos…
The developmentThe Polars project has released version 2.0, changing lazy-query execution defaults and adding initial out-of-core support.

Streaming Changes Query Defaults

The shift to streaming by default changes a central behavior for users of Polars’ lazy API, not just the set of available features. It may help queries handle larger workloads with less memory pressure, while making row order a setting users need to check when it matters to their results. That trade-off is relevant to existing applications as well as new projects adopting version 2.0.

Disk spilling offers another way to complete some memory-intensive work when data exceeds RAM, but current support is limited to specified operations. Users should not assume every query can spill, particularly joins and group-bys. The new Map dtype may also require attention in workflows that previously handled Arrow maps through Polars’ list-of-struct representation.

The release presents SQL as a first-class interface and reports favorable benchmark performance. For readers choosing an analytical engine, the results are a useful signal, but not an independent or universal comparison: outcomes depend on hardware, queries, data and configuration.

Amazon

Polars 2.0 data analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Polars Tested SQL Performance

Polars says it compared its SQL engine with DuckDB 1.5.6, a DuckDB 2.0 alpha build and DataFusion 54.0.0 using data derived from TPC-H and TPC-DS. Tests ran on two AWS machine types: a c7a.4xlarge with 16 vCPUs and 32 GB of memory, and a c7a.metal with 192 vCPUs and 384 GB. Each query ran five times in a hot setting, and the fastest run was used. The project compared both summed and geometric-mean query times.

According to Polars, its default configuration was fastest in all but one of the reported benchmarks. The project also reported that Polars and both DuckDB versions completed all queries, while DataFusion timed out on TPC-DS query 72, timed out once on query 67 and ran out of memory on TPC-H query 18 on the smaller machine. Those affected queries were excluded from results for all engines.

The company disclosed a scaling limitation: Polars has overhead when using 192 threads, which hurts some small-data queries. It said a 32-core Polars configuration was competitive or leading in all benchmarks and that it hopes to address the overhead in a later release. Polars shared a benchmark repository so others can replicate the test; the release post does not provide independent validation.

Amazon

out-of-core data processing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits in Spill and Benchmarks

The release post does not specify a calendar date for the announcement in the supplied material, nor does it give a general performance guarantee for users’ workloads. The claimed benchmark advantage is based on Polars-run tests; the post provides test conditions and a replication repository, but the supplied source contains no independent results. Polars also says the 192-thread overhead remains unresolved and may be addressed in a future release.

Out-of-core execution is not yet available for joins or group-bys, and the project says its approximate 80% RAM threshold may need tuning. The post does not establish when those operations will gain disk-spill support. Users should also verify whether row order is part of their application’s requirements when upgrading.

Amazon

disk spill data analytics

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Out-of-Core Support

Polars says it plans to add out-of-core support for joins and group-bys, but gives no delivery date. It also says it hopes to fix the performance overhead observed when scaling to 192 threads in a subsequent release. Until those changes arrive, users can check which operations support spilling and reproduce the published SQL benchmarks using the project’s shared repository.

For teams moving to 2.0, the immediate next step is to test representative lazy queries, check any dependence on row ordering, and review Map data handling. The release post does not list a separate upgrade timetable or provide a date for a follow-up release.

Amazon

SQL data analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main change in Polars 2.0?

LazyFrame.collect now defaults to the streaming engine. Polars also enables initial spill-to-disk support by default and adds SQL, optimizer and data-type changes.

Does the new streaming default preserve row order?

Not for certain operations by default, according to Polars. The release post names joins, group-bys and unpivots as examples and says users who need observable ordering can opt in with maintain_order=True.

Which operations can spill data to disk?

At launch, Polars lists sorting, window functions and many expressions as supported. Joins and group-bys are not yet supported for out-of-core execution, though the project says it plans to add them.

Did Polars prove it is faster than DuckDB and DataFusion?

No independent proof is included in the supplied source. Polars reports that its default configuration was fastest in all but one of its tested TPC-H and TPC-DS benchmarks, while also disclosing test limitations and sharing a repository for replication.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

I’m a USB-C Maximalist

A tech enthusiast proclaims unwavering support for USB-C, emphasizing its versatility and future-proofing in a recent online statement.

More Tailscale Tricks For Your Jailbroken Kindle

Developers reveal advanced Tailscale configurations for jailbroken Kindles, expanding their network capabilities and security options.

Opinionated And Easy Pi.dev Configuration

Pi.dev introduces a new, opinionated, and easy-to-use configuration system, streamlining setup for developers. Details are emerging on its impact and adoption.

Does Intelligence Need A Hard Cap?

At The Curve conference, speakers reportedly discussed capping AI capabilities amid concerns about self-improving systems and weak enforcement tools.