Repository navigation
Can't call open_mfdataset without creating chunked dask arrays #9038
Description
Activity
- addedtopic-chunked-arraysManaging different chunked backends, e.g. daskManaging different chunked backends, e.g. dask
on May 21, 2024 Passing chunks=None to open_mfdataset should return lazily-indexed numpy arrays, like open_dataset does.
Can't do this without virtual concat machinery (#4628) which someone decided to implement elsewhere 🙄 ;)
We could change the default to
chunks={}in anticipation though.Can't do this without virtual concat machinery (#4628) which someone decided to implement elsewhere 🙄 ;)
😅
It's still broken at the moment though - I had a (ridiculous) case where I don't care that concat will load everything in memory, I just want to completely avoid creating
dask.arrayobjects, and right now there is no possible input option toopen_mfdatasetto do that.We could change the default to chunks={} in anticipation though.
That's probably more useful, as well as actually being consistent.
Reacted by James Mineau and Wei JiSee #5704 for changing
chunks={}and more discussion.Passing
chunks=Nonetoxr.open_dataset/open_mfdatasetis supposed to avoid using dask at all, returning lazily-indexed numpy arrays even if dask is installed.Agree with having the ability to totally avoid dask and/or any chunking manager, and have the possibility of letting the backend engine handle things on its own. Will also call out that the lazy-indexing shouldn't just return NumPy arrays, but allow for other Array API types (e.g. CuPy arrays).
To elaborate on my use-case, I'm hitting into issues at xarray-contrib/cupy-xarray#81 (comment) because
xr.open_mfdatasetcan't avoid using dask at all, and dask defaults to returning numpy.arrays, but I want CuPy arrays 🙂 (edit: went down the 🐇🕳️ and found #8733 (comment) that mentioned cupy!). Usingxr.open_dataset(..., chunks=None)works fine though, it is just thatchunks=chunks or {}part that is tripping things up inxr.open_mfdataset.- added a commit that references this issue
on Jun 18, 2026
What happened?
Passing
chunks=Nonetoxr.open_dataset/open_mfdatasetis supposed to avoid using dask at all, returning lazily-indexed numpy arrays even if dask is installed. Howeverchunks=Nonedoesn't currently work forxr.open_mfdatasetas it gets silently coerced internally tochunks={}, which creates dask chunks aligned with the on-disk files.Offending line of code:
xarray/xarray/backends/api.py
Line 1040 in 12123be
What did you expect to happen?
Passing
chunks=Nonetoopen_mfdatasetshould return lazily-indexed numpy arrays, likeopen_datasetdoes.Minimal Complete Verifiable Example
MVCE confirmation
Relevant log output
Anything else we need to know?
As the default is
None, changing this without changing the default would be a breaking change. But the current behaviour is also not intended.Environment
main