github huggingface/datasets 2.13.0

latest releases: 3.1.0, 3.0.2, 3.0.1...
17 months ago

Dataset Features

  • Add IterableDataset.from_spark by @maddiedawson in #5770

    • Stream the data from your Spark DataFrame directly to your training pipeline
    from datasets import IterableDataset
    from torch.utils.data import DataLoader
    
    ids = IterableDataset.from_spark(df)
    ids = ids.map(...).filter(...).with_format("torch")
    for batch in DataLoader(ids, batch_size=16, num_workers=4):
        ...
  • IterableDataset formatting for PyTorch, TensorFlow, Jax, NumPy and Arrow:

    from datasets import load_dataset
    
    ids = load_dataset("c4", "en", split="train", streaming=True)
    ids = ids.map(...).with_format("torch")  # to get PyTorch tensors - also works with tf, np, jax etc.
  • Add IterableDataset.from_file to load local dataset as iterable by @mariusz-jachimowicz-83 in #5893

    from datasets import IterableDataset
    
    ids = IterableDataset.from_file("path/to/data.arrow")
  • Arrow dataset builder to be able to load and stream Arrow datasets by @mariusz-jachimowicz-83 in #5944

    from datasets import load_dataset
    
    ds = load_dataset("arrow", data_files={"train": "train.arrow", "test": "test.arrow"})

Experimental

  • Add parallel module using joblib for Spark by @es94129 in #5924

General improvements and bug fixes

New Contributors

Full Changelog: 2.12.0...zef

Don't miss a new datasets release

NewReleases is sending notifications on new releases.