write_webdataset#

Dataset.write_webdataset(path: str, *, filesystem: pyarrow.fs.FileSystem | None = None, try_create_dir: bool = True, arrow_open_stream_args: Dict[str, Any] | None = None, filename_provider: FilenameProvider | None = None, min_rows_per_file: int | None = None, ray_remote_args: Dict[str, Any] = None, encoder: bool | str | Callable[[Dict[str, Any]], Dict[str, Any]] | List[bool | str | Callable[[Dict[str, Any]], Dict[str, Any]]] | None = True, concurrency: int | None = None, num_rows_per_file: int | None = None, mode: SaveMode = SaveMode.APPEND) → None[source]#

Writes the dataset to WebDataset tar archives.

Each row is written as a WebDataset sample.

This is only supported for datasets convertible to Arrow records. To control the number of files, use Dataset.repartition().

Unless a custom filename provider is given, generated output filenames end in .tar.

Note

This operation will trigger execution of the lazy transformations performed on this dataset.

Examples

import ray

ds = ray.data.from_items(
    [
        {"__key__": f"{i:06d}", "txt": str(i)}
        for i in range(100)
    ]
)
ds.write_webdataset("s3://bucket/folder/")

Time complexity: O(dataset size / parallelism)

Parameters:
  • path (str) – The path to the destination root directory, where WebDataset tar archives are written.

  • filesystem (pyarrow.fs.FileSystem | None) – The filesystem implementation to write to.

  • try_create_dir (bool) – If True, attempts to create all directories in the destination path. Does nothing if all directories already exist. Defaults to True.

  • arrow_open_stream_args (Dict[str, Any] | None) – kwargs passed to pyarrow.fs.FileSystem.open_output_stream

  • filename_provider (FilenameProvider | None) – A FilenameProvider implementation. Use this parameter to customize what your filenames look like.

  • min_rows_per_file (int | None) – [Experimental] The target minimum number of rows to write to each file. If None, Ray Data writes a system-chosen number of rows to each file. If the number of rows per block is larger than the specified value, Ray Data writes the number of rows per block to each file. The specified value is a hint, not a strict limit. Ray Data might write more or fewer rows to each file.

  • ray_remote_args (Dict[str, Any]) – Kwargs passed to ray.remote() in the write tasks.

  • encoder (bool | str | Callable[[Dict[str, Any]], Dict[str, Any]] | List[bool | str | Callable[[Dict[str, Any]], Dict[str, Any]]] | None) – Controls how dataset rows are encoded into WebDataset samples. A boolean or string selects the built-in encoder, which automatically handles common data types based on column-name extensions. A callable receives and returns a sample dictionary, and a list applies multiple encoder specifications in sequence. Set this to None to skip encoding; in that case, each non-special value that isn’t None must already be bytes or a string.

  • concurrency (int | None) – The maximum number of Ray tasks to run concurrently. Set this to control number of tasks to run concurrently. This doesn’t change the total number of tasks run. By default, concurrency is dynamically decided based on the available resources.

  • num_rows_per_file (int | None) – [Deprecated] Use min_rows_per_file instead.

  • mode (SaveMode) – Determines how to handle existing files. Valid modes are “overwrite”, “error”, “ignore”, “append”. Defaults to “append”. NOTE: This method isn’t atomic. “Overwrite” first deletes all the data before writing to path.

PublicAPI (alpha): This API is in alpha and may change before becoming stable.