/usr/local/lib/python3.6/site-packages/datasets/__pycache__
NameSizeModeActions
arrow_dataset.cpython-36.pyc1898500644editdlrm
arrow_reader.cpython-36.pyc224830644editdlrm
arrow_writer.cpython-36.pyc221100644editdlrm
builder.cpython-36.pyc527790644editdlrm
combine.cpython-36.pyc55080644editdlrm
config.cpython-36.pyc51810644editdlrm
dataset_dict.cpython-36.pyc833190644editdlrm
data_files.cpython-36.pyc301530644editdlrm
fingerprint.cpython-36.pyc189790644editdlrm
info.cpython-36.pyc173590644editdlrm
inspect.cpython-36.pyc185830644editdlrm
iterable_dataset.cpython-36.pyc629320644editdlrm
keyhash.cpython-36.pyc33410644editdlrm
load.cpython-36.pyc629220644editdlrm
metric.cpython-36.pyc232170644editdlrm
naming.cpython-36.pyc25710644editdlrm
search.cpython-36.pyc313390644editdlrm
splits.cpython-36.pyc225220644editdlrm
streaming.cpython-36.pyc41190644editdlrm
table.cpython-36.pyc796120644editdlrm
__init__.cpython-36.pyc22820644editdlrm
Edit: /usr/local/lib/python3.6/site-packages/datasets/__pycache__/builder.cpython-36.pyc (52779B)
3 <%EgW@sdZddlZddlZddlZddlZddlZddlZddlZddlZddl Z ddl Z ddl m Z ddl mZddlmZmZmZmZmZddlmZmZddlmZdd lmZmZmZmZmZdd l m!Z!m"Z"dd l#m$Z$m%Z%dd l&m'Z'm(Z(dd l)m*Z*ddl+m,Z,m-Z-ddl.m/Z/ddl0m1Z1ddl2m3Z3ddl4m5Z5ddl6m7Z7m8Z8m9Z9ddl:m;Z;mm?Z?ddl@mAZAmBZBddlCmDZDmEZEmFZFddlGmHZHddlmIZIddlJmKZKmLZLddlMmNZNddlOmPZPmQZQmRZRddlSmTZTmUZUmVZVmWZWmXZXmYZYeIjZe[Z\Gddde]Z^Gd d!d!e_Z`Gd"d#d#e`Zae Gd$d%d%ZbGd&d'd'ZcGd(d)d)ecZdGd*d+d+ecZeGd,d-d-e]ZfGd.d/d/ecZgdS)0zDatasetBuilder base class.N) dataclass)partial)DictMappingOptionalTupleUnion)configutils)Dataset)HF_GCP_BASE_URL ArrowReaderDatasetNotOnHfGcsErrorMissingFilesOnHfGcsErrorReadInstruction) ArrowWriter BeamWriter) DataFilesDictsanitize_patterns) DatasetDictIterableDatasetDict)DownloadConfig)DownloadManager DownloadMode)MockDownloadManager)StreamingDownloadManager)Features)Hasher) DatasetInfoDatasetInfosDictPostProcessedInfo)ExamplesIterableIterableDataset&_generate_examples_from_tables_wrapper)DuplicatedKeysError)"INVALID_WINDOWS_CHARACTERS_IN_PATHcamelcase_to_snakecase)Split SplitDictSplitGenerator)$extend_dataset_builder_for_streaming)logging) cached_path is_remote_url)FileLock)get_size_checksum_dictverify_checksums verify_splits) classpropertyhas_sufficient_disk_space map_nestedmemoizesize_strtemporary_assignmentc@s eZdZdS)InvalidConfigNameN)__name__ __module__ __qualname__r=r=:/usr/local/lib/python3.6/site-packages/datasets/builder.pyr9Isr9c@s eZdZdS)DatasetBuildErrorN)r:r;r<r=r=r=r>r?Msr?c@s eZdZdS)ManualDownloadErrorN)r:r;r<r=r=r=r>r@Qsr@c@seZdZUdZdZeejdZe e ejefdZ e e dZ e e  dZe eddZddZd ee eed d d ZdS) BuilderConfiga Base class for :class:`DatasetBuilder` data configuration. DatasetBuilder subclasses with data configuration options should subclass :class:`BuilderConfig` and add their own properties. Attributes: name (:obj:`str`, default ``"default"``): version (:class:`Version` or :obj:`str`, optional): data_dir (:obj:`str`, optional): data_files (:obj:`str` or :obj:`Sequence` or :obj:`Mapping`, optional): Path(s) to source data file(s). description (:obj:`str`, optional): defaultz0.0.0NcCs\x.tD]&}||jkrtdtd|jdqW|jdk rXt|jt rXtd|jdS)Nz Bad characters from black list 'z ' found in 'z\'. They could create issues when creating a directory for this config on Windows filesystem.z/Expected a DataFilesDict in data_files but got )r&namer9 data_files isinstancer ValueError)selfZ invalid_charr=r=r> __post_init__js   zBuilderConfig.__post_init__cs>tjjtjjkr dStfddjjDS)NFc3s*|]"}|t|f|t|fkVqdS)N)getattr).0k)orGr=r> zsz'BuilderConfig.__eq__..)set__dict__keysall)rGrLr=)rLrGr>__eq__uszBuilderConfig.__eq__) config_kwargscustom_featuresreturncs"d}|jjddjdddkrDddkrDjddrfddtDtddjDrd jd djD}t|d krtj }n tj }|dk rt}|r|j ||j ||j }|r|j d |}t|t jkr|j d tj |}|S|j SdS) a/ The config id is used to build the cache directory. By default it is equal to the config name. However the name of a config is not sufficient to have a unique identifier for the dataset being generated since it doesn't take into account: - the config kwargs that can be used to overwrite attributes - the custom features used to write the dataset - the data_files for json/text/csv/pandas datasets Therefore the config id is just the config name with an optional suffix based on these. NrCversiondata_dircsi|]}||qSr=r=)rJrK)config_kwargs_to_add_to_suffixr=r> sz2BuilderConfig.create_config_id..css |]}t|ttttfVqdS)N)rEstrboolintfloat)rJvr=r=r>rMsz1BuilderConfig.create_config_id..,css.|]&\}}t|dtjjt|VqdS)=N)rZurllibparse quote_plus)rJrKr^r=r=r>rMs -)copypopsortedrQvaluesjoinitemslenrhashupdate hexdigestrCr Z%MAX_DATASET_CONFIG_ID_READABLE_LENGTH)rGrSrTsuffixm config_idr=)rXr>create_config_id|s6          zBuilderConfig.create_config_id)N)r:r;r<__doc__rCrZr VersionrVrrrWrDr descriptionrHrRdictrrsr=r=r=r>rAUs      rAc@seZdZdZdZeZgZdZdZ dWe e e e e e e e e e e e e eee fe e e ee eeefe e d ddZdd Zd d Zee e d d dZeed ddZe d ddZdXeee fd ddZeeeddZeddZ dYe d ddZ!ddZ"e#j$e d ddZ%ed d!Z&dZe e'e e(eee e)e e e eee fd"d#d$Z*d%d&Z+e'd'd(d)Z,d*d+Z-d,d-Z.e d d.d/Z/d0d1Z0d2d3Z1d4d5Z2d[e e3ee4e5fd6d7d8Z6d\ee e7e3feeed9d:d;Z8e3j9dfee7e3fee4d<d=d>Z:ee7e3fe d6d?d@Z;d]e e e e eee=d dDdEZ?e4e@e e fe e4dFdGdHZAe e Parameter `name` was renamed to `config_name`. hash (`str`, *optional*): Hash specific to the dataset code. Used to update the caching directory when the dataset loading script code is updated (to avoid reusing old data). The typical caching directory (defined in ``self._relative_data_dir``) is: ``name/version/hash/``. base_path (`str`, *optional*): Base path for relative paths that are used to download files. This can be a remote URL. features ([`Features`], *optional*): Features types to use with this dataset. It can be used to change the Features types of a dataset, for example. use_auth_token (`str` or `bool`, *optional*): String or boolean to use as Bearer token for remote files on the Datasets Hub. If `True`, will get token from ``"~/.huggingface"``. repo_id (`str`, *optional*): ID of the dataset repository. Used to distinguish builders with the same name but not coming from the same namespace, for example "squad" and "lhoestq/squad" repo IDs. In the latter, the builder name would be "lhoestq___squad". data_files (`str` or `Sequence` or `Mapping`, *optional*): Path(s) to source data file(s). For builders like "csv" or "json" that need the user to specify data files. They can be either local or remote files. For convenience, you can use a DataFilesDict. data_dir (`str`, *optional*): Path to directory containing source data file(s). Use only if `data_files` is not passed, in which case it is equivalent to passing ``os.path.join(data_dir, "**")`` as `data_files`. For builders that require manual download, it must be the path to the local directory containing the manually downloaded data. name (`str`): Configuration name for the dataset. Use `config_name` instead. **config_kwargs (additional keyword arguments): Keyword arguments to be passed to the corresponding builder configuration class, set on the class attribute [`DatasetBuilder.BUILDER_CONFIG_CLASS`]. The builder configuration class is [`BuilderConfig`] or a subclass of it. NF deprecated) cache_dir config_namerm base_pathinfofeaturesuse_auth_tokenrepo_idrDrWc  Ks| dkrtjdtd| }t|jjdd|_||_||_||_ ||_ | dk rpt | t  rpt j t| ||d} dtj|jjjkr|dk r|| d<| dk r| | d<| dk r| | d <|j|fd |i| \|_|_|dkr|j}|j|j|j|_|jj|_|jj|_||_|dk r&||j_t|p2tj |_!t"|j!rJ|j!n t#j$j%|j!|_!t"|j!rlt&j'nt#j$j'} |r| |j!tj(nttj)|_*t"|j*r|j*n t#j$j%|j*|_*|j+|_,t"|j!sxt#j-|j!d d t#j$j'|j!|j,j.t#j/d d}t0|nt#j$j1|j,rnt2t#j3|j,dkrFt4jdt5j6|j,|_n(t4j7d|j,d|jdt#j8|j,WdQRXd|_9d|_:t;|dS)Nryz\Parameter 'name' was renamed to 'config_name' in version 2.3.0 and will be removed in 3.0.0.)category.r )r|rr~rDrWrTT)exist_ok_z.lockrz2Overwrite dataset info from restored data version.zOld caching folder z for dataset z. exists but not data were found. Removing it. F)<warningswarn FutureWarningr'r;splitrCrmr|rrrErZfrom_local_or_remoterinspect signatureBUILDER_CONFIG_CLASS__init__ parameters_create_builder_configr rrget_exported_dataset_inforn_info builder_namer{rVr}r~rZZHF_DATASETS_CACHE_cache_dir_rootr.ospath expanduser posixpathrjZDOWNLOADED_DATASETS_DIRZDOWNLOADED_DATASETS_PATH_cache_downloaded_dir_build_cache_dir _cache_dirmakedirsreplacesepr/existsrllistdirloggerrfrom_directorywarningrmdir dl_manager _record_infosr+)rGrzr{rmr|r}r~rrrDrWrCrS path_join lock_pathr=r=r>rsl    "      zDatasetBuilder.__init__cCs|jS)N)rO)rGr=r=r> __getstate__eszDatasetBuilder.__getstate__cCs||_t|dS)N)rOr+)rGdr=r=r> __setstate__hszDatasetBuilder.__setstate__)rUcCsdS)Nr=)rGr=r=r>manual_download_instructionsssz+DatasetBuilder.manual_download_instructionscCs2tjj|jtj}tjj|r.tj|jSiS)a"Empty dict if doesn't exist Example: ```py >>> from datasets import load_dataset_builder >>> ds_builder = load_dataset_builder('rotten_tomatoes') >>> ds_builder.get_all_exported_dataset_infos() {'default': DatasetInfo(description="Movie Review Dataset. This is a dataset of containing 5,331 positive and 5,331 negative processed sentences from Rotten Tomatoes movie reviews. This data was first used in Bo Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.'', Proceedings of the ACL, 2005. ", citation='@InProceedings{Pang+Lee:05a, author = {Bo Pang and Lillian Lee}, title = {Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales}, booktitle = {Proceedings of the ACL}, year = 2005 } ', homepage='http://www.cs.cornell.edu/people/pabo/movie-review-data/', license='', features={'text': Value(dtype='string', id=None), 'label': ClassLabel(num_classes=2, names=['neg', 'pos'], id=None)}, post_processed=None, supervised_keys=SupervisedKeysData(input='', output=''), task_templates=[TextClassification(task='text-classification', text_column='text', label_column='label')], builder_name='rotten_tomatoes_movie_review', config_name='default', version=1.0.0, splits={'train': SplitInfo(name='train', num_bytes=1074810, num_examples=8530, dataset_name='rotten_tomatoes_movie_review'), 'validation': SplitInfo(name='validation', num_bytes=134679, num_examples=1066, dataset_name='rotten_tomatoes_movie_review'), 'test': SplitInfo(name='test', num_bytes=135972, num_examples=1066, dataset_name='rotten_tomatoes_movie_review')}, download_checksums={'https://storage.googleapis.com/seldon-datasets/sentence_polarity_v1/rt-polaritydata.tar.gz': {'num_bytes': 487770, 'checksum': 'a05befe52aafda71d458d188a1c54506a998b1308613ba76bbda2e5029409ce9'}}, download_size=487770, post_processing_size=None, dataset_size=1345461, size_in_bytes=1833231)} ``` ) rrrjget_imported_module_dirr ZDATASETDICT_INFOS_FILENAMErr r)clsZdset_infos_file_pathr=r=r>get_all_exported_dataset_infosws  z-DatasetBuilder.get_all_exported_dataset_infoscCs|jj|jjtS)aEmpty DatasetInfo if doesn't exist Example: ```py >>> from datasets import load_dataset_builder >>> ds_builder = load_dataset_builder('rotten_tomatoes') >>> ds_builder.get_exported_dataset_info() DatasetInfo(description="Movie Review Dataset. This is a dataset of containing 5,331 positive and 5,331 negative processed sentences from Rotten Tomatoes movie reviews. This data was first used in Bo Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.'', Proceedings of the ACL, 2005. ", citation='@InProceedings{Pang+Lee:05a, author = {Bo Pang and Lillian Lee}, title = {Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales}, booktitle = {Proceedings of the ACL}, year = 2005 } ', homepage='http://www.cs.cornell.edu/people/pabo/movie-review-data/', license='', features={'text': Value(dtype='string', id=None), 'label': ClassLabel(num_classes=2, names=['neg', 'pos'], id=None)}, post_processed=None, supervised_keys=SupervisedKeysData(input='', output=''), task_templates=[TextClassification(task='text-classification', text_column='text', label_column='label')], builder_name='rotten_tomatoes_movie_review', config_name='default', version=1.0.0, splits={'train': SplitInfo(name='train', num_bytes=1074810, num_examples=8530, dataset_name='rotten_tomatoes_movie_review'), 'validation': SplitInfo(name='validation', num_bytes=134679, num_examples=1066, dataset_name='rotten_tomatoes_movie_review'), 'test': SplitInfo(name='test', num_bytes=135972, num_examples=1066, dataset_name='rotten_tomatoes_movie_review')}, download_checksums={'https://storage.googleapis.com/seldon-datasets/sentence_polarity_v1/rt-polaritydata.tar.gz': {'num_bytes': 487770, 'checksum': 'a05befe52aafda71d458d188a1c54506a998b1308613ba76bbda2e5029409ce9'}}, download_size=487770, post_processing_size=None, dataset_size=1345461, size_in_bytes=1833231) ``` )rgetr rCr)rGr=r=r>rs z(DatasetBuilder.get_exported_dataset_infoc Ks@d}|dkr|jr| r|jdk rL|jj|j}tjd|jd|jnrt|jdkrd|jd|jdjd}td t |jj d |d |jd}tj d |jd|jt |t r|jj|}|dko|jrtd |dt |jj |sR|dk r||d<d|krDt|drD|jrD|j|d<|jf|}nZtj|}xN|jD]B\}}|dk rft||std |d|dt|||qfW|jstd|j|j||d}||jk} | rtjd|nD||j|jkrtdt |jj |js8td |jd||fS)aCreate and validate BuilderConfig object as well as a unique config id for this config. Raises ValueError if there are multiple builder configs and name and DEFAULT_CONFIG_NAME are None. config_kwargs override the defaults kwargs in config Nz$No config specified, defaulting to: /r zload_dataset('z', 'rz')zEConfig name is missing. Please pick one among the available configs: z Example of usage: ``z6No config specified, defaulting to the single config: zBuilderConfig z not found. Available: rCrVVERSIONz doesn't have a 'z' key.z$BuilderConfig must have a name, got )rTz Using custom data configuration zvCannot name a custom BuilderConfig the same as an available BuilderConfig. Change the name. Available BuilderConfigs: z must have a version)BUILDER_CONFIGSDEFAULT_CONFIG_NAMEbuilder_configsrrrrCrlrFlistrPr}rErZhasattrrrrfdeepcopyrksetattrrsrV) rGrCrTrSbuilder_configZexample_of_usagekeyvaluerrZ is_customr=r=r>rsT          z%DatasetBuilder._create_builder_configcCsDdd|jD}t|t|jkr@dd|jD}td||S)z:Pre-defined list of configurations for this builder class.cSsi|] }||jqSr=)rC)rJr r=r=r>rYsz2DatasetBuilder.builder_configs..cSsg|] }|jqSr=)rC)rJr r=r=r> sz2DatasetBuilder.builder_configs..z5Names in BUILDER_CONFIGS must not be duplicated. Got )rrlrF)rZconfigsnamesr=r=r>rs zDatasetBuilder.builder_configscCs|jS)N)r)rGr=r=r>rzszDatasetBuilder.cache_dirTc Cs|jr&|jjddkr&|jjddnd}|dkr8|jn|d|j}|j}|j}|r`tjjnt j}|rv|||j }|r||t |jj }|r|rt |t r|||}|S)aoRelative path of this dataset in cache_dir: Will be: self.name/self.config.version/self.hash/ or if a repo_id with a namespace has been specified: self.namespace___self.name/self.config.version/self.hash/ If any of these element is missing or if ``with_version=False`` the corresponding subfolders are dropped. rrNZ___)rcountrrCr rmrrrjrrrrZrVrE) rG with_version with_hashis_local namespacebuilder_data_dirrrmrr=r=r>_relative_data_dirs*  z!DatasetBuilder._relative_data_dirc st|j }|rtjjntj}||j|jd|d||j|jd|d}fdd}ts|}|r|dd}||jjkrdt |d|j d |jd t |jjd }t j ||S) z2Return the data directory for the current version.F)rrTc sftjjsgSg}x@tjD]2}y|jtj||fWq tk rPYq Xq W|jdd|S)z"Returns previous versions on disk.T)reverse) rrrrappendr rurFsort)Zversion_dirnamesdir_name)rr=r>_other_versions_on_disk s   z@DatasetBuilder._build_cache_dir.._other_versions_on_diskrzFound a different version z of dataset z in cache_dir z". Using currently defined version r) r.rrrrjrrr rVrZrCrr)rGrrZversion_data_dirrZ version_dirsZ other_versionZwarn_msgr=)rr>rs    . zDatasetBuilder._build_cache_dircCstdS)a Construct the DatasetInfo object. See `DatasetInfo` for details. Warning: This function is only called once and the result is cached for all following .info() calls. Returns: info: (DatasetInfo) The dataset information N)NotImplementedError)rGr=r=r>r*s zDatasetBuilder._infocCstjjtjtj|S)z8Return the path of the module of this class or subclass.)rrdirnamergetfile getmodule)rr=r=r>r6sz&DatasetBuilder.get_imported_module_dir)download_config download_modeignore_verificationstry_from_hf_gcsrr|rc#Kst|p tj}| } |dk r |n|j}|dkr|dkr^t|jt|tjkt|tjkd|d}t|j||j j ||j s~|j p| ndd}nt |trd}||_t|j } | rtjj|j|jjtjdd} | rt| ntj| r>tjj|j} | r>|tjkr>tjd|jd|jd |j|_|j |dStjd |jd|jd | rt!|jj"pnd |jd st#d t$|jj"pd dt$|jj%pd dt$|jj&pd dt$|jj'pd d tj(dd} |jj"rLt)d|jj*d|jj+dt$|jj%dt$|jj&dt$|jj'dt$|jj"d|jdn&t)d|jj*d|jj+d|jd|j,|| |j}t-|d|d}|ry|j.|j/d}WnBt0t1fk rtjdYn t2k rtjdYnX|s|j3f|| d|t4dd|jj5j6D|j_&|j7|j_8|jj&|jj%|j_"|j9WdQRXWdQRX|j |t)d |jd!|jd"WdQRXdS)#aDownloads and prepares dataset for reading. Args: download_config (:class:`DownloadConfig`, optional): specific download configuration parameters. download_mode (:class:`DownloadMode`, optional): select the download/generate mode - Default to ``REUSE_DATASET_IF_EXISTS`` ignore_verifications (:obj:`bool`): Ignore the verifications of the downloaded/processed dataset information (checksums/size/splits/...) try_from_hf_gcs (:obj:`bool`): If True, it will try to download the already prepared dataset from the Hf google cloud storage dl_manager (:class:`DownloadManager`, optional): specific Download Manger to use base_path (:obj:`str`, optional): base path for relative paths that are used to download files. This can be a remote url. If not specified, the value of the `base_path` attribute (`self.base_path`) will be used instead. use_auth_token (:obj:`Union[str, bool]`, optional): Optional string or boolean to use as Bearer token for remote files on the Datasets Hub. If True, will get token from ~/.huggingface. **download_and_prepare_kwargs (additional keyword arguments): Keyword arguments. Example: ```py >>> from datasets import load_dataset_builder >>> builder = load_dataset_builder('rotten_tomatoes') >>> ds = builder.download_and_prepare() ``` NF)rzZforce_downloadZ force_extractZuse_etagr) dataset_namerrWr|record_checksumsrz.lockzReusing dataset z ()zGenerating dataset r) directoryzNot enough disk space. Needed: z (download: z , generated: z, post-processed: c sspt|r|Vn\|d}tj|ddz,|Vtjj|rDtj|tj||Wdtjj|rjtj|XdS)z4Create temporary dir for dirname and rename on exit.z .incompleteT)rN) r.rrrisdirshutilrmtreerenamer)rtmp_dirr=r=r>incomplete_dirs   z;DatasetBuilder.download_and_prepare..incomplete_dirz"Downloading and preparing dataset rz , total: z) to z...z to rTzJDataset not on Hf google storage. Downloading and preparing it from sourcezGHF google storage unreachable. Downloading and preparing it from source)r verify_infoscss|] }|jVqdS)N) num_bytes)rJrr=r=r>rMsz6DatasetBuilder.download_and_prepare..zDataset z downloaded and prepared to z(. Subsequent calls will reuse this data.):rZREUSE_DATASET_IF_EXISTSr|rrr[ZFORCE_REDOWNLOADrrCr rW$SKIP_CHECKSUM_COMPUTATION_BY_DEFAULTrrErrr.rrrrjrrrr/ contextlibZ nullcontextrrr _load_infor}"download_post_processing_resourcesr4 size_in_bytesOSErrorr7 download_size dataset_sizepost_processing_sizecontextmanagerprintrr{_check_manual_downloadr8_download_prepared_from_hf_gcsrrrConnectionError_download_and_preparesumsplitsriget_recorded_sizes_checksumsdownload_checksums _save_info)rGrrrrrr|rZdownload_and_prepare_kwargsrrrZ data_existsrZ tmp_data_dirZdownloaded_from_gcsr=r=r>download_and_prepare;s!        X ^$     z#DatasetBuilder.download_and_preparec CsJ|jdk rF|jdkrFttjd|jd|jjd|jd|jd dS)Nz The dataset z with config zp requires manual data. Please follow the manual download instructions: za Manual data can be loaded with: datasets.load_dataset("z$", data_dir=""))rZ manual_dirr@textwrapdedentrCr )rGrr=r=r>rsz%DatasetBuilder._check_manual_download)rc Cs|jddd}t|j|j}|j||tj|j}|jj|td|j t j d}x|jj D]}x|j |jD]p}t j |krtd|y,t|d|}tj|t jj|j|Wqttk rtjd|dYqtXqtWq`WtjddS) NTF)rrrz+Resources shouldn't be in a sub-directory: z Couldn't download resourse file z from Hf google storage.z*Dataset downloaded from Hf google storage.)rrrr}Zdownload_from_hf_gcsrrrnr rrrr_post_processing_resourcesrirFr-rmoverrjrr) rGrZrelative_data_dirreaderZdownloaded_infoZremote_cache_dirrresource_file_name resource_pathr=r=r>rs      z-DatasetBuilder._download_prepared_from_hf_gcsc KsVt|jd}|j|}|j|f|}|rB|jrBt|jj|jdx|D]}t |j jj dkrht dt jd|j jd|j|j y|j|f|Wntk r}z&td|jpdd t |d WYd d }~XnBtk r}z$t|j|jd |jd d d WYd d }~XnX|jqHW|r@t|jj|||j_|j|j_d S)aDownloads and prepares dataset for reading. This is the internal implementation to overwrite called when user calls `download_and_prepare`. It should download all required data and generate the pre-processed datasets files. Args: dl_manager: (DownloadManager) `DownloadManager` used to download and cache data. verify_infos: bool, if False, do not perform checksums and size tests. prepare_split_kwargs: Additional options. )rzdataset source filesrQz{`all` is a special split keyword corresponding to the union of all splits, so cannot be used as key in ._split_generator().z Generating z splitzCannot find data file. z Original error: Nz7To avoid duplicate keys, please fix the dataset script z.py)Zfix_msg)r)rC_make_split_generators_kwargs_split_generatorsrr1r}rrrZ split_infolowerrFradd_prepare_splitrrr%rZduplicate_key_indicesZmanage_extracted_filesr2rZdownloaded_sizer) rGrrprepare_split_kwargs split_dictsplit_generators_kwargsZsplit_generatorssplit_generatorer=r=r>rs:      z$DatasetBuilder._download_and_preparecCsx|jjD]}x|j|jD]p\}}tj|kr>td|tjj|j |}tjj |s|j |||}|rt jd|d|t j||qWq WdS)Nz+Resources shouldn't be in a sub-directory: z$Downloaded post-processing resource z as )r}rrrkrrrFrrjrr#_download_post_processing_resourcesrrr)rGrr resource_namerrZdownloaded_resource_pathr=r=r>r1s   z1DatasetBuilder.download_post_processing_resourcescCs tj|jS)N)rrr)rGr=r=r>r?szDatasetBuilder._load_infoc CsFtjj|j|jjtjdd}t||jj |jWdQRXdS)Nrz.lock) rrrjrrrrr/r}write_to_directory)rGrr=r=r>rBs  zDatasetBuilder._save_infoc CsVtjj|j|jjtjdd}t|$tf|j j |j ij |j WdQRXdS)Nrz.lock)rrrjrrrrr/r r rCr}r r)rGrr=r=r> _save_infosGs  zDatasetBuilder._save_infoscCs~iS)zFGet kwargs for `self._split_generators()` from `prepare_split_kwargs`.r=)rGrr=r=r>rLsz,DatasetBuilder._make_split_generators_kwargs)rrUcCstjj|js(td|jd|jdtjd|p>dj |j j d|j|dkrhdd |j j D}t t |j|||d |d tj d }t|trt|}|S) aReturn a Dataset for the specified split. Args: split (`datasets.Split`): Which subset of the data to return. run_post_process (bool, default=True): Whether to run post-processing dataset transforms and/or add indexes. ignore_verifications (bool, default=False): Whether to ignore the verifications of the downloaded/processed dataset information (checksums/size/splits/...). in_memory (bool, default=False): Whether to copy the data in-memory. Returns: datasets.Dataset Example: ```py >>> from datasets import load_dataset_builder >>> builder = load_dataset_builder('rotten_tomatoes') >>> ds = builder.download_and_prepare() >>> ds = builder.as_dataset(split='train') >>> ds Dataset({ features: ['text', 'label'], num_rows: 8530 }) ``` zDataset z: could not find data in z. Please make sure to call builder.download_and_prepare(), or pass download=True to datasets.load_dataset() before trying to access the Dataset object.zConstructing Dataset for split z, z, from NcSsi|] }||qSr=r=)rJsr=r=r>rYzsz-DatasetBuilder.as_dataset..)run_post_processr in_memoryT) map_tupleZ disable_tqdm)rrrrAssertionErrorrCrrdebugrjr}rr5r_build_single_datasetr,is_progress_bar_enabledrErwr)rGrr rr datasetsr=r=r> as_datasetQs$& zDatasetBuilder.as_dataset)rr rr csJ| }t|ts:t|}|dkr2djjjj}t|}j||d}|rFx.j |j D]}t j |kr^t d|q^Wfddj |jD}j||} | dk rF| }i} d} x$|jD]\} } t| }|| | <qW|o| r,jjdks jjjdkrd}njjjj|}t|| d jjdkrDtj_jjjdkr^ijj_| jjjt|<td d jjjj Dj_jjdk rȈjjdk rȈjjjjjjj_jjj|j_jj|j_jj|j_jjjdk rFjjjj|jjkr8t d jjjd |jnjjj|j_|S)zas_dataset for a single split.rQ+)rr z+Resources shouldn't be in a sub-directory: cs"i|]\}}tjjj||qSr=)rrrjr)rJrr)rGr=r>rYsz8DatasetBuilder._build_single_dataset..NFzpost processing resourcescss$|]}|jD]}|dVqqdS)rN)ri)rJZsplit_checksums_dictsZchecksums_dictr=r=r>rMsz7DatasetBuilder._build_single_dataset..z:Post-processed features info don't match the dataset: Got z but expected something like )rErrZrjr}rrPr( _as_datasetrrirrrFrk _post_processr0post_processedZresources_checksumsrr1r!rrrrrrrr~type)rGrr rr rZdsrresources_pathsrZrecorded_checksumsrrrZ size_checksumZexpected_checksumsr=)rGr>rs`             z$DatasetBuilder._build_single_dataset)rr rUcCsBt|j|jj|j||jjj|d}|j|}tfd|i|S)aConstructs a `Dataset`. This is the internal implementation to overwrite called when user calls `as_dataset`. It should read the pre-processed datasets files and generate the `Dataset` object. Args: split: `datasets.Split` which subset of the data to read. in_memory (bool, default False): Whether to copy the data in-memory. Returns: `Dataset` )rCZ instructionsZ split_infosr  fingerprint) rrr}readrCrri_get_dataset_fingerprintr )rGrr Zdataset_kwargsrr=r=r>rs  zDatasetBuilder._as_datasetcCs8t}|j|jjtjd|jt||j}|S)zThe dataset fingerprint is the hash of the relative directory dataset_name/config_name/version/hash, as well as the split specs.r)rrnrrrrrZro)rGrZhasherrr=r=r>rs z'DatasetBuilder._get_dataset_fingerprint)rr|rUcCst|ttfs td|jdt|p*|jt|jd|j|j j d}|j |dd|j |D}|dkrp|}n*||kr||}ntd|dt |t|j|d d }t|trt|}|S) NzBuilder z is not streamable.)r)r|rrrWcSsi|] }||jqSr=)rC)rJZsgr=r=r>rYsz7DatasetBuilder.as_streaming_dataset..z Bad split: z. Available splits: T)r)rEGeneratorBasedBuilderArrowBasedBuilderrFrCrr|rrr rWrrrr5_as_streaming_dataset_singlerwr)rGrr|rZsplits_generatorssplits_generatorrr=r=r>as_streaming_datasets*     z#DatasetBuilder.as_streaming_datasetcCs4|j|}|jr|j|jini}t||j|j|dS)N)r}rtoken_per_repo_id) _get_examples_iterable_for_splitrrr#r}rC)rGr!Z ex_iterabler#r=r=r>r s z+DatasetBuilder._as_streaming_dataset_single)datasetrrUcCsdS)z%Run dataset transforms or add indexesNr=)rGr%rr=r=r>rszDatasetBuilder._post_processcCsiS)z+Mapping resource_name -> resource_file_namer=)rGrr=r=r>r#sz)DatasetBuilder._post_processing_resources)rrrrUcCsdS)zPDownload the resource using the download manager and return the downloaded path.Nr=)rGrrrr=r=r>r'sz2DatasetBuilder._download_post_processing_resources)rcCs tdS)aSpecify feature dictionary generators and dataset splits. This function returns a list of `SplitGenerator`s defining how to generate data and what splits to use. Example:: return [ datasets.SplitGenerator( name=datasets.Split.TRAIN, gen_kwargs={'file': 'train_data.zip'}, ), datasets.SplitGenerator( name=datasets.Split.TEST, gen_kwargs={'file': 'test_data.zip'}, ), ] The above code will first call `_generate_examples(file='train_data.zip')` to write the train data, then `_generate_examples(file='test_data.zip')` to write the test data. Datasets are typically split into different subsets to be used at various stages of training and evaluation. Note that for datasets without a `VALIDATION` split, you can use a fraction of the `TRAIN` data for evaluation as you iterate on your model so as not to overfit to the `TEST` data. For downloads and extractions, use the given `download_manager`. Note that the `DownloadManager` caches downloads, so it is fine to have each generator attempt to download the source data. A good practice is to download all data in this function, and then distribute the relevant parts to each split with the `gen_kwargs` argument Args: dl_manager: (DownloadManager) Download manager to download the data Returns: `list`. N)r)rGrr=r=r>r-s,z DatasetBuilder._split_generators)rcKs tdS)aGenerate the examples and record them on disk. Args: split_generator: `SplitGenerator`, Split generator to process **kwargs: Additional kwargs forwarded from _download_and_prepare (ex: beam pipeline) N)r)rGrkwargsr=r=r>r[s zDatasetBuilder._prepare_split)rrUcCs tdS)zGenerate the examples on the fly. Args: split_generator: `SplitGenerator`, Split generator to process N)r)rGrr=r=r>r$fsz/DatasetBuilder._get_examples_iterable_for_split) NNNNNNNNNNry)NN)TTT)NNFTNNN)NTFF)F)NN)Ir:r;r<rtrrArrrrrrZrrrr[rrwrrrrpropertyr classmethodrrrrr3r6rrzrrabcabstractmethodrrrrrrrrrrrrr rr(r rrrrZTRAINrrrr#r"r rrrrrr*rr"r$r=r=r=r>rxs;^X F ( 8 >> A" !- rxcs`eZdZdZdZdZddfdd ZejddZ d d Z fd d Z e e d ddZZS)rawBase class for datasets with data generation based on dict generators. `GeneratorBasedBuilder` is a convenience class that abstracts away much of the data writing and reading of `DatasetBuilder`. It expects subclasses to implement generators of feature dictionaries across the dataset splits (`_split_generators`). See the method docstrings for details. TN)writer_batch_sizecstj|||p|j|_dS)N)superrDEFAULT_WRITER_BATCH_SIZE_writer_batch_size)rGr+argsr&) __class__r=r>rszGeneratorBasedBuilder.__init__cKs tdS)aWDefault function generating examples for each `SplitGenerator`. This function preprocess the examples from the raw data to the preprocessed dataset files. This function is called once for each `SplitGenerator` defined in `_split_generators`. The examples yielded here will be written on disk. Args: **kwargs (additional keyword arguments): Arguments forwarded from the SplitGenerator.gen_kwargs Yields: key: `str` or `int`, a unique deterministic example identification key. * Unique: An error will be raised if two examples are yield with the same key. * Deterministic: When generating the dataset twice, the same example should have the same key. Good keys can be the image id, or line number if examples are extracted from a text file. The key will be hashed and sorted to shuffle examples deterministically, such as generating the dataset multiple times keep examples in the same order. example: `dict`, a feature dictionary ready to be encoded and written to disk. The example will be encoded with `self.info.features.encode_example({...})`. N)r)rGr&r=r=r>_generate_examplessz(GeneratorBasedBuilder._generate_examplesc Cs|jjdk r|jj|j}n|j}|jd|jd}tjj|j|}|jf|j }t |jj ||j |j|dj}zTxNt j|d|jdt j d|jddD]"\}} |jj j| } |j| |qWWd|j\} } XWdQRX| |j_| |j_dS) Nrez.arrow)r~rr+Z hash_saltZcheck_duplicatesz examplesFz Generating z split)unittotalleavedisabledesc)r}rrCrrrrjrr1 gen_kwargsrr~r.r,tqdm num_examplesrencode_examplewritefinalizer) rGrcheck_duplicate_keysrfnamefpath generatorwriterrrecordZexampler9rr=r=r>rs4 z$GeneratorBasedBuilder._prepare_splitcstj|||ddS)N)r=)r,r)rGrr)r0r=r>rsz+GeneratorBasedBuilder._download_and_prepare)rrUcCst|j|jS)N)r"r1r7)rGrr=r=r>r$sz6GeneratorBasedBuilder._get_examples_iterable_for_split)r:r;r<rttest_dummy_datar-rr)r*r1rrr*r"r$ __classcell__r=r=)r0r>ros# rc@s:eZdZdZdZejddZddZe e ddd Z d S) rzaBase class for datasets with data generation based on Arrow loading functions (CSV/JSON/Parquet).TcKs tdS)aDefault function generating examples for each `SplitGenerator`. This function preprocess the examples from the raw data to the preprocessed dataset files. This function is called once for each `SplitGenerator` defined in `_split_generators`. The examples yielded here will be written on disk. Args: **kwargs (additional keyword arguments): Arguments forwarded from the SplitGenerator.gen_kwargs Yields: key: `str` or `int`, a unique deterministic example identification key. * Unique: An error will be raised if two examples are yield with the same key. * Deterministic: When generating the dataset twice, the same example should have the same key. Good keys can be the image id, or line number if examples are extracted from a text file. The key will be hashed and sorted to shuffle examples deterministically, such as generating the dataset multiple times keep examples in the same order. example: `pyarrow.Table`, a feature table ready to be encoded and written to disk. N)r)rGr&r=r=r>_generate_tablessz"ArrowBasedBuilder._generate_tablesc Cs|jd|jd}tjj|j|}|jf|j}t|jj |dB}x.t j |ddt j dD]\}}|j |q^W|j\}} WdQRX||j_| |j_|jj dkr|j|j_ dS)Nrez.arrow)r~rz tablesF)r2r4r5)rCrrrjrrEr7rr}r~r,r8rZ write_tabler<rr9rZ _features) rGrr>r?r@rArtabler9rr=r=r>rs z ArrowBasedBuilder._prepare_split)rrUcCstt|j|jdS)N)r&)r"r$rEr7)rGrr=r=r>r$sz2ArrowBasedBuilder._get_examples_iterable_for_splitN) r:r;r<rtrCr)r*rErr*r"r$r=r=r=r>rs rc@s eZdZdS)MissingBeamOptionsN)r:r;r<r=r=r=r>rG srGcsbeZdZdZdZdddfdd ZddZejd d Z fd d Z fd dZ ddZ Z S)BeamBasedBuilderzBeam based Builder.FN) beam_runner beam_optionscs$||_||_i|_tj||dS)N) _beam_runner _beam_options _beam_writersr,r)rGrIrJr/r&)r0r=r>rszBeamBasedBuilder.__init__cCs.i}tj|jjj}d|kr*|d|d<|S)Npipeline)rrrrrP)rGrrZsplit_generators_arg_namesr=r=r>rs  z.BeamBasedBuilder._make_split_generators_kwargscKs tdS)a"Build the beam pipeline examples for each `SplitGenerator`. This function extracts examples from the raw data with parallel transforms in a Beam pipeline. It is called once for each `SplitGenerator` defined in `_split_generators`. The examples from the PCollection will be encoded and written to disk. Warning: When running in a distributed setup, make sure that the data which will be read (download_dir, manual_dir,...) and written (cache_dir) can be accessed by the workers jobs. The data should be located in a shared filesystem, like GCS. Args: pipeline ([`utils.beam_utils.BeamPipeline`]): Apache Beam pipeline. **kwargs (additional keyword arguments): Arguments forwarded from the SplitGenerator.gen_kwargs. Returns: `beam.PCollection`: Apache Beam PCollection containing the example to send to `self.info.features.encode_example(...)`. Example: ``` def _build_pcollection(pipeline, extracted_dir=None): return ( pipeline | beam.Create(gfile.io.listdir(extracted_dir)) | beam.Map(_process_file) ) ``` N)r)rGrNr&r=r=r>_build_pcollection%s#z#BeamBasedBuilder._build_pcollectioncs ddl}ddljj}|j}|j}| rT| rTd|jd|jjd}td|d|pb|j j j }d|j |j j j _|j||d}tj|d|d |j} | j| j} |jj} xP|jjD]B\} } |jjj| d }| j| j|\}}| | }||_||_qWdS) Nrzload_dataset('z', 'z', beam_runner='DirectRunner')a*Trying to generate a dataset using Apache Beam, yet no Beam Runner or PipelineOptions() has been provided in `load_dataset` or in the builder arguments. For big datasets it has to run on large-scale data processing tools like Dataflow, Spark, etc. More information about Apache Beam runners at https://beam.apache.org/documentation/runners/capability-matrix/ If you really want to run it locally because you feel like the Dataset is small enough, you can use the local beam runner called `DirectRunner` (you may run out of memory). Example of usage: `rF)runneroptions)rrN)r) apache_beamZdatasets.utils.beam_utilsr beam_utilsrKrLrCr rGrQZpipeline_optionsZPipelineOptionsZview_asZ TypeOptionsZpipeline_type_checkZ BeamPipeliner,rrunZwait_until_finishmetricsr}rrMrkZ MetricsFilterZwith_namespacer<queryr9r)rGrrbeamrSrIrJZ usage_examplerNZpipeline_resultsrUr split_name beam_writerZm_filterr9rr)r0r=r>rJs6   z&BeamBasedBuilder._download_and_preparecstjj|jrtjnzddl}|jjj }|j tjj |jt j }|jj|WdQRX|jjr|j tjj |jt j}|jj|WdQRXdS)Nr)rrrrr,rrRioZ filesystemsZ FileSystemscreaterjr ZDATASET_INFO_FILENAMEr}Z _dump_infolicenseZLICENSE_FILENAMEZ _dump_license)rGrWfsf)r0r=r>r~s  zBeamBasedBuilder._save_infocsddljj}jd|d}tjjj|}tjj ||jdj |<jj j j fdd}|||?B}dS)Nrrez.arrow)r~rrrzcs4j|fj}|djfdd?O}j|S)z+PTransformation which build a single split.ZEncodecs|d|dfS)Nrr r=)Zkey_ex)r:r=r>szMBeamBasedBuilder._prepare_split.._build_pcollection..)rOr7ZMapZwrite_from_pcollection)rNZpcoll_examples)rWrYr:rGrr=r>rOsz;BeamBasedBuilder._prepare_split.._build_pcollection) rRrrCrrrjrrr}r~rMr:Z ptransform_fn)rGrrNrXr>r?rOrr=)rWrYr:rGrr>rs   zBeamBasedBuilder._prepare_split)r:r;r<rtrCrrr)r*rOrrrrDr=r=)r0r>rHs % 4 rH)hrtr)rrfrrrrrrarZ dataclassesr functoolsrtypingrrrrrrr r Z arrow_datasetr Z arrow_readerr rrrrZ arrow_writerrrrDrrZ dataset_dictrrZdownload.download_configrZdownload.download_managerrrZdownload.mock_download_managerrZ#download.streaming_download_managerrr~rrrr}rr r!Ziterable_datasetr"r#r$Zkeyhashr%Znamingr&r'rr(r)r*Z streamingr+r,Zutils.file_utilsr-r.Zutils.filelockr/Zutils.info_utilsr0r1r2Zutils.py_utilsr3r4r5r6r7r8 get_loggerr:rrFr9 Exceptionr?r@rArxrrrGrHr=r=r=r>sj             `Ab: