A dataset without documentation is a file of unknown origin. You cannot tell what it is for, how it was made, or when it will mislead you. For synthetic data the gap is wider, because the generation process is part of what the data is. A good dataset card closes that gap before the data is used.
Where the idea comes from
Gebru and colleagues proposed "datasheets for datasets" in work first circulated in 2018 and published in Communications of the ACM in 2021. Their framework is a set of questions covering a dataset's motivation, composition, collection process, preprocessing and labeling, recommended uses, distribution and maintenance. It was modelled on the datasheets that accompany electronic components, and it sits alongside Mitchell et al.'s "model cards" for documenting trained models. Dataset cards on public model hubs follow the same spirit.
The standard sections
- Purpose. The task the dataset was built for, and who it was built for.
- Composition. What a record contains, the schema, the label set, the size and the splits.
- Collection and processing. How the data was obtained, cleaned and labeled.
- Uses. Intended uses, and uses it should not be put to.
- Maintenance. Who maintains it, how to report problems, and how updates are versioned.
What synthetic data adds
Generated data needs sections that real-world collections do not:
- Generation method. The generator models and versions, prompt templates, seeds and sampling settings, so the process can be understood and repeated.
- Seed provenance. Where seed examples came from, whether they contained personal data, and how they were de-identified.
- Synthetic-to-real ratio. How much of the data is generated and how much is real, since this affects the risks described in our article on model collapse.
- Filtering log. Each filter applied, its threshold, and how many records it removed.
- Quality evidence. The human audit sample size and results, judge calibration, duplication rate, diversity measures, and comparison against real data.
- Terms of use. The terms attached to the output of the generator models vary between providers, so record which apply to this dataset.
- Known limitations. Scenarios that are thin, languages that were reviewed less thoroughly, and biases that were measured but not removed.
- Version history. What changed in each release, so results from different versions stay comparable.
How to read one as a buyer
A short test: can you find the answers to these four questions in a minute? What real data is this anchored to? How was quality measured, on how many examples, and by whom? What is it known to be bad at? Could you regenerate or extend it? A card that cannot answer them describes a dataset you cannot yet trust.
The takeaway
Documentation is a quality control, not a formality. Writing down the limitations is also the clearest sign that the people who made the dataset measured it.
Sources
- Gebru, T., Morgenstern, J., Vecchione, B., et al. "Datasheets for Datasets." Communications of the ACM, 2021 (first circulated 2018).
- Mitchell, M., Wu, S., Zaldivar, A., et al. "Model Cards for Model Reporting." Proceedings of FAT*, 2019.