When people ask for a synthetic dataset, the first question is usually how many examples. It is the wrong first question. Ten thousand examples that say the same thing in slightly different words teach a model about as much as a few hundred do, and they can cause harm beyond wasted compute.
Lee and colleagues studied this in "Deduplicating Training Data Makes Language Models Better" (ACL 2022). They showed that common text datasets contain many near-duplicates and long repeated passages. Models trained on deduplicated data emitted memorized text roughly ten times less often, and reached the same or better accuracy in fewer training steps. Deduplication also reduced overlap between training and validation data, which affected more than 4% of the validation sets of standard datasets, so evaluation became more reliable.
Three separate benefits follow: less memorization, less wasted training, and cleaner evaluation.
Generators repeat themselves. They reuse the same openings, the same sentence structures and the same example names, and a template-driven prompt can yield thousands of paraphrases of one underlying case. Early instruction-generation work recognized this. The Self-Instruct method (Wang et al., ACL 2023) adds a newly generated instruction to its pool only if its ROUGE-L similarity to existing instructions stays below 0.7, which is a simple form of similarity filtering built into the generation loop.
Removing duplicates is not the same as having variety. A few measurements make diversity visible:
Similarity thresholds trade one error for another. Set too tight, they remove legitimate variants, such as the same question asked about different amounts or products. Set too loose, they leave the repetition in. Check the effect on the examples you care about, keep a minimum count per scenario, and judge the threshold by the downstream result rather than by the number of rows removed.
Duplicates inside a dataset are only part of the problem. Overlap between training and test data is what quietly inflates scores. For every engagement we check the new dataset against the client's existing training data and held-out tests, and check benchmarks against public corpora, and we report the rates.
Ask any data supplier for the duplication rate and the diversity measures, not just the row count. A smaller dataset that is varied and clean is usually the better purchase.
Synthetic datasets and evaluation benchmarks for teams shipping LLMs and AI agents, delivered with the evidence.