Logrine
Synthetic data Evaluation Blog About Trust Careers Book a scoping call
← All articles
Data quality 11 October 2026 · 7 min read

Deduplication and diversity matter more than dataset size

When people ask for a synthetic dataset, the first question is usually how many examples. It is the wrong first question. Ten thousand examples that say the same thing in slightly different words teach a model about as much as a few hundred do, and they can cause harm beyond wasted compute.

What the deduplication research found

Lee and colleagues studied this in "Deduplicating Training Data Makes Language Models Better" (ACL 2022). They showed that common text datasets contain many near-duplicates and long repeated passages. Models trained on deduplicated data emitted memorized text roughly ten times less often, and reached the same or better accuracy in fewer training steps. Deduplication also reduced overlap between training and validation data, which affected more than 4% of the validation sets of standard datasets, so evaluation became more reliable.

Three separate benefits follow: less memorization, less wasted training, and cleaner evaluation.

Why synthetic data is especially exposed

Generators repeat themselves. They reuse the same openings, the same sentence structures and the same example names, and a template-driven prompt can yield thousands of paraphrases of one underlying case. Early instruction-generation work recognized this. The Self-Instruct method (Wang et al., ACL 2023) adds a newly generated instruction to its pool only if its ROUGE-L similarity to existing instructions stays below 0.7, which is a simple form of similarity filtering built into the generation loop.

Three levels of duplication

  1. Exact. The same text after normalization. Hash it and drop repeats. Cheap and always worth doing.
  2. Near-duplicate. Texts that differ by a few words. Methods based on n-gram overlap, such as MinHash with locality-sensitive hashing, find these at scale.
  3. Semantic. Different words, same meaning. Embed the examples and look for pairs above a similarity threshold, or cluster them and cap how many items each cluster may contribute. This is the level that catches paraphrased repeats.

Measure diversity directly

Removing duplicates is not the same as having variety. A few measurements make diversity visible:

Do not over-deduplicate

Similarity thresholds trade one error for another. Set too tight, they remove legitimate variants, such as the same question asked about different amounts or products. Set too loose, they leave the repetition in. Check the effect on the examples you care about, keep a minimum count per scenario, and judge the threshold by the downstream result rather than by the number of rows removed.

Deduplicate across boundaries, too

Duplicates inside a dataset are only part of the problem. Overlap between training and test data is what quietly inflates scores. For every engagement we check the new dataset against the client's existing training data and held-out tests, and check benchmarks against public corpora, and we report the rates.

The takeaway

Ask any data supplier for the duplication rate and the diversity measures, not just the row count. A smaller dataset that is varied and clean is usually the better purchase.

Sources
  • Lee, K., Ippolito, D., Nystrom, A., et al. "Deduplicating Training Data Makes Language Models Better." ACL 2022.
  • Wang, Y., Kordi, Y., Mishra, S., et al. "Self-Instruct: Aligning Language Models with Self-Generated Instructions." ACL 2023.
Keep reading
Model collapse: why synthetic data needs a real-data anchor
What belongs in a dataset card for synthetic data
Logrine

Synthetic datasets and evaluation benchmarks for teams shipping LLMs and AI agents, delivered with the evidence.

Product
Synthetic data Evaluation benchmarks Trust and data handling
Company
About Careers Contact hello@logrine.com
Resources
Blog Model collapse LLM-as-a-judge Dataset cards
© 2026 MaxxLabs. Logrine is a MaxxLabs product.
Bengaluru, India