Abstract

Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original passages but falls to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set. We then compare two ways of selecting generated examples over three rounds of generation and retraining. One uses source information. The other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models. The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment therefore separates source identity, criterion-facing selection, and recursive training outcome: neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Armstrong, J. (2026). Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution. https://omanscience.com/en/articles/source-identification-is-not-fitness-testing-measuring-the-limits-of-synthetic-data-attribution

MLA 9

Armstrong, Joss. "Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution." https://omanscience.com/en/articles/source-identification-is-not-fitness-testing-measuring-the-limits-of-synthetic-data-attribution.

Chicago (author–date)

Armstrong, Joss. 2026. "Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution." https://omanscience.com/en/articles/source-identification-is-not-fitness-testing-measuring-the-limits-of-synthetic-data-attribution.

Harvard

Armstrong, J. (2026) 'Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution', Available at: https://omanscience.com/en/articles/source-identification-is-not-fitness-testing-measuring-the-limits-of-synthetic-data-attribution.

Vancouver

Armstrong J. Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution. https://omanscience.com/en/articles/source-identification-is-not-fitness-testing-measuring-the-limits-of-synthetic-data-attribution

IEEE

J. Armstrong, "Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution," https://omanscience.com/en/articles/source-identification-is-not-fitness-testing-measuring-the-limits-of-synthetic-data-attribution.