Beyond Aggregate Fidelity: A Comparative Study of Distributional, Downstream-Utility, and Per-Instance Measures for Evaluating Synthetic Time-Series Data

DOI:

https://doi.org/10.63646/LVFT3833

Categories

Keywords:

Synthetic time series; generative models; evaluation measures; downstream utility; per-instance quality; data augmentation

Abstract

Synthetic time series produced by generative models are now widely used for data augmentation, privacy-preserving data sharing, and the simulation of rare operating conditions, yet their value depends entirely on how faithfully and usefully they stand in for real data. Deciding whether a synthetic dataset is good, however, remains unsettled: a broad and growing family of evaluation measures exists, spanning distributional fidelity, downstream utility, diversity and coverage, per-instance realism, and privacy leakage, but each measure emphasises a different facet of quality, and the most widely used measures are aggregate, computed over a whole set, and post-hoc, requiring access to real data. This article assembles a structured taxonomy of these measures and applies a common panel of them, on equal footing, to several generators of physiological sensor time series. Three findings recur. Measures disagree on which generator is best, so the verdict depends on the measure chosen; distributional fidelity is only weakly related to downstream utility, so a realistic-looking dataset need not be a useful one; and aggregate scores conceal a heavy tail of low-quality individual samples that a single average cannot reveal. We argue that evaluating synthetic time series calls for a multi-faceted, per-instance, and task-aware protocol rather than any single score, and we set out practical guidance for assembling one.

How to Cite

Zheng, L., Luo, Z., & Gao, Y. (2026). Beyond Aggregate Fidelity: A Comparative Study of Distributional, Downstream-Utility, and Per-Instance Measures for Evaluating Synthetic Time-Series Data. Data Science & Big Data Technology, 1(1), 241-262. https://doi.org/10.63646/LVFT3833

References

Alaa, A., van Breugel, B., Saveliev, E. S., & van der Schaar, M. (2022). How faithful is your synthetic data? Sample-level metrics for evaluating and auditing generative models. In Proceedings of the 39th International Conference on Machine Learning (pp. 290–306). PMLR.

Borji, A. (2019). Pros and cons of GAN evaluation measures. Computer Vision and Image Understanding, 179, 41–65. https://doi.org/10.1016/j.cviu.2018.10.009

Borji, A. (2022). Pros and cons of GAN evaluation measures: New developments. Computer Vision and Image Understanding, 215, 103329. https://doi.org/10.1016/j.cviu.2021.103329

Brophy, E., Wang, Z., She, Q., & Ward, T. (2023). Generative adversarial networks in time series: A systematic literature review. ACM Computing Surveys, 55(10), 1–31. https://doi.org/10.1145/3559540

Dankar, F. K., & Ibrahim, M. (2021). Fake it till you make it: Guidelines for effective synthetic data generation. Applied Sciences, 11(5), 2158. https://doi.org/10.3390/app11052158

Esteban, C., Hyland, S. L., & Rätsch, G. (2017). Real-valued (medical) time series generation with recurrent conditional GANs. arXiv preprint arXiv:1706.02633.

Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In Advances in Neural Information Processing Systems (Vol. 27, pp. 2672–2680).

Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., & Smola, A. (2012). A kernel two-sample test. Journal of Machine Learning Research, 13(25), 723–773.

Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., & Hochreiter, S. (2017). GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing Systems (Vol. 30, pp. 6626–6637).

Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (Vol. 33, pp. 6840–6851).

Iqbal, T., & Qureshi, S. (2022). The survey: Text generation models in deep learning. Journal of King Saud University - Computer and Information Sciences, 34(6), 2515–2528. https://doi.org/10.1016/j.jksuci.2020.04.001

Ismail Fawaz, H., Forestier, G., Weber, J., Idoumghar, L., & Muller, P. A. (2019). Deep learning for time series classification: A review. Data Mining and Knowledge Discovery, 33(4), 917–963. https://doi.org/10.1007/s10618-019-00619-1

Jordon, J., Yoon, J., & van der Schaar, M. (2019). PATE-GAN: Generating synthetic data with differential privacy guarantees. In International Conference on Learning Representations.

Kingma, D. P., & Welling, M. (2014). Auto-encoding variational Bayes. In International Conference on Learning Representations.

Kynkäänniemi, T., Karras, T., Laine, S., Lehtinen, J., & Aila, T. (2019). Improved precision and recall metric for assessing generative models. In Advances in Neural Information Processing Systems (Vol. 32, pp. 3927–3936).

Lopez-Paz, D., & Oquab, M. (2017). Revisiting classifier two-sample tests. In International Conference on Learning Representations.

Lu, W., Lu, Y., Li, J., Sigov, A., Ratkin, L., & Ivanov, L. A. (2024). Quantum machine learning: Classifications, challenges, and solutions. Journal of Industrial Information Integration, 42, 100736. https://doi.org/10.1016/j.jii.2024.100736

Lu, Y. (2019). Artificial intelligence: A survey on evolution, models, applications and future trends. Journal of Management Analytics, 6(1), 1–29. https://doi.org/10.1080/23270012.2019.1570365

Lu, Y., & Xu, L. D. (2019). Internet of Things (IoT) cybersecurity research: A review of current research topics. IEEE Internet of Things Journal, 6(2), 2103–2115. https://doi.org/10.1109/JIOT.2018.2869847

Lucic, M., Kurach, K., Michalski, M., Gelly, S., & Bousquet, O. (2018). Are GANs created equal? A large-scale study. In Advances in Neural Information Processing Systems (Vol. 31, pp. 700–709).

Naeem, M. F., Oh, S. J., Uh, Y., Choi, Y., & Yoo, J. (2020). Reliable fidelity and diversity metrics for generative models. In Proceedings of the 37th International Conference on Machine Learning (pp. 7176–7185). PMLR.

Sajjadi, M. S. M., Bachem, O., Lucic, M., Bousquet, O., & Gelly, S. (2018). Assessing generative models via precision and recall. In Advances in Neural Information Processing Systems (Vol. 31, pp. 5228–5237).

Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., & Chen, X. (2016). Improved techniques for training GANs. In Advances in Neural Information Processing Systems (Vol. 29, pp. 2234–2242).

Stenger, M., Leppich, R., Foster, I., Kounev, S., & Bauer, A. (2024). Evaluation is key: A survey on evaluation measures for synthetic time series. Journal of Big Data, 11(1), 66. https://doi.org/10.1186/s40537-024-00924-7

Theis, L., van den Oord, A., & Bethge, M. (2016). A note on the evaluation of generative models. In International Conference on Learning Representations.

van Breugel, B., Kyono, T., Berrevoets, J., & van der Schaar, M. (2021). DECAF: Generating fair synthetic data using causally-aware generative networks. In Advances in Neural Information Processing Systems (Vol. 34, pp. 22221–22233).

Wang, Z., She, Q., & Ward, T. E. (2021). Generative adversarial networks in computer vision: A survey and taxonomy. ACM Computing Surveys, 54(2), 1–38. https://doi.org/10.1145/3439723

Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. In Advances in Neural Information Processing Systems (Vol. 32, pp. 7335–7345).

Yoon, J., Jarrett, D., & van der Schaar, M. (2019). Time-series generative adversarial networks. In Advances in Neural Information Processing Systems (Vol. 32, pp. 5508–5518).

Zhang, C., & Lu, Y. (2021). Study on artificial intelligence: The state of the art and future prospects. Journal of Industrial Information Integration, 23, 100224. https://doi.org/10.1016/j.jii.2021.100224