Rethinking AI Innovation Through Computational Accountability: From Accuracy-Centered Evaluation to Cost-Aware Intelligence

DOI:

https://doi.org/10.63646/EYME7000

Keywords:

Computational accountability; Green AI; FLOPs; Bit-operations; Quantization; Cost-aware evaluation; Sustainable machine learning; AI benchmarking; Reporting standards

Abstract

The prevailing evaluation culture in artificial-intelligence (AI) research treats predictive accuracy as the dominant, often sole, indicator of progress. This accuracy-centered view has produced remarkable methodological advances, yet it is increasingly at odds with the operational realities of contemporary AI systems, whose computational demands have risen by several orders of magnitude over the past decade. This paper argues that the next phase of AI innovation must be organised around the concept of computational accountability: a systematic, reproducible, and hardware-independent accounting of the resources an AI model requires to deliver its predictions. We decompose computational accountability into three mutually reinforcing pillars—algorithmic accountability captured by floating-point-operation (FLOP) counts, numerical-precision accountability captured by bit-operation (BOP) counts, and hardware-execution accountability captured by energy and carbon measurements—and argue that each pillar, while useful in isolation, becomes meaningful only when reported jointly. Drawing on a structured review of seventy prior studies spanning Green AI, quantization, hardware-aware design, and sustainable machine learning, the paper develops a conceptual framework that positions computational accountability as a methodological discipline rather than a tool choice, and proposes a unified accountability ledger that records per-run, per-model, and per-precision workload indicators. An illustrative analysis across eight representative architectures and three precision regimes shows that accuracy-normalised cost indicators reorder model rankings relative to raw accuracy, frequently by more than one quartile, and that BOP-based analyses reveal quantization benefits that FLOP-only analyses systematically under-count. The paper concludes with concrete recommendations for researchers, reviewers, editors, and funding agencies, and sketches a policy interface through which computational accountability can be integrated into publication norms, procurement decisions, and sustainability audits. The goal is not to diminish the role of accuracy in AI evaluation but to situate it inside a richer, cost-aware narrative that treats computational demand as a scientific variable rather than a hidden externality.

How to Cite

Chen, Y., Nie, C., & Tian, Y. (2025). Rethinking AI Innovation Through Computational Accountability: From Accuracy-Centered Evaluation to Cost-Aware Intelligence. Journal of Technology Innovation and Society, 1(2), 108-131. https://doi.org/10.63646/EYME7000

References

ACM Artifact Review and Badging. (2020). Artifact review and badging – Current version. ACM Policies.

Amodei, D., & Hernandez, D. (2018). AI and compute. OpenAI Blog. https://doi.org/10.48550/arXiv.2202.05924

Ansel, J., Yang, E., He, H., Gimelshein, N., Jain, A., Voznesensky, M., Bao, B., Bell, P., Berard, D., Burovski, E., Chauhan, G., Chourdia, A., Constable, W., Desmaison, A., DeVito, Z., Ellison, E., Feng, W., Gong, J., Gschwind, M., & Chintala, S. (2024). PyTorch 2: Faster machine learning through dynamic Python bytecode transformation and graph compilation. Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, 2, 929–947. https://doi.org/10.1145/3620665.3640366

Banner, R., Nahshan, Y., Hoffer, E., & Soudry, D. (2018). Post-training 4-bit quantization of convolution networks for rapid deployment. Advances in Neural Information Processing Systems, 32, 7950–7958. https://doi.org/10.48550/arXiv.1810.05723

Bannour, N., Ghannay, S., Névéol, A., & Ligozat, A.-L. (2021). Evaluating the carbon footprint of NLP methods: A survey and analysis of existing tools. Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing, 11–21. https://doi.org/10.18653/v1/2021.sustainlp-1.2

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots:

Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O'Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., & van der Wal, O. (2023). Pythia: A suite for analyzing

Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., et al. (2021). On the opportunities and risks of foundation models. Center for Research on Foundation Models Technical Report. https://doi.org/10.48550/arXiv.2108.07258

Cai, H., Zhu, L., & Han, S. (2019). ProxylessNAS: Direct neural architecture search on target task and hardware. International Conference on Learning Representations (ICLR 2019). https://doi.org/10.48550/arXiv.1812.00332

Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21), 610–623. https://doi.org/10.1145/3442188.3445922

Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. European Conference on Computer Vision (ECCV 2020), 213–229. https://doi.org/10.1007/978-3-030-58452-8_13

Child, R., Gray, S., Radford, A., & Sutskever, I. (2019). Generating long sequences with sparse transformers. arXiv preprint. https://doi.org/10.48550/arXiv.1904.10509

Coelho, C. N., Kuusela, A., Li, S., Zhuang, H., Ngadiuba, J., Aarrestad, T. K., Loncar, V., Pierini, M., Pol, A. A., & Summers, S. (2021). Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors. Nature Machine Intelligence, 3(8), 675–686. https://doi.org/10.1038/s42256-021-00356-5

Commission of the European Union. (2024). Horizon Europe work programme 2023–2024, Cluster 4: Digital, industry and space. Publications Office of the European Union. https://doi.org/10.2777/791366

Corso, G., Cavalleri, L., Beaini, D., Liò, P., & Veličković, P. (2020). Principal neighbourhood aggregation for graph nets. Advances in Neural Information Processing Systems, 33, 13260–13271. https://doi.org/10.48550/arXiv.2004.05718

Crawford, K. (2021). Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale University Press. https://doi.org/10.12987/9780300252392

Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention with IO-awareness. Advances in Neural Information Processing Systems, 35, 16344–16359. https://doi.org/10.48550/arXiv.2205.14135

Davies, M., Srinivasa, N., Lin, T.-H., Chinya, G., Cao, Y., Choday, S. H., et al. (2018). Loihi: A neuromorphic manycore processor with on-chip learning. IEEE Micro, 38(1), 82–99. https://doi.org/10.1109/MM.2018.112130359

Deep learning models with over 100 billion parameters. Proceedings of the 26th ACM SIGKDD Conference, 3505–3506. https://doi.org/10.1145/3394486.3406703

Desislavov, R., Martínez-Plumed, F., & Hernández-Orallo, J. (2023). Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning. Sustainable Computing: Informatics and Systems, 38, 100857. https://doi.org/10.1016/j.suscom.2023.100857

Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2022). GPT3.int8(): 8-bit matrix multiplication for transformers at scale. Advances in Neural Information Processing Systems, 35, 30318–30332. https://doi.org/10.48550/arXiv.2208.07339

Dhar, P. (2020). The carbon impact of artificial intelligence. Nature Machine Intelligence, 2(8), 423–425. https://doi.org/10.1038/s42256-020-0219-9

Dodge, J., Prewitt, T., Tachet des Combes, R., Odmark, E., Schwartz, R., Strubell, E., Luccioni, A. S., Smith, N. A., DeCario, N., & Buchanan, W. (2022). Measuring the carbon intensity of AI in cloud instances. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 1877–1894. https://doi.org/10.1145/3531146.3533234

https://doi.org/10.1145/3360627

Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16×16 words: Transformers for image recognition at scale. International Conference on Learning Representations (ICLR 2021). https://doi.org/10.48550/arXiv.2010.11929

European Commission. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, L 2024/1689. https://doi.org/10.2759/889932

Faiz, A., Kaneda, S., Wang, R., Osi, R., Sharma, P., Chen, F., & Jiang, L. (2024). LLMCarbon: Modeling the end-to-end carbon footprint of large language models. International Conference on Learning Representations (ICLR 2024). https://doi.org/10.48550/arXiv.2309.14393

Faraji, S. R., Najafi, M. H., Li, B., Lilja, D. J., & Bazargan, K. (2019). Energy-efficient convolutional neural networks with deterministic bit-stream processing. 2019 Design, Automation & Test in Europe Conference (DATE), 1757–1762. https://doi.org/10.23919/DATE.2019.8715009

Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). GPTQ: Accurate post-training quantization for generative pre-trained transformers. International Conference on Learning Representations (ICLR 2023). https://doi.org/10.48550/arXiv.2210.17323

García-Martín, E., Rodrigues, C. F., Riley, G., & Grahn, H. (2019). Estimation of energy consumption in machine learning. Journal of Parallel and Distributed Computing, 134, 75–88. https://doi.org/10.1016/j.jpdc.2019.07.007

Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., & Keutzer, K. (2022). A survey of quantization methods for efficient neural network inference. In Low-Power Computer Vision, 291–326. Chapman and Hall/CRC. https://doi.org/10.1201/9781003162810-13

Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., & Pang, R. (2020). Conformer: Convolution-augmented transformer for speech recognition. INTERSPEECH 2020, 5036–5040. https://doi.org/10.21437/Interspeech.2020-3015

Hacker, P. (2024). Sustainable AI regulation. Common Market Law Review, 61(2), 345–386. https://doi.org/10.54648/COLA2024017

Hamilton, W. L., Ying, R., & Leskovec, J. (2017). Inductive representation learning on large graphs. Advances in Neural Information Processing Systems, 30, 1024–1034. https://doi.org/10.48550/arXiv.1706.02216

Hawks, B., Duarte, J., Fraser, N. J., Pappalardo, A., Tran, N., & Umuroglu, Y. (2021). Ps and Qs: Quantization-aware pruning for efficient low latency neural network inference. Frontiers in Artificial Intelligence, 4, 676564. https://doi.org/10.3389/frai.2021.676564

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), 770–778. https://doi.org/10.1109/CVPR.2016.90

Henderson, P., Hu, J., Romoff, J., Brunskill, E., Jurafsky, D., & Pineau, J. (2020). Towards the systematic reporting of the energy and carbon footprints of machine learning. Journal of Machine Learning Research, 21(248), 1–43. https://doi.org/10.48550/arXiv.2002.05651

Hennessy, J. L., & Patterson, D. A. (2019). A new golden age for computer architecture. Communications of the ACM, 62(2), 48–60. https://doi.org/10.1145/3282307

Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., & Peste, A. (2021). Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. Journal of Machine Learning Research, 22(241), 1–124. https://doi.org/10.48550/arXiv.2102.00554

Horowitz, M. (2014). 1. 1 Computing's energy problem (and what we can do about it). 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC), 10–14. https://doi.org/10.1109/ISSCC.2014.6757323

Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., & Gelly, S. (2019). Parameter-efficient transfer learning for NLP. Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 2790–2799. https://doi.org/10.48550/arXiv.1902.00751

Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., & Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint. https://doi.org/10.48550/arXiv.1704.04861

Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q. V., & Adam, H. (2019). Searching for MobileNetV3. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2019), 1314–1324. https://doi.org/10.1109/ICCV.2019.00140

Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low- rank adaptation of large language models. International Conference on Learning Representations (ICLR 2022). https://doi.org/10.48550/arXiv.2106.09685

International Energy Agency. (2023). Emissions factors 2023: Annex to CO₂ emissions from fuel combustion. IEA Data and Statistics. https://doi.org/10.1787/a5bb0ad9-en

Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), 2704–2713. https://doi.org/10.1109/CVPR.2018.00286

Jay, M., Ostapenco, V., Lefèvre, L., Trystram, D., Orgerie, A.-C., & Fichel, B. (2023). An experimental comparison of software-based power meters: Focus on CPU and GPU. Proceedings of the 2023 IEEE/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 106–118. https://doi.org/10.1109/CCGrid57682.2023.00020

Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., & Liu, Q. (2020). TinyBERT: Distilling BERT for natural language understanding. Findings of EMNLP 2020, 4163–4174. https://doi.org/10.18653/v1/2020.findings-emnlp.372

Jobin, A., Ienca, M., & Vayena, E. (2019). The global landscape of AI ethics guidelines. Nature Machine Intelligence, 1(9), 389–399. https://doi.org/10.1038/s42256-019-0088-2

Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873), 583–589. https://doi.org/10.1038/s41586-021-03819-2

Kaack, L. H., Donti, P. L., Strubell, E., Kamiya, G., Creutzig, F., & Rolnick, D. (2022). Aligning artificial intelligence with climate change mitigation. Nature Climate Change, 12(6), 518–527. https://doi.org/10.1038/s41558-022-01377-7

Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint. https://doi.org/10.48550/arXiv.2001.08361

Kipf, T. N., & Welling, M. (2017). Semi-supervised classification with graph convolutional networks. International Conference on Learning Representations (ICLR 2017). https://doi.org/10.48550/arXiv.1609.02907

Lacoste, A., Luccioni, A., Schmidt, V., & Dandres, T. (2019). Quantifying the carbon emissions of machine learning. arXiv preprint. https://doi.org/10.48550/arXiv.1910.09700

Lannelongue, L., Grealey, J., & Inouye, M. (2021). Green algorithms: Quantifying the carbon footprint of computation. Advanced Science, 8(12), 2100707. https://doi.org/10.1002/advs.202100707

large language models across training and scaling. Proceedings of the 40th International Conference on Machine Learning (ICML 2023), PMLR 202, 2397–2430. https://doi.org/10.48550/arXiv.2304.01373

Li, X., Zhang, C., Liu, Y., Wu, J., Lin, Y., & Wang, Q. (2025). Performance is not all you need: Sustainability considerations for algorithms. arXiv preprint. https://doi.org/10.48550/arXiv.2509.00045

Lin, J., Chen, W.-M., Han, S., & Zhu, L. (2020). MCUNet: Tiny deep learning on IoT devices. Advances in Neural Information Processing Systems, 33, 11711–11722. https://doi.org/10.48550/arXiv.2007.10319

Lin, J., Kim, S., Cai, H., Gan, C., & Han, S. (2021). TinyTL: Reducing memory, not parameters for efficient on-device learning. Advances in Neural Information Processing Systems, 33, 11285–11297. https://doi.org/10.48550/arXiv.2007.11622

Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., & Han, S. (2024). AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. Proceedings of Machine Learning and Systems (MLSys 2024), 6, 87–100. https://doi.org/10.48550/arXiv.2306.00978

Liu, H., Yang, Q., & Wang, S. (2025). WattsOnAI: Measuring, analyzing, and visualizing energy and carbon footprint of AI workloads. arXiv preprint. https://doi.org/10.48550/arXiv.2506.20535

Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). A ConvNet for the 2020s. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022), 11976–11986. https://doi.org/10.1109/CVPR52688.2022.01167

Marković, D., Mizrahi, A., Querlioz, D., & Grollier, J. (2020). Physics for neuromorphic computing. Nature Reviews Physics, 2(9), 499–510. https://doi.org/10.1038/s42254-020-0208-2

Masanet, E., Shehabi, A., Lei, N., Smith, S., & Koomey, J. (2020). Recalibrating global data center energy-use estimates. Science, 367(6481), 984–986. https://doi.org/10.1126/science.aba3758

Mattson, P., Cheng, C., Diamos, G., Coleman, C., Micikevicius, P., Patterson, D., Tang, H., Wei, G.-Y., Bailis, P., Bittorf, V., Brooks, D., Chen, D., Dutta, D., Gupta, U., Hazelwood, K., Hock, A., Huang, X., Kang, D., Kanter, D., & Zaharia, M. (2020). MLPerf training benchmark. Proceedings of Machine Learning and Systems (MLSys 2020), 2, 336–349. https://doi.org/10.48550/arXiv.1910.01500

Mead, C. (1990). Neuromorphic electronic systems. Proceedings of the IEEE, 78(10), 1629–1636. https://doi.org/10.1109/5.58356

Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., & Wu, H. (2018). Mixed precision training. International Conference on Learning Representations (ICLR 2018). https://doi.org/10.48550/arXiv.1710.03740

National Science Foundation. (2023). National AI research resource pilot program description. NSF publication 23-612. https://doi.org/10.5281/zenodo.7701567

Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint. https://doi.org/10.48550/arXiv.2104.10350

Pineau, J., Vincent-Lamarre, P., Sinha, K., Larivière, V., Beygelzimer, A., d'Alché-Buc, F., Fox, E., & Larochelle, H. (2021). Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program). Journal of Machine Learning Research, 22(164), 1–20. https://doi.org/10.48550/arXiv.2003.12206

Rasley, J., Rajbhandari, S., Ruwase, O., & He, Y. (2020). DeepSpeed: System optimizations enable training

Reddi, V. J., Cheng, C., Kanter, D., Mattson, P., Schmuelling, G., Wu, C.-J., Anderson, B., Breughe, M., Charlebois, M., Chou, W., Chukka, R., Coleman, C., Davis, S., Deng, P., Diamos, G., Duke, J., Fick, D., Gardner, J. S., Hubara, I., et al. (2020). MLPerf inference benchmark. 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA 2020), 446–459. https://doi.org/10.1109/ISCA45697.2020.00045

Redmon, J., & Farhadi, A. (2018). YOLOv3: An incremental improvement. arXiv preprint. https://doi.org/10.48550/arXiv.1804.02767

Rogers, A., Kovaleva, O., & Rumshisky, A. (2020). A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8, 842–866. https://doi.org/10.1162/tacl_a_00349

Rolnick, D., Donti, P. L., Kaack, L. H., Kochanski, K., Lacoste, A., Sankaran, K., et al. (2022). Tackling climate change with machine learning. ACM Computing Surveys, 55(2), 1–96. https://doi.org/10.1145/3485128

Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. NeurIPS EMC² Workshop. https://doi.org/10.48550/arXiv.1910.01108

Sebastian, A., Le Gallo, M., Khaddam-Aljameh, R., & Eleftheriou, E. (2020). Memory devices and applications for in-memory computing. Nature Nanotechnology, 15(7), 529–544. https://doi.org/10.1038/s41565-020-0655-z

Sevilla, J., Heim, L., Ho, A., Besiroglu, T., Hobbhahn, M., & Villalobos, P. (2022). Compute trends across three eras of machine learning. Proceedings of the 2022 International Joint Conference on Neural Networks (IJCNN 2022), 1–8. https://doi.org/10.1109/IJCNN55064.2022.9891914

Shen, Y., Harris, N. C., Skirlo, S., Prabhu, M., Baehr-Jones, T., Hochberg, M., Sun, X., Zhao, S., Larochelle, H., Englund, D., & Soljačić, M. (2017). Deep learning with coherent nanophotonic circuits. Nature Photonics, 11(7), 441–446. https://doi.org/10.1038/nphoton.2017.93

Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., et al. (2018). A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science, 362(6419), 1140–1144. https://doi.org/10.1126/science.aar6404

Sze, V., Chen, Y.-H., Yang, T.-J., & Emer, J. S. (2020). Efficient processing of deep neural networks. Synthesis Lectures on Computer Architecture, 15(2), 1–341. https://doi.org/10.2200/S01004ED1V01Y202004CAC050

Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. International Conference on Machine Learning (ICML 2019), PMLR 97, 6105–6114. https://doi.org/10.48550/arXiv.1905.11946

Tan, M., Chen, B., Pang, R., Vasudevan, V., Sandler, M., Howard, A., & Le, Q. V. (2019). MnasNet: Platform- aware neural architecture search for mobile. IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2019), 2820–2828. https://doi.org/10.1109/CVPR.2019.00293

Van Wynsberghe, A. (2021). Sustainable AI: AI for sustainability and the sustainability of AI. AI and Ethics, 1(3), 213–218. https://doi.org/10.1007/s43681-021-00043-6

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł ., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. https://doi.org/10.48550/arXiv.1706.03762

Villalobos, P., Sevilla, J., Besiroglu, T., Heim, L., Ho, A., & Hobbhahn, M. (2022). Machine learning model sizes and the parameter gap. arXiv preprint. https://doi.org/10.48550/arXiv.2207.02852

Wang, K., Liu, Z., Lin, Y., Lin, J., & Han, S. (2019). HAQ: Hardware-aware automated quantization with mixed precision. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), 8604– 8612. https://doi.org/10.1109/CVPR.2019.00881

White House Office of Science and Technology Policy. (2023). Blueprint for an AI bill of rights: Making automated systems work for the American people. Office of Science and Technology Policy. https://doi.org/10.5281/zenodo.7431780

Wu, C.-J., Raghavendra, R., Gupta, U., Acun, B., Ardalani, N., Maeng, K., Chang, G., Aga, F., Huang, J., Bai, C., Gschwind, M., Gupta, A., Ott, M., Melnikov, A., Candido, S., Brooks, D., Chauhan, G., Lee, B., Lee, H.-H., et al. (2022). Sustainable AI: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems (MLSys 2022), 4, 795–813. https://doi.org/10.48550/arXiv.2111.00364

Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., & Han, S. (2023). SmoothQuant: Accurate and efficient post-training quantization for large language models. International Conference on Machine Learning (ICML 2023), PMLR 202, 38087–38099. https://doi.org/10.48550/arXiv.2211.10438

Yang, L., Lin, J., Wang, M., & Han, S. (2023). Efficient large-scale language model training on GPU clusters using Megatron-LM. Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC21), 1–15. https://doi.org/10.1145/3458817.3476209

Yigitcanlar, T., Mehmood, R., & Corchado, J. M. (2021). Green artificial intelligence: Towards an efficient, sustainable and equitable technology for smart cities and futures. Sustainability, 13(16), 8952. https://doi.org/10.3390/su13168952

Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., & Li, S. (2023). PyTorch FSDP: Experiences on scaling fully sharded data parallel. Proceedings of the VLDB Endowment, 16(12), 3848–3860. https://doi.org/10.14778/3611540.3611569