Research Article

Assessing the role of key and value dimensionality in multi-head attention mechanisms

Number: Advanced Online Publication Early Pub Date: July 14, 2026
TR EN

Assessing the role of key and value dimensionality in multi-head attention mechanisms

Abstract

ContextThe transformer architecture has occupied a prominent position in recent research due to its attention-driven design; however, the role of asymmetric key and value dimensions within its attention mechanism remains largely underexplored in the literature despite their potential impact on model behavior and performance.

ObjectiveThis study aims to systematically analyze the effects of key architectural parameters in the Transformer architecture on neural machine translation performance. These parameters include asymmetric key and value dimensions within the multi-head attention mechanism, embedding dimension, number of attention heads and sentence length.

MethodThis study analyses the architectural components of the Transformer model across three neural machine translation datasets: English–Turkish, French–Turkish, and German–Turkish. Six fundamental hyperparameters—source and target embedding dimensions, key and value dimensions, the number of attention heads, and sentence length—were assessed at two levels each. To address the computational complexity of the Transformer's modular structure, a fractional factorial experimental design was employed, reducing the 64 possible configurations to 16 experimental runs. This approach significantly reduced the use of computational resources while ensuring a comprehensive analysis of both individual parameters and their pairwise interactions. All models were trained for 150 epochs using PyTorch with Distributed Data Parallel (DDP) and Automated Mixed Precision (AMP) in an HPC environment with NVIDIA A100 GPUs. Finally, statistical evaluation and performance visualization were conducted using Minitab.

ResultsThe results show that while establishing equivalent key and value dimensions serves as a stable default, it does not guarantee optimal performance across linguistic contexts. The model’s response to asymmetrical configurations was highly language-specific, with notable variability observed in the interaction between the key dimension and the number of attention heads. Furthermore, increasing the target embedding dimension consistently improved generalization, whereas a larger source embedding dimension led to overfitting across all language pairs. These effects resulted from specific combinations of parameters that improved training accuracy but reduced validation performance, reflecting limitations in the generalization of language-dependent processes.

ConclusionThese results emphasize that optimal Transformer performance cannot be achieved through a universal hyperparameter configuration and must instead be tailored to the linguistic characteristics of each dataset.

Keywords

Supporting Institution

This work has been supported by the Scientific Research Projects Coordination Unit of the Sivas University of Science and Technology.

Project Number

2024-DTP-Müh-0004

Ethical Statement

Ethics committee approval was not required for this study because of there was no study on animals or humans.

Thanks

This work has been supported by the Scientific Research Projects Coordination Unit of the Sivas University of Science and Technology. Project Number: 2024-DTP-Müh-0004. Computing resources used in this work were provided by the National Center for High Performance Computing of Turkey (UHeM) under grant number 5020092024. The research utilized computational resources provided by the TUBITAK ULAKBIM High Performance and Grid Computing Center (TRUBA) and the Lütfi Albay Artificial Intelligence and Robotics Laboratory at Sivas University of Science and Technology.

References

  1. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, “Attention is all you need”, Proceedings of the 31st Annual Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, USA, 04-09 December 2017.
  2. Y. Zhang, M. X. Tuo, Q. Y. Yin, L. Qi, X. X. Wang, T. Liu, “Keywords extraction with deep neural network model”, Neurocomputing, 383, 113–121, 2020. https://doi.org/10.1016/j.neucom.2019.11.083.
  3. J. Devlin, M. W. Chang, K. Lee, K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Proceedings of the 17th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Minneapolis, MN, USA, 02-07 June 2019, 4171-4186.
  4. Y. Guan, J. Whitehill, “Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation”, arXiv, 2025. https://doi.org/10.48550/arXiv.2509.17930.
  5. J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, S. Sanghai, “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023), Singapore, Singapore, 06-10 December 2023, 4895-4901. https://doi.org/10.18653/v1/2023.emnlp-main.298.
  6. N. Kalchbrenner, P. Blunsom, “Recurrent continuous translation models”, Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP 2013), Seattle, Washington, USA, 18-21 October 2013, 1700-1709. https://doi.org/10.18653/v1/d13-1176.
  7. I. Sutskever, O. Vinyals, Q. V. Le, “Sequence to sequence learning with neural networks”, Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 2014), Montreal, Canada, 08-13 December 2014.
  8. D. Bahdanau, K. Cho, Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate”, arXiv, 2014. https://doi.org/10.48550/arXiv.1409.0473.

Details

Primary Language

English

Subjects

Big Data, Natural Language Processing

Journal Section

Research Article

Early Pub Date

July 14, 2026

Publication Date

-

Submission Date

March 31, 2026

Acceptance Date

June 26, 2026

Published in Issue

Year 2026 Number: Advanced Online Publication

APA
Katırcı, R., Çelik, H., & Gön, N. (2026). Assessing the role of key and value dimensionality in multi-head attention mechanisms. Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi, Advanced Online Publication. https://doi.org/10.65206/pajes.1920115
AMA
1.Katırcı R, Çelik H, Gön N. Assessing the role of key and value dimensionality in multi-head attention mechanisms. Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi. 2026;(Advanced Online Publication). doi:10.65206/pajes.1920115
Chicago
Katırcı, Ramazan, Hilal Çelik, and Nusret Gön. 2026. “Assessing the Role of Key and Value Dimensionality in Multi-Head Attention Mechanisms”. Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi, no. Advanced Online Publication. https://doi.org/10.65206/pajes.1920115.
EndNote
Katırcı R, Çelik H, Gön N (July 1, 2026) Assessing the role of key and value dimensionality in multi-head attention mechanisms. Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi Advanced Online Publication
IEEE
[1]R. Katırcı, H. Çelik, and N. Gön, “Assessing the role of key and value dimensionality in multi-head attention mechanisms”, Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi, no. Advanced Online Publication, July 2026, doi: 10.65206/pajes.1920115.
ISNAD
Katırcı, Ramazan - Çelik, Hilal - Gön, Nusret. “Assessing the Role of Key and Value Dimensionality in Multi-Head Attention Mechanisms”. Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi. Advanced Online Publication (July 1, 2026). https://doi.org/10.65206/pajes.1920115.
JAMA
1.Katırcı R, Çelik H, Gön N. Assessing the role of key and value dimensionality in multi-head attention mechanisms. Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi. 2026. doi:10.65206/pajes.1920115.
MLA
Katırcı, Ramazan, et al. “Assessing the Role of Key and Value Dimensionality in Multi-Head Attention Mechanisms”. Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi, no. Advanced Online Publication, July 2026, doi:10.65206/pajes.1920115.
Vancouver
1.Ramazan Katırcı, Hilal Çelik, Nusret Gön. Assessing the role of key and value dimensionality in multi-head attention mechanisms. Pamukkale Üniversitesi Mühendislik Bilimleri Dergisi. 2026 Jul. 1;(Advanced Online Publication). doi:10.65206/pajes.1920115