Assessing the role of key and value dimensionality in multi-head attention mechanisms
Öz
Context—The transformer architecture has occupied a prominent position in recent research due to its attention-driven design; however, the role of asymmetric key and value dimensions within its attention mechanism remains largely underexplored in the literature despite their potential impact on model behavior and performance.
Objective—This study aims to systematically analyze the effects of key architectural parameters in the Transformer architecture on neural machine translation performance. These parameters include asymmetric key and value dimensions within the multi-head attention mechanism, embedding dimension, number of attention heads and sentence length.
Method—This study analyses the architectural components of the Transformer model across three neural machine translation datasets: English–Turkish, French–Turkish, and German–Turkish. Six fundamental hyperparameters—source and target embedding dimensions, key and value dimensions, the number of attention heads, and sentence length—were assessed at two levels each. To address the computational complexity of the Transformer's modular structure, a fractional factorial experimental design was employed, reducing the 64 possible configurations to 16 experimental runs. This approach significantly reduced the use of computational resources while ensuring a comprehensive analysis of both individual parameters and their pairwise interactions. All models were trained for 150 epochs using PyTorch with Distributed Data Parallel (DDP) and Automated Mixed Precision (AMP) in an HPC environment with NVIDIA A100 GPUs. Finally, statistical evaluation and performance visualization were conducted using Minitab.
Results—The results show that while establishing equivalent key and value dimensions serves as a stable default, it does not guarantee optimal performance across linguistic contexts. The model’s response to asymmetrical configurations was highly language-specific, with notable variability observed in the interaction between the key dimension and the number of attention heads. Furthermore, increasing the target embedding dimension consistently improved generalization, whereas a larger source embedding dimension led to overfitting across all language pairs. These effects resulted from specific combinations of parameters that improved training accuracy but reduced validation performance, reflecting limitations in the generalization of language-dependent processes.
Conclusion—These results emphasize that optimal Transformer performance cannot be achieved through a universal hyperparameter configuration and must instead be tailored to the linguistic characteristics of each dataset.
Anahtar Kelimeler
- Asymmetrical Dimensions
- Deep Learning
- Hyperparameter Optimization
- Key dimension
- Multi-Head Attention
- Neural machine translation
- Transformer Architecture
- Value dimension
Destekleyen Kurum
Proje Numarası
Etik Beyan
Teşekkür
Kaynakça
- A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, “Attention is all you need”, Proceedings of the 31st Annual Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, USA, 04-09 December 2017.
- Y. Zhang, M. X. Tuo, Q. Y. Yin, L. Qi, X. X. Wang, T. Liu, “Keywords extraction with deep neural network model”, Neurocomputing, 383, 113–121, 2020. https://doi.org/10.1016/j.neucom.2019.11.083.
- J. Devlin, M. W. Chang, K. Lee, K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Proceedings of the 17th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Minneapolis, MN, USA, 02-07 June 2019, 4171-4186.
- Y. Guan, J. Whitehill, “Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation”, arXiv, 2025. https://doi.org/10.48550/arXiv.2509.17930.
- J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, S. Sanghai, “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023), Singapore, Singapore, 06-10 December 2023, 4895-4901. https://doi.org/10.18653/v1/2023.emnlp-main.298.
- N. Kalchbrenner, P. Blunsom, “Recurrent continuous translation models”, Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP 2013), Seattle, Washington, USA, 18-21 October 2013, 1700-1709. https://doi.org/10.18653/v1/d13-1176.
- I. Sutskever, O. Vinyals, Q. V. Le, “Sequence to sequence learning with neural networks”, Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 2014), Montreal, Canada, 08-13 December 2014.
- D. Bahdanau, K. Cho, Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate”, arXiv, 2014. https://doi.org/10.48550/arXiv.1409.0473.
Ayrıntılar
Birincil Dil
İngilizce
Konular
Büyük Veri, Doğal Dil İşleme
Bölüm
Araştırma Makalesi
Yazarlar
Ramazan Katırcı
0000-0003-2448-011X
Türkiye
Hilal Çelik
*
0000-0001-5428-3411
Türkiye
Nusret Gön
0009-0004-4074-3391
Türkiye
Erken Görünüm Tarihi
14 Temmuz 2026
Yayımlanma Tarihi
-
Gönderilme Tarihi
31 Mart 2026
Kabul Tarihi
26 Haziran 2026
Yayımlandığı Sayı
Yıl 2026 Sayı: Advanced Online Publication