ADAPTIVE KNOWLEDGE REGULARIZATION FOR CONTINUAL LEARNING IN TRANSFORMER ARCHITECTURES
Abstract
The subject matter of the article is development of latent representation regularization mechanism in transformer-based architecture under conditions of continuous learning with domain shifts. Modern language models achieve high quality in static learning scenarios, but they remain limited in long-term operation cases, incremental adaptation to new domains and lacking resistance to catastrophic forgetting, especially in the absence of access to previously observed data. This paper explores the possibility of overcoming these limitations by combining uncertainty inspired regularization and forgetting attention mechanisms into one transformer architecture. The goal of the study is to design, implement and validate transformer-based architecture with multi-level representation regularization mechanism that can help transformer-based language models efficiently adapt to alternative data distribution while retaining previously acquired knowledge. The proposed approach aims to achieve an equilibrium between model adaptability and stability in continual learning without requiring complete model retraining or legacy data retention. The tasks to be solved in this study include: formalization of an unified conceptual latent representations regularization method that combines Bayesian uncertainty inspired latent representations regularization with head-wise attention scaling in attention mechanism; implementation of this method into the transformer model; creating the experimental case of continuous language modeling with sequential domain shifts; give quantitative estimation of model forgetting and stability in prediction quality; compare the proposed model to classical naive fine-tuning, LoRA and parameter regularization methods. The conclusions demonstrate that the proposed method achieves lower forgetting, lower perplexity on previously learned domains and a better stability–plasticity trade-off than naive fine-tuning, LoRA and Elastic Weight Consolidation, while requiring comparable computational resources. The scientific novelty of proposed approach consists in development of layer-selective latent regularization framework for continual language modeling which integrates an attention with forgetting mechanism with preserving domain-invariant representations through statistical alignment in latent space for reducing forgetting in continual learning scenarios. Unlike existing approaches to continual learning that consider model regularization either on parameter or memory level (by using previous data), the proposed approach moves regularization into representation space, where it uses both direct regularization via proposed uncertainty-based regularization and indirect regularization via attention with forget gate, ensuring the models’ possibility of stable continual language modeling in non-stationary environments.
Keywords
References
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J. & Zhang, Y. An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-Tuning. IEEE Transactions on Audio, Speech and Language Processing. 2025, vol. 33, pp.3776–3786. DOI: 10.1109/taslpro.2025.3606231.
Jiang, M., Fan, J. & Li, F. Advances in continual learning: A comprehensive review. Expert Systems with Applications, 2025, vol. 294, article no.128739. DOI: 10.1016/j.eswa.2025.128739.
Shin, H., Lee, J.K., Kim, J. & Kim, J. (2017). Continual Learning with Deep Generative Replay. Neural Information Processing Systems. 2017, vol. 30, pp 2990-2999. DOI: 10.48550/arXiv.1705.08690.
He, J., Guo, H., Zhu, K., Zhao, Z., Tang, M. & Wang, J. SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp.3254–3266. DOI: 10.18653/v1/2024.emnlp-main.190.
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A.A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D. & Hadsell, R. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 2017, vol. 114, no. 13, pp. 3521–3526, DOI: 10.1073/pnas.1611835114.
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., Laroussilhe, Q.D., Gesmundo, A., Attariyan, M. & Gelly, S. Parameter-Efficient Transfer Learning for NLP, Proceedings of the 36th International Conference on Machine Learning, 2019, vol. 97, pp. 2790-2799, DOI: 10.48550/arXiv.1902.00751.
Hu, E.J., Shen, Y., Wallis, P., Zeyuan Allen-Zhu, Li, Y., Wang, S., Wang, L. & Chen, W. LoRA: Low-Rank Adaptation of Large Language Models, Available at: https://openreview.net/forum?id=nZeVKeeFYf9 (accessed 18 March 2026).
Li, H., Tan, Z., Li, X. & Huang, W. ATLAS: Adapter-Based Multi-Modal Continual Learning with a Two-Stage Learning Strategy. Avaliable at: https://arxiv.org/abs/2410.10923 (accessed 18 March 2026).
Hu, G., Zhang, W., Ding, H. & Zhu, W. Gradient Episodic Memory with a Soft Constraint for Continual Learning. Avaliable at: https://arxiv.org/abs/2011.07801 (accessed 18 March 2026).
Buzzega, P., Boschini, M., Porrello, A., Abati, D., & Calderara, S. Dark Experience for General Continual Learning: a Strong, Simple Baseline, Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, article no. 1335. DOI: 10.48550/arXiv.2004.07211.
Nguyen, C.V., Li, Y., Bui, T.D. & Turner, R.E. Variational Continual Learning. Available at: https://openreview.net/forum?id=BkQqq0gRb (accessed 18 March 2026).
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Gomez, A., Kaiser, Ł., Poloskuhin, I. Attention Is All You Need, Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017, vol. 30, pp 6000-6010. DOI: 10.48550/arXiv.1706.03762.
Bonnet, D., Kellian Cottart, Tifenn Hirtzlin, Tarcisius Januel, Dalgaty, T., Vianello, E. & Querlioz, D. Bayesian continual learning and forgetting in neural networks. Nature Communications, vol. 16, pp.9614–9614. DOI: 10.1038/s41467-025-64601-w.
Cover, T.M. & Thomas, J.A. Elements of Information Theory. Hoboken, NJ, USA: John Wiley & Sons, Inc., 2005. 748 p. DOI: 10.1002/047174882x.
Kingma, D. P., Welling, M. An Introduction to Variational Autoencoders, Foundations and Trends in Machine Learning, 2019, vol. 12, no. 4, pp. 307–392, DOI: 10.1561/2200000056.
Lin, Z., Nikishin, E., He, X. O., Courville, A. Forgetting Transformer: Softmax Attention with a Forget Gate. Available at: https://openreview.net/forum?id=q2Lnyegkr8.
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L. & Liu, T. On Layer Normalization in the Transformer Architecture, Proceedings of the 37th International Conference on Machine Learning, 2020, vol. 119, pp. 10524–10533. DOI: 10.48550/arXiv.2002.04745.
Wikimedia Foundation. Wikipedia dataset. In: Hugging Face Datasets. Available at: https://huggingface.co/datasets/wikimedia/wikipedia (accessed 18 March 2026).
Vladimir Blagojevic. CC-News dataset. Hugging Face, n.d. Available at: https://huggingface.co/datasets/vblagoje/cc_news (accessed 18 March 2026).
Pietro Lesci. EURLEX-57K dataset. Hugging Face, n.d. Available at: https://huggingface.co/datasets/pietrolesci/eurlex-57k (accessed 18 March 2026).
Hospedales, T. M., Antoniou, A., Micaelli, P., &. Storkey, A. J Meta-Learning in Neural Networks: A Survey, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 9, pp. 5149–5169, 2022, DOI: 10.1109/TPAMI.2021.3079209
DOI: https://doi.org/10.32620/reks.2026.2.07
Refbacks
- There are currently no refbacks.
