Orthogonalising gradients to speed up neural network optimisation

Autor:	Tuddenham, Mark, Prügel-Bennett, Adam, Hare, Jonathan
Rok vydání:	2022
Předmět:	Computer Science - Machine Learning
Druh dokumentu:	Working Paper
Popis:	The optimisation of neural networks can be sped up by orthogonalising the gradients before the optimisation step, ensuring the diversification of the learned representations. We orthogonalise the gradients of the layer's components/filters with respect to each other to separate out the intermediate representations. Our method of orthogonalisation allows the weights to be used more flexibly, in contrast to restricting the weights to an orthogonalised sub-space. We tested this method on ImageNet and CIFAR-10 resulting in a large decrease in learning time, and also obtain a speed-up on the semi-supervised learning BarlowTwins. We obtain similar accuracy to SGD without fine-tuning and better accuracy for na\"ively chosen hyper-parameters.
Databáze:	arXiv
Externí odkaz:	http://arxiv.org/abs/2202.07052 Zobrazit plný text záznamu View this record from Arxiv