How Do BERT embeddings organize linguistic knowledge?

Puccetti, Giovanni; Miaschi, Alessio; Dell'Orletta, Felice

doi:10.18653/v1/2021.deelio-1.6

Several studies investigated the linguistic information implicitly encoded in Neural Language Models. Most of these works focused on quantifying the amount and type of information available within their internal representations and across their layers. In line with this scenario, we proposed a different study, based on Lasso regression, aimed at understanding how the information encoded by BERT sentence-level representations is arranged within its hidden units. Using a suite of several probing tasks, we showed the existence of a relationship between the implicit knowledge learned by the model and the number of individual units involved in the encodings of this competence. Moreover, we found that it is possible to identify groups of hidden units more relevant for specific linguistic properties. © 2021 Association for Computational Linguistics.

How Do BERT embeddings organize linguistic knowledge?

Puccetti, Giovanni;Miaschi, Alessio;Dell'Orletta, Felice

2021

Abstract

Several studies investigated the linguistic information implicitly encoded in Neural Language Models. Most of these works focused on quantifying the amount and type of information available within their internal representations and across their layers. In line with this scenario, we proposed a different study, based on Lasso regression, aimed at understanding how the information encoded by BERT sentence-level representations is arranged within its hidden units. Using a suite of several probing tasks, we showed the existence of a relationship between the implicit knowledge learned by the model and the number of individual units involved in the encodings of this competence. Moreover, we found that it is possible to identify groups of hidden units more relevant for specific linguistic properties. © 2021 Association for Computational Linguistics.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno di pubblicazione
	
				2021
			
	Settore Scientifico Disciplinare (validi fino a 24/06/2024)
	
				Settore INF/01 - Informatica
Settore ING-INF/05 - Sistemi di Elaborazione delle Informazioni
			
	Titolo del Convegno
	
				Deep Learning Inside Out: 2nd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, DeeLIO 2021
			
	Periodo del Convegno
	
				2021
			
	Titolo del Volume
	
				Proceedings of Deep Learning Inside Out (DeeLIO): The 2nd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures
			
	Editore
	
				Association for Computational Linguistics
			
	ISBN
	
				978-195408530-5
			
	DOI
	
				https://dx.doi.org/10.18653/v1/2021.deelio-1.6
			
	Parole chiave
	
				Computational linguistics, Embeddings
			
	Appare nelle tipologie:
	
				4.1 Contributo in Atti di convegno

File in questo prodotto:

File	Dimensione	Formato
How Do BERT Embeddings Organize Linguistic Knowledge - 2021.deelio-1.6.pdf accesso aperto Tipologia: Published version Licenza: Creative Commons Dimensione 2.54 MB Formato Adobe PDF	2.54 MB	Adobe PDF