http://repositorio.unb.br/handle/10482/42729| Arquivo | Descrição | Tamanho | Formato | |
|---|---|---|---|---|
| 2021_JoséReinaldodaCunhaSantosArosoVieiradaSilvaNeto.pdf | 896,44 kB | Adobe PDF | Visualizar/Abrir |
| Título: | Deep Active Learning Approaches to the task of Named Entity Recognition |
| Autor(es): | Silva Neto, José Reinaldo da Cunha Santos Aroso Vieira da |
| Orientador(es): | Faleiros, Thiago de Paulo |
| Assunto: | Aprendizagem ativa Auto-aprendizagem Classificação sequencial Redes neurais profundas Reconhecimento de entidades nomeadas |
| Data de publicação: | 11-Jan-2022 |
| Data de defesa: | 4-Nov-2021 |
| Referência: | SILVA NETO, José Reinaldo da Cunha S. A. V. da. Deep Active Learning Approaches to the task of Named Entity Recognition. 2021. 83 f., il. Dissertação (Mestrado em Informática)—Universidade de Brasília, Brasília, 2021. |
| Abstract: | Deep neural networks are the current state-of-the-art for a variety of challenging tasks in fields such as natural language processing and computer vision, but they rely on big labeled datasets to be trained to achieve such results. Deep active learning algorithms have been designed to reduce the amount of labeled data to train these models. This dissertation identifies shortcomings of the current works from the literature on deep active learning algorithms applied to the task of named entity recognition, and proposes potential solutions to them. In particular, current works from the literature rely on validation sets to apply early stopping of the model training during the active learning process. In low resource scenarios, however, separating labeled samples in order to create a validation set is undesirable. Therefore, we propose the Dynamic Update of Training Epochs (DUTE) strategy that acts as an unsupervised early stopping technique. Experimental results suggest that the proposed DUTE strategy is capable of maintaining the trained model’s performance, when compared to traditional early stopping techniques, while not relying on validation sets. We also investigate self-labeling as a viable option to further reduce the annotation costs in active learning scenarios. In particular, we experiment with sentence-level and token-level self-labeling strategies. It was observed that despite significant efforts, sentence-level self-labeling did not incur a significant improvement over previous works from the literature. However, token-level self-labeling has shown promising results by training models that achieve similar performance to the current state-of-the-art works on deep active learning from the literature while requiring significantly less hand annotated data. More specifically, experiments performed on the CoNLL2003 dataset have shown that the proposed token-level self-labeling strategy trained a neural model to near peak performance using 29.24% less hand annotated data. |
| Unidade Acadêmica: | Instituto de Ciências Exatas (IE) Departamento de Ciência da Computação (IE CIC) |
| Informações adicionais: | Dissertação (mestrado)—Universidade de Brasília, Instituto de Ciências Exatas, Departamento de Ciência da Computação, 2021. |
| Programa de pós-graduação: | Programa de Pós-Graduação em Informática |
| Licença: | A concessão da licença deste item refere-se ao termo de autorização impresso assinado pelo autor com as seguintes condições: Na qualidade de titular dos direitos de autor da publicação, autorizo a Universidade de Brasília e o IBICT a disponibilizar por meio dos sites www.bce.unb.br, www.ibict.br, http://hercules.vtls.com/cgi-bin/ndltd/chameleon?lng=pt&skin=ndltd sem ressarcimento dos direitos autorais, de acordo com a Lei nº 9610/98, o texto integral da obra disponibilizada, conforme permissões assinaladas, para fins de leitura, impressão e/ou download, a título de divulgação da produção científica brasileira, a partir desta data. |
| Aparece nas coleções: | Teses, dissertações e produtos pós-doutorado |
Os itens no repositório estão protegidos por copyright, com todos os direitos reservados, salvo quando é indicado o contrário.