Метод аугментації текстів про стан масивів вод на основі інтелектуальної прив`язки до багатозв`язних геоінформаційних систем іменованих сутностей
Вантажиться...
Файли
Дата
Назва журналу
Номер ISSN
Назва тому
Анотація
The article is dedicated to the augmentation of Ukrainian-language texts about the state of surface water bodies in a river basin for the training of machine learning models that should automatically annotate these texts, i. e. referencing in space
and time and performing their classification.
The authors describe the progress made in creating the "Water Information System with Spatial and Temporal Referencing for the Southern Bug Basin" ("WISEST-SBB"), which is being populated with annotated data on the state of water bodies
in the river basin using technologies and algorithms developed by the authors earlier. It is noted that the experience has
shown a lack of information for training machine learning models intended for automating its annotation. An analysis of
modern methods of text data augmentation applicable to Ukrainian texts has been conducted, highlighting their drawbacks,
primarily the high probability of synthesizing unreliable information.
The proposed approach suggests augmenting data on water bodies of a river network, considering the propagation of
reliable information about one water body to others located upstream or downstream or otherwise connected to them. To
formalize and automate this process, a new formalization of the river network in the form of a multi-related geoinformation
system of named entities (MGISNE) is proposed, which involves identifying named entities among all objects and then establishing spatial relationships between them. Examples of MGISNE are described, including hydrographic or ecological
networks, networks of administrative entities, and others. The previously proposed recursive algorithm for referencing water
body data with named entities in MGISNE is improved, and its formalized description is developed. After referencing texts
with water bodies, the augmentation of the texts is proposed with subsequent verification of the results in a semi-automated
manner, which can later be made more automated.
The results of the proposed method, algorithm, and approaches in the WISEST-SBB system are characterized, demonstrating their effectiveness. The findings of this work can be extended to other types of MGISNE, both for basins of other
rivers and systems of a different character.
Опис
УДК
Тип документа
Мова
ISSN
Бібліографічний опис
Метод аугментації текстів про стан масивів вод на основі інтелектуальної прив`язки до багатозв`язних геоінформаційних систем іменованих сутностей [Текст] / В. Б. Мокін, К. О. Бондалєтов, Є. М. Крижановський, В. О. Караваєв // Вісник Вінницького політехнічного інcтитуту. – 2023. – Вип. 3. – С. 55–65.
Схвалення
Рецензія
Доповнено
Цитується в
Список використаної літератури (6)
- Directive 2000/60/ec of the European Parliament and of the Council. EUR-Lex – Access to European Union Law. [Electronic resource]. Available: https://eur-lex.europa.eu/resource.html?uri=cellar:5c835afb-2ec6-4577-bdf8- 756d3d694eeb.0004.02/DOC_1&format=PDF . Access: 07.06.2023
- Oleh Bisikalo, and Alexander Yahimovich, Keyword search based on lexical relationships in the text, Mauritius: Lap Lambert Academic Publishing, 2019, 57 p. ISBN 978-620-0-00314-0
- A. Fiori, Trends and Applications of Text Summarization Techniques. IGI Global, 2019.
- Vitalii Mokin, “NLP for UA : BERT CLS & 10 Classifiers,” Kaggle: Your Machine Learning and Data Science Community. [Electronic resource]. Available: https://www.kaggle.com/code/vbmokin/nlp-for-ua-bert-cls-10-classifiers. Access: 07.06.2023.
- “Environmental indicators: typology and overview,” European Environment Agency. [Electronic resource]. Available: https://www.eea.europa.eu/publications/TEC25
- Vitalii Mokin, and Kostiantyn Bondaletov, “SpaCy for Ukrainian text similarity,” Kaggle: Your Machine Learning and Data Science Community. [Electro