AI-stödd analys av namnlikhet: En jämförelse av metoder för identifiering av likartade företagsnamn
2026 (Swedish)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE credits
Student thesis
Abstract [sv]
Företagsinformation kan enkelt manipuleras, förvanskas eller utnyttjas för bedrägerier. Företagsnamn som liknar varandra kan inte bara uppfattas förvirrande, namnen kan också vara en säkerhetsrisk. Syftet med arbetet är att undersöka hur en automatiserad metod kan jämföra företagsnamn på ett effektivt och tillförlitligt sätt. För att undersöka detta utvecklas en hybridmodell med kombinerad metod av token och embeddings. Dessa utför en kvantitativ jämförelse där hybridmodellen testas och jämförs mot Bolagsverkets befintliga modell och metoderna token och embeddings var för sig. Modellen tränas och testas med korsvalidering för att få ett så tillförlitligt resultat som möjligt. Därefter utvärderas dessa modeller utifrån mätvärdena precision, recall och F1 score. Resultatet av dessa tester visar att den nuvarande modellen har ett högt precision och ett lågt recall vilket tyder på en obalanserad modell. Hybridmodellens resultat är däremot mer balanserade och uppnår det högsta sammanlagda värdet, F1-score, för samtliga testade modeller. Med detta som grund konstateras att den nya hybridmodellen med kombinationen av tokens-metod och embeddings-metod ger en högre och mer balanserad träffsäkerhet jämfört med Bolagsverkets nuvarande modell. Detta då en kombinerad metod av dessa kan finna likheter som är visuella och semantiska. Dessutom är hybridmodellen mer balanserad än modellerna som består av separat token-metod och embeddings metod.
Abstract [en]
Company information can be easily manipulated, distorted or used for fraud. Company names that are similar to each other can not only be seen as confusing, the names can also be a security risk. The purpose of the work is to investigate how an automated method can compare company names efficiently and reliably. To investigate this, a hybrid model with a combined method of tokens and embeddings is developed. These performed a quantitative comparison where the hybrid model is tested and compared against Bolagsverkets existing model and also methods of tokens and embeddings separately. The model is trained and tested with cross-validation to obtain the most reliable result possible. These models are then evaluated based on the metrics precision, recall and F1-score. The results of these tests show that Bolagsverkets current model have a high precision and a low recall, which indicates an unbalanced model. The results of the hybrid model, on the other hand, are more balanced and achieve the highest F1 score, for all tested models. Based on this, it is concluded that a hybrid model with the combination of the tokens method and the embeddings method provides a higher and more balanced accuracy than previous models. This may be because the combined method of these can find similarities in company names, that are both visual and semantic.
Place, publisher, year, edition, pages
2026. , p. 57
Keywords [en]
AI, company name, name similarity, hybrid model, embeddings, token, machine learning
Keywords [sv]
AI, företagsnamn, namnlikhet, hybridmodell, embeddings, token, maskininlärning
National Category
Software Engineering
Identifiers
URN: urn:nbn:se:miun:diva-57936Local ID: DT-V26-G3-025OAI: oai:DiVA.org:miun-57936DiVA, id: diva2:2080379
Subject / course
Computer Engineering DT1
Educational program
Computer Science TDATG 180 higher education credits
Supervisors
Examiners
2026-06-262026-06-262026-06-26Bibliographically approved