Klasifikace autorství textu s neznámým autorem

Dolník, Karel

Text authorship classification with unknown authors

dc.contributor.advisor	Hajič, Jan
dc.creator	Dolník, Karel
dc.date.accessioned	2024-11-29T14:03:04Z
dc.date.available	2024-11-29T14:03:04Z
dc.date.issued	2024
dc.identifier.uri	http://hdl.handle.net/20.500.11956/192068
dc.description.abstract	Přiřazení autorství pomocí statistických a výpočetních metod je hojně zkoumaným tématem literární vědy, ovšem jen málo prací se zabývá řešením problému, kdy klasifikovaný text nenapsal nikdo z autorů, které model viděl při trénování. Tato práce hledá způsob, jak takového neznámého autora detekovat v rámci stejných metod strojového učení, které se pro přiřazení autorství běžně používají, zejména klasifikátoru SVM. Zavádíme zde upravené klasifikační schéma One-versus-Rest-and-None které rozšiřuje schéma One-versus-Rest o trénování pomocí dat, která nepatří žádnému klasifikovanému autorovi. K tomu lze využít synteticky vytvořená data, nebo data od autorů, u kterých je jisté, že s klasifikovanými texty nejsou nijak spojeni. Ukázalo se, že právě při použití syntetických dat dojde k nejmenšímu snížení přesnosti oproti klasifikaci bez detekce neznámého autora.	cs_CZ
dc.description.abstract	Statistical and computational authorship attribution is a widely researched topic in literary science, but few works deal with solving the problem when the classified text does not belong to any of the authors the model saw during training. This work seeks a way to detect such an unknown author using machine learning methods commonly used for authorship attribution, especially the SVM classifier. He we introduce a modified One-versus-Rest-and-None classification scheme, which extends the One-versus-Rest scheme by training with data that does not belong to any classified author. This can be done using synthetically produced data or data from authors who are certain to have no connection to the classified texts. It turned out that the smallest decrease in accuracy occurred when synthetic data is used, compared to the classification without detection of an unknown author.	en_US
dc.language	Čeština	cs_CZ
dc.language.iso	cs_CZ
dc.publisher	Univerzita Karlova, Matematicko-fyzikální fakulta	cs_CZ
dc.subject	authorship classification\|machine learning\|computational literary science\|stylometry	en_US
dc.subject	klasifikace autorství\|strojové učení\|výpočetní literární věda\|stylometrie	cs_CZ
dc.title	Klasifikace autorství textu s neznámým autorem	cs_CZ
dc.type	bakalářská práce	cs_CZ
dcterms.created	2024
dcterms.dateAccepted	2024-06-28
dc.description.department	Institute of Formal and Applied Linguistics	en_US
dc.description.department	Ústav formální a aplikované lingvistiky	cs_CZ
dc.description.faculty	Matematicko-fyzikální fakulta	cs_CZ
dc.description.faculty	Faculty of Mathematics and Physics	en_US
dc.identifier.repId	269513
dc.title.translated	Text authorship classification with unknown authors	en_US
dc.contributor.referee	Mírovský, Jiří
thesis.degree.name	Bc.
thesis.degree.level	bakalářské	cs_CZ
thesis.degree.discipline	Computer Science with specialisation in Artificial Intelligence	en_US
thesis.degree.discipline	Informatika se specializací Umělá inteligence	cs_CZ
thesis.degree.program	Computer Science	en_US
thesis.degree.program	Informatika	cs_CZ
uk.thesis.type	bakalářská práce	cs_CZ
uk.taxonomy.organization-cs	Matematicko-fyzikální fakulta::Ústav formální a aplikované lingvistiky	cs_CZ
uk.taxonomy.organization-en	Faculty of Mathematics and Physics::Institute of Formal and Applied Linguistics	en_US
uk.faculty-name.cs	Matematicko-fyzikální fakulta	cs_CZ
uk.faculty-name.en	Faculty of Mathematics and Physics	en_US
uk.faculty-abbr.cs	MFF	cs_CZ
uk.degree-discipline.cs	Informatika se specializací Umělá inteligence	cs_CZ
uk.degree-discipline.en	Computer Science with specialisation in Artificial Intelligence	en_US
uk.degree-program.cs	Informatika	cs_CZ
uk.degree-program.en	Computer Science	en_US
thesis.grade.cs	Velmi dobře	cs_CZ
thesis.grade.en	Very good	en_US
uk.abstract.cs	Přiřazení autorství pomocí statistických a výpočetních metod je hojně zkoumaným tématem literární vědy, ovšem jen málo prací se zabývá řešením problému, kdy klasifikovaný text nenapsal nikdo z autorů, které model viděl při trénování. Tato práce hledá způsob, jak takového neznámého autora detekovat v rámci stejných metod strojového učení, které se pro přiřazení autorství běžně používají, zejména klasifikátoru SVM. Zavádíme zde upravené klasifikační schéma One-versus-Rest-and-None které rozšiřuje schéma One-versus-Rest o trénování pomocí dat, která nepatří žádnému klasifikovanému autorovi. K tomu lze využít synteticky vytvořená data, nebo data od autorů, u kterých je jisté, že s klasifikovanými texty nejsou nijak spojeni. Ukázalo se, že právě při použití syntetických dat dojde k nejmenšímu snížení přesnosti oproti klasifikaci bez detekce neznámého autora.	cs_CZ
uk.abstract.en	Statistical and computational authorship attribution is a widely researched topic in literary science, but few works deal with solving the problem when the classified text does not belong to any of the authors the model saw during training. This work seeks a way to detect such an unknown author using machine learning methods commonly used for authorship attribution, especially the SVM classifier. He we introduce a modified One-versus-Rest-and-None classification scheme, which extends the One-versus-Rest scheme by training with data that does not belong to any classified author. This can be done using synthetically produced data or data from authors who are certain to have no connection to the classified texts. It turned out that the smallest decrease in accuracy occurred when synthetic data is used, compared to the classification without detection of an unknown author.	en_US
uk.file-availability	V
uk.grantor	Univerzita Karlova, Matematicko-fyzikální fakulta, Ústav formální a aplikované lingvistiky	cs_CZ
thesis.grade.code	2
uk.publication-place	Praha	cs_CZ
uk.thesis.defenceStatus	O