Introducing RONEC -- the Romanian Named Entity Corpus
We present RONEC - the Named Entity Corpus for the Romanian language. The corpus contains over 26000 entities in ~5000 annotated sentences, belonging to 16 distinct classes. The sentences have been extracted from a copy-right free newspaper, covering several styles. This corpus represents the first...
Gespeichert in:
Hauptverfasser: | , |
---|---|
Format: | Artikel |
Sprache: | eng |
Schlagworte: | |
Online-Zugang: | Volltext bestellen |
Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Zusammenfassung: | We present RONEC - the Named Entity Corpus for the Romanian language. The
corpus contains over 26000 entities in ~5000 annotated sentences, belonging to
16 distinct classes. The sentences have been extracted from a copy-right free
newspaper, covering several styles. This corpus represents the first initiative
in the Romanian language space specifically targeted for named entity
recognition. It is available in BRAT and CoNLL-U Plus formats, and it is free
to use and extend at github.com/dumitrescustefan/ronec . |
---|---|
DOI: | 10.48550/arxiv.1909.01247 |