Skip to content

Latest commit

ย 

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

Awesome

A list of Romanian NLP Datasets

A curated list of open source and open access Romanian Language NLP Datasets. For the moment we don't add parallel copora to the list.

For additions or any other changes please submit a pull request.

Table of contents

Unlabeled text Corpora

The FuLG dataset is a comprehensive Romanian language corpus comprising
150 billion tokens, carefully extracted from Common Crawl. 

arXiv

Part of a large multilanguage corpus originated from Common Crawl.
It's a raw, unannotated corpus. It has roughly 50 GB of Romanian text
in 4.5 million documnets. For details check its homepage 
and the paper

arXiv Homepage

 Similar to Oscar, part of a multilanguage corpus also based on Common Crawl
 from 2018. Romanian text is 16GB large

arXiv Homepage

  Romanian language wikipedia dump. 
  A collection of varoius unannotated corpora collected around 2018-2019.
  Includes books, scraped newspapers and juridical documents  
  A collection of written and spoken text from various
  sources: Articles, Fairy tales, Fiction, History, Theatre, News
 Romanian national legilation from  1881 to 2021. The corpus
 includes mainly: governmental decisions, ministerial orders,
 decisions, decrees and laws.
 Automatically annotated for Named Entities

ACL Homepage

Mega-COV is a billion-scale dataset from Twitter for studying COVID-19. It is available in over 100+ languages, Romanian being one of them. Tweets need to be rehydrated

arXiv Medium

A corpus of Romanian tweets related to COVID and vaccination against COVID, created and collected between January 2021 and February 2022. It contains 19319 tweets.

Minutes of the Sittings of the Chamber of Deputies of Romania (2016-2018)
Unannotated corpus
contains 500k+ instances of speech from the parliament podium from
1996 to 2018. Sentence splitting and deduplication onm sentence level
have been applied as processing steps
Unannotated corpus
Romanian presidential discouses (1990-2020) split in 4 files
one for each president. Unannotated corpus
Monolingual Romanian corpus, including content from public websites related to culture

Monolingual (ron) corpus, containing 38063991 tokens and 854096 lexical types in the law domain.

Monolingual Romanian corpus, containing 360833 sentences (9064764 words) in the public administration domain.

The New Civil Procedure Code in Romanian (monolingual) comprising 297888 words.

The Romanian updated criminal code: text with law content.

news articles dataset from romanian newssites title, summary and article

multi-language corpus from online available news sources. It contains also 43mil words in Romanian language from Twitter, Blogs and Newspapers

Homepage

The Romanian novel collection for ELTeC, the European Literary Text Collection Sources: Biblioteca Metropolitana din Bucuresti, Biblioteca Universitara "Mihai Eminescu" din Iasi, Biblioteca Judeteana din Botosani, personal micro-collections uploaded on Zenodo under the following labels: "Hajduks Library"; "RomanianNovel Library"; "CityMysteries Library"; "BibliotecaDHL_Iasi"

Public dataset of 1447 manually annotated Romanian business-oriented emails. The corpus is annotated with 5 token-related labels, as well as 5 sequence-related classes

MDPI

The corpus consists of texts written by Romanian authors between 19th century and present, representing stories, short-stories, fairy tales and sketches. The current version contains 19 authors, 1263 full texts and 12516 paragraphs of around 200 words each, preserving paragraphs integrity.

A dataset containing 400 Romanian texts written by 10 authors The dataset contains stories, short stories, fairy tales, novels, articles, and sketches written by Ion Creangฤƒ, Barbu ลžtefฤƒnescu Delavrancea, Mihai Eminescu, Nicolae Filimon, Emil Gรขrleanu, Petre Ispirescu, Mihai Oltean, Emilia Plugaru, Liviu Rebreanu, Ioan Slavici.

MDPI

891 Cooking Recipes in Romanian Language

A large-scale Romanian pretraining corpus derived from web data through quality and diversity filtering, totalling billions of tokens. Public releases include fineweb2-ro-llm and fineweb2-ro-bert, providing large-scale corpora for training Romanian language models.

ACL HuggingFace

A temporally-aware Romanian legal corpus, part of the GRAF package (alongside JuRO and Law-RoG). Designed for legal reasoning and retrieval tasks, covering Romanian legislation with temporal metadata.

ACL

Approximately 708 million Romanian sentence-level records, assembled by splitting existing public text datasets. Intended for masked language modeling and embedding pretraining.

Approximately 19.9 million Romanian documents assembled from existing sources including mC4, OSCAR and Wikipedia. Source licenses differ; check the dataset card and selected configuration before reuse, especially for commercial purposes.

Semantic Textual Similarity / Paraphrasing

Semantic Textual Similarity dataset for the Romanian language RO-STS contains 8,628 sentence pairs with their similarity scores

NeurIPS

A paraphrase corpus created from 10 different Romanian language Bible versions. The final dataset contains 904,815 similar records and 218,977 non matching records, totaling 1,123,927

Around ~100k examples of paraphrases. No clear explanation on how the dataset was built

A multi-language paraphrase corpus for 73 languages extracted from the Tatoeba database. It has ~ 2000 romanian phrases totaling 941 paraphrase groups.

ACL Homepage

Approximately 70,600 Romanian sentence pairs with semantic-similarity scores for training and evaluating sentence embeddings. Combines translated, existing Romanian and manually generated examples.

Lexical Simplification / Complexity Prediction

Resources for Romanian lexical complexity prediction and lexical simplification, including 3,921 human-annotated word-in-context complexity samples and synonym simplification judgments. Download the dataset archive from the repository; the archive password is documented in its README.

ACL

Natural Language Inference

We introduce the first Romanian NLI corpus (RoNLI) comprising 58K training sentence pairs, which are obtained via distant supervision, and 6K validation and test sentence pairs, which are manually annotated with the correct labels. ACL

The repository seems to be just an attempt at starting to build the dataset

Summarization

Around ~72k Full texts and their summary. Source seems to be news websites. No description or explanation available

A large-scale Romanian news dataset with 615k+ articles for summarization, headline generation, and keyword extraction. Includes rich metadata such as summaries, topics, and dialect labels distinguishing Romania vs Moldova text.

ACL HuggingFace

Dialect and regional speech identification

varied compilation of speech samples from five distinct regions of Romania, covering both urban and rural environments. Around 2800 records labeled with age, gender and type of dialect

arXiv

MOROCO: The Moldavian and Romanian Dialectal Corpus The MOROCO data set contains Moldavian and Romanian samples of text collected from the news domain. The samples belong to one of the following six topics: culture, finance, politics, science, sports, tech totaling over 32.000 labeled records

arXiv

A speech dataset for Romanian dialect identification containing over 93 hours of audio and ~88k samples across Romania and Moldova. Captures intra-language variation rather than multilingual diversity, making it ideal for ASR robustness and dialect modeling.

ACL

Named Entity Recognition (NER)

A manually annotated Romanian legal-domain NER corpus identifying people, organizations, locations, time expressions and references to legal resources.

Romanian Named Entity Corpus (version 2.0): about 12,330 sentences annotated with 15 entity classes, with train, validation and test splits.

A multilingual Wikipedia-derived named entity recognition corpus with Romanian data and person, organization and location labels.

A Romanian medical treebank with 4,239 sentences from cardiology, diabetes and endocrinology, annotated for syntax and biomedical named entities.

The dataset contains 323k tokens of text, covering more than half of the 19th century (i.e., 1817) until the late part of the 20th century (i.e., 1990). The samples belong to one of the following four historical regions of Romania, namely Bessarabia, Moldavia, Transylvania, and Wallachia.

arXiv

Autorship Attribution

Romanian literary texts by ten authors across several genres; also listed in the unlabeled-text corpora section above.

Sentiment Analysis

A Romanian sentiment classification dataset based on product and movie reviews, distributed in the processed form used for Romanian Transformers experiments.

A multilingual sentiment lexicon with Romanian word-level sentiment entries, generated by graph propagation over a knowledge graph.

A Romanian sentiment dataset of 15,000 reviews, with 7,500 positive and 7,500 negative examples labeled from star ratings.

Sentiment Analysis for Romanian Tweets: a labeled Romanian Twitter sentiment dataset with training, validation and test CSV files.

The Romanian Emotions Dataset, a resource for studying and classifying emotion expressed in Romanian text.

About 5,000 Romanian web articles annotated for topic and sentiment. The repository distributes URLs and a collection script rather than the article text.

A Romanian-language movie-review dataset from Cinemagia for sentiment classification.

Romanian aspect-based sentiment analysis dataset containing 9,590 annotated entries. Includes aspect categories and polarity labels for opinion mining from customer reviews.

Emotion and Mental Health Language

Romanian-language survey corpus with 205 respondent records, open-ended responses about emotions and social-media experiences, and associated PHQ-9/GAD-7 questionnaire scores. The dataset uses a custom license; consult its terms before reuse.

arXiv

Dependency Parsing

The CoNLL shared tasks on multilingual dependency parsing using Universal Dependencies treebanks, including Romanian; this link points to the task archive rather than a standalone Romanian dataset.

A multilingual collection of treebanks derived semi-automatically from Universal Dependencies, enriched with deep-syntactic and semantic annotations, including Romanian.

A Romanian corpus distributed through ELRC-SHARE as a resource for linguistic and dependency-parsing research.

Harmonized Multi-LanguagE Dependency Treebank: dependency annotations converted to a common format across languages, including Romanian.

A Romanian lexical-semantic network of word senses and their relationships, distributed with data and a Python API; it is not itself a dependency treebank.

A reference Universal Dependencies treebank for standard Romanian, containing about 9,500 annotated sentences from multiple genres.

Diacritics Restoration / Grammar Correction

Romanian text prepared for training and evaluating models that restore missing diacritics.

A Romanian language-correction resource linked as a downloadable archive; consult the accompanying files for its format and annotation scheme.

A Romanian diacritics-restoration dataset of approximately 340,925 examples, suitable for training and evaluating models that recover missing diacritics.

Fake News / Clickbait / Satirical News

A Romanian fake-news research dataset containing more than 14,000 news items for studying misinformation and news credibility.

A Romanian science-and-technology news dataset intended for identifying clickbait headlines.

A Romanian satire-detection corpus of 55,608 news articles labeled as regular or satirical, with train, validation and test CSV files.

Offensive Language

A Romanian offensive-language detection corpus of annotated comments from a local sports news website.

manually annotated 4,052 comments on a Romanian local news website into one of the following classes: non-offensive, targeted insults, racist, homophobic, and sexist.

arXiv

4455 organic generated comments from Facebook live broadcasts annotated not binary offensive language detection tasks and for fine-grained offensive language detection

IEEE

4800 Romanian comments annotated with offensive text spans Offensive span detection

MDPI

3860 labeled hate speech records

Dataset consists of 5000 tweets, from which 924 were labeled as offensive (18.48 %) and 4076 tweets as non-offensive.

ACL

The corpus contains 39 245 tweets, annotated by multiple annotators, following the sexist label set of a recent study.

ACL

LLM Safety and Evaluation

A Romanian-language benchmark with 953 prompts for evaluating LLM safety, bias, hallucinations, over-refusal, jailbreak robustness, and Romanian-English consistency. The dataset includes potentially harmful evaluation prompts.

Questions and Answers

A Romanian Wikipedia-derived fill-in-the-blank dataset with 72,541 examples across 45 academic domains, provided with train, validation, and test splits.

This dataset is just the translation of the gsm8k dataset. GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. There is no information on the quality of the translation

RoCode, a competitive programming dataset, consisting of 2,642 problems written in Romanian, 11k solutions in C, C++ and Python and comprehensive testing suites for each problem. The purpose of RoCode is to provide a benchmark for evaluating the code intelligence of language models trained on Romanian / multilingual text as well as a fine-tuning set for pretrained Romanian models.

arXiv

Romanian IT Dataset (RoITD) resembling SQuAD 1.1. RoITD consists of 9575 Romanian QA pairs formulated by crowd workers. QA pairs are based on 5043 articles from Romanian Wikipedia articles describing IT and household products. Of the total number of questions, 5103 are possible (i.e. the correct answer can be found within the paragraph) and 4472 are not possible (i.e. the given answer is a "plausible answer" and not correct)

A dataset of 3,574 annotated question-answer pairs from journalist-president exchanges, labeled for reply clarity and evasion. Contains training and validation splits.

The dataset comprises 102,646 high-quality QA pairs from real-world clinical records of 1,011 oncology patients (796 patients with breast cancer and 215 patients with lung cancer). The QA pairs are the results of a manual annotation process carried out by physicians specialized in oncology and radiotherapy RoMedQA includes 76,416 QA pairs about breast cancer patients and 26,230 about lung cancer patients, with questions grounded in medical case summaries (epicrises).

arXiv

Romanian legal MCQA dataset, comprising 10,836 questions from three examinations. Each entry essentially consists of a body in which a theoretical question is posed regarding a legal aspect, along with three possible answer choices labeled A, B, and C, out of which at mosttwo answers are correct.

ACL

A domain-specific multiple-choice QA dataset with ~14k biology questions in Romanian, designed for evaluating and fine-tuning LLMs in educational contexts aligned with the national curriculum.

ACL

A multimodal Romanian benchmark for driving-license exam reasoning, combining text, images, and legal references. Includes over 1,100 questions across tasks like QA, visual QA, and retrieval.

ACL HuggingFace

A longitudinal dataset of Romanian math exams spanning 1895โ€“2025, with over 10k problems and 600+ exam sets. Supports educational AI, retrieval, and reasoning tasks with rare historical depth.

arXiv HuggingFace

Romanian question-answer data extracted from "Who Wants to Be a Millionaire?" videos, with original and translated JSON files available for download. Published in 2025; the linked public Zenodo deposit was created in July 2026.

ACL

Spelling, Dictionaries and Gramatical Errors

Synthetic dataset with ~1.9M records. Altered and correct statement as columns

Romanian Archaisms Regionalisms Lexicon containing ~ 1940 Word definitions

Romanian Rules for Dialects - 1940 regionalisms, meanings and the region of provenience

The dataset was developed mainly for speech processing applications, yet its applicability extends beyond this domain. RoLEX includes over 330,000 curated entries with information regarding lemma, morphosyntactic description, syllabification, lexical stress and phonemic transcription.

Cambridge

The first Romanian legal-domain grammatical error correction dataset with ~350k annotated sentence pairs. Enables research in correction, normalization, and document processing for Romanian legal text.

arXiv

A Romanian lexical resource containing 9,836 words grouped into four school-grade vocabulary lists, useful for readability analysis and age-adapted lexical simplification.

Automatic Speech Recognition

Underrepresented Speech Dataset from Open Data. Duration of the dataset is 4h 18m 55s. Distribution according to platforms: 83% of the content comes from YouTube, 12% from SoundCloud, and 5% from Vimeo Dataset covers primarily under-represented speech groups (outside the 19โ€“29 male category)

MDPI

A Romanian multi-speaker audio-and-transcript dataset with approximately 247,000 records for speech recognition and speech synthesis research. Research and educational use only; consult the dataset card for restrictions relating to the source audio.

About

A list of Romanian NLP Datasets

Topics

Resources

Contributing

Stars

62 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors