A curated list of open source and open access Romanian Language NLP Datasets. For the moment we don't add parallel copora to the list.
For additions or any other changes please submit a pull request.
- Unlabeled text Corpora
- Semantic Textual Similarity / Paraphrasing
- Lexical Simplification / Complexity Prediction
- Natural Language Inference
- Summarization
- Dialect and regional speech identification
- Named Entity Recognition (NER)
- Autorship Attribution
- Sentiment Analysis
- Emotion and Mental Health Language
- Dependency Parsing
- Diacritics Restoration / Grammar Correction
- Fake News / Clickbait / Satirical News
- Offensive Language
- LLM Safety and Evaluation
- Questions and Answering
- Spelling, Dictionaries and Gramatical Errors
- Automatic Speech Recognition (ASR)
The FuLG dataset is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully extracted from Common Crawl.
Part of a large multilanguage corpus originated from Common Crawl. It's a raw, unannotated corpus. It has roughly 50 GB of Romanian text in 4.5 million documnets. For details check its homepage and the paper
Similar to Oscar, part of a multilanguage corpus also based on Common Crawl from 2018. Romanian text is 16GB large
Romanian language wikipedia dump.
A collection of varoius unannotated corpora collected around 2018-2019. Includes books, scraped newspapers and juridical documents
A collection of written and spoken text from various sources: Articles, Fairy tales, Fiction, History, Theatre, News
Romanian national legilation from 1881 to 2021. The corpus includes mainly: governmental decisions, ministerial orders, decisions, decrees and laws. Automatically annotated for Named Entities
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19. It is available in over 100+ languages, Romanian being one of them. Tweets need to be rehydrated
A corpus of Romanian tweets related to COVID and vaccination against COVID, created and collected between January 2021 and February 2022. It contains 19319 tweets.
Minutes of the Sittings of the Chamber of Deputies of Romania (2016-2018) Unannotated corpus
contains 500k+ instances of speech from the parliament podium from 1996 to 2018. Sentence splitting and deduplication onm sentence level have been applied as processing steps Unannotated corpus
Romanian presidential discouses (1990-2020) split in 4 files one for each president. Unannotated corpus
Monolingual Romanian corpus, including content from public websites related to culture
Monolingual (ron) corpus, containing 38063991 tokens and 854096 lexical types in the law domain.
Monolingual Romanian corpus, containing 360833 sentences (9064764 words) in the public administration domain.
The New Civil Procedure Code in Romanian (monolingual) comprising 297888 words.
The Romanian updated criminal code: text with law content.
news articles dataset from romanian newssites title, summary and article
multi-language corpus from online available news sources. It contains also 43mil words in Romanian language from Twitter, Blogs and Newspapers
The Romanian novel collection for ELTeC, the European Literary Text Collection Sources: Biblioteca Metropolitana din Bucuresti, Biblioteca Universitara "Mihai Eminescu" din Iasi, Biblioteca Judeteana din Botosani, personal micro-collections uploaded on Zenodo under the following labels: "Hajduks Library"; "RomanianNovel Library"; "CityMysteries Library"; "BibliotecaDHL_Iasi"
Public dataset of 1447 manually annotated Romanian business-oriented emails. The corpus is annotated with 5 token-related labels, as well as 5 sequence-related classes
The corpus consists of texts written by Romanian authors between 19th century and present, representing stories, short-stories, fairy tales and sketches. The current version contains 19 authors, 1263 full texts and 12516 paragraphs of around 200 words each, preserving paragraphs integrity.
A dataset containing 400 Romanian texts written by 10 authors The dataset contains stories, short stories, fairy tales, novels, articles, and sketches written by Ion Creangฤ, Barbu ลtefฤnescu Delavrancea, Mihai Eminescu, Nicolae Filimon, Emil Gรขrleanu, Petre Ispirescu, Mihai Oltean, Emilia Plugaru, Liviu Rebreanu, Ioan Slavici.
891 Cooking Recipes in Romanian Language
A large-scale Romanian pretraining corpus derived from web data through quality and diversity filtering, totalling billions of tokens. Public releases include fineweb2-ro-llm and fineweb2-ro-bert, providing large-scale corpora for training Romanian language models.
A temporally-aware Romanian legal corpus, part of the GRAF package (alongside JuRO and Law-RoG). Designed for legal reasoning and retrieval tasks, covering Romanian legislation with temporal metadata.
Approximately 708 million Romanian sentence-level records, assembled by splitting existing public text datasets. Intended for masked language modeling and embedding pretraining.
Approximately 19.9 million Romanian documents assembled from existing sources including mC4, OSCAR and Wikipedia. Source licenses differ; check the dataset card and selected configuration before reuse, especially for commercial purposes.
Semantic Textual Similarity dataset for the Romanian language RO-STS contains 8,628 sentence pairs with their similarity scores
A paraphrase corpus created from 10 different Romanian language Bible versions. The final dataset contains 904,815 similar records and 218,977 non matching records, totaling 1,123,927
Around ~100k examples of paraphrases. No clear explanation on how the dataset was built
A multi-language paraphrase corpus for 73 languages extracted from the Tatoeba database. It has ~ 2000 romanian phrases totaling 941 paraphrase groups.
Approximately 70,600 Romanian sentence pairs with semantic-similarity scores for training and evaluating sentence embeddings. Combines translated, existing Romanian and manually generated examples.
Resources for Romanian lexical complexity prediction and lexical simplification, including 3,921 human-annotated word-in-context complexity samples and synonym simplification judgments. Download the dataset archive from the repository; the archive password is documented in its README.
We introduce the first Romanian NLI corpus (RoNLI) comprising 58K training sentence pairs, which are obtained via distant supervision, and 6K validation and test sentence pairs, which are manually annotated with the correct labels.
The repository seems to be just an attempt at starting to build the dataset
Around ~72k Full texts and their summary. Source seems to be news websites. No description or explanation available
A large-scale Romanian news dataset with 615k+ articles for summarization, headline generation, and keyword extraction. Includes rich metadata such as summaries, topics, and dialect labels distinguishing Romania vs Moldova text.
varied compilation of speech samples from five distinct regions of Romania, covering both urban and rural environments. Around 2800 records labeled with age, gender and type of dialect
MOROCO: The Moldavian and Romanian Dialectal Corpus The MOROCO data set contains Moldavian and Romanian samples of text collected from the news domain. The samples belong to one of the following six topics: culture, finance, politics, science, sports, tech totaling over 32.000 labeled records
A speech dataset for Romanian dialect identification containing over 93 hours of audio and ~88k samples across Romania and Moldova. Captures intra-language variation rather than multilingual diversity, making it ideal for ASR robustness and dialect modeling.
A manually annotated Romanian legal-domain NER corpus identifying people, organizations, locations, time expressions and references to legal resources.
Romanian Named Entity Corpus (version 2.0): about 12,330 sentences annotated with 15 entity classes, with train, validation and test splits.
A multilingual Wikipedia-derived named entity recognition corpus with Romanian data and person, organization and location labels.
A Romanian medical treebank with 4,239 sentences from cardiology, diabetes and endocrinology, annotated for syntax and biomedical named entities.
The dataset contains 323k tokens of text, covering more than half of the 19th century (i.e., 1817) until the late part of the 20th century (i.e., 1990). The samples belong to one of the following four historical regions of Romania, namely Bessarabia, Moldavia, Transylvania, and Wallachia.
Romanian literary texts by ten authors across several genres; also listed in the unlabeled-text corpora section above.
A Romanian sentiment classification dataset based on product and movie reviews, distributed in the processed form used for Romanian Transformers experiments.
A multilingual sentiment lexicon with Romanian word-level sentiment entries, generated by graph propagation over a knowledge graph.
A Romanian sentiment dataset of 15,000 reviews, with 7,500 positive and 7,500 negative examples labeled from star ratings.
Sentiment Analysis for Romanian Tweets: a labeled Romanian Twitter sentiment dataset with training, validation and test CSV files.
The Romanian Emotions Dataset, a resource for studying and classifying emotion expressed in Romanian text.
About 5,000 Romanian web articles annotated for topic and sentiment. The repository distributes URLs and a collection script rather than the article text.
A Romanian-language movie-review dataset from Cinemagia for sentiment classification.
Romanian aspect-based sentiment analysis dataset containing 9,590 annotated entries. Includes aspect categories and polarity labels for opinion mining from customer reviews.
Romanian-language survey corpus with 205 respondent records, open-ended responses about emotions and social-media experiences, and associated PHQ-9/GAD-7 questionnaire scores. The dataset uses a custom license; consult its terms before reuse.
The CoNLL shared tasks on multilingual dependency parsing using Universal Dependencies treebanks, including Romanian; this link points to the task archive rather than a standalone Romanian dataset.
A multilingual collection of treebanks derived semi-automatically from Universal Dependencies, enriched with deep-syntactic and semantic annotations, including Romanian.
A Romanian corpus distributed through ELRC-SHARE as a resource for linguistic and dependency-parsing research.
Harmonized Multi-LanguagE Dependency Treebank: dependency annotations converted to a common format across languages, including Romanian.
A Romanian lexical-semantic network of word senses and their relationships, distributed with data and a Python API; it is not itself a dependency treebank.
A reference Universal Dependencies treebank for standard Romanian, containing about 9,500 annotated sentences from multiple genres.
Romanian text prepared for training and evaluating models that restore missing diacritics.
A Romanian language-correction resource linked as a downloadable archive; consult the accompanying files for its format and annotation scheme.
A Romanian diacritics-restoration dataset of approximately 340,925 examples, suitable for training and evaluating models that recover missing diacritics.
A Romanian fake-news research dataset containing more than 14,000 news items for studying misinformation and news credibility.
A Romanian science-and-technology news dataset intended for identifying clickbait headlines.
A Romanian satire-detection corpus of 55,608 news articles labeled as regular or satirical, with train, validation and test CSV files.
A Romanian offensive-language detection corpus of annotated comments from a local sports news website.
manually annotated 4,052 comments on a Romanian local news website into one of the following classes: non-offensive, targeted insults, racist, homophobic, and sexist.
4455 organic generated comments from Facebook live broadcasts annotated not binary offensive language detection tasks and for fine-grained offensive language detection
4800 Romanian comments annotated with offensive text spans Offensive span detection
3860 labeled hate speech records
Dataset consists of 5000 tweets, from which 924 were labeled as offensive (18.48 %) and 4076 tweets as non-offensive.
The corpus contains 39 245 tweets, annotated by multiple annotators, following the sexist label set of a recent study.
A Romanian-language benchmark with 953 prompts for evaluating LLM safety, bias, hallucinations, over-refusal, jailbreak robustness, and Romanian-English consistency. The dataset includes potentially harmful evaluation prompts.
A Romanian Wikipedia-derived fill-in-the-blank dataset with 72,541 examples across 45 academic domains, provided with train, validation, and test splits.
This dataset is just the translation of the gsm8k dataset. GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. There is no information on the quality of the translation
RoCode, a competitive programming dataset, consisting of 2,642 problems written in Romanian, 11k solutions in C, C++ and Python and comprehensive testing suites for each problem. The purpose of RoCode is to provide a benchmark for evaluating the code intelligence of language models trained on Romanian / multilingual text as well as a fine-tuning set for pretrained Romanian models.
Romanian IT Dataset (RoITD) resembling SQuAD 1.1. RoITD consists of 9575 Romanian QA pairs formulated by crowd workers. QA pairs are based on 5043 articles from Romanian Wikipedia articles describing IT and household products. Of the total number of questions, 5103 are possible (i.e. the correct answer can be found within the paragraph) and 4472 are not possible (i.e. the given answer is a "plausible answer" and not correct)
A dataset of 3,574 annotated question-answer pairs from journalist-president exchanges, labeled for reply clarity and evasion. Contains training and validation splits.
The dataset comprises 102,646 high-quality QA pairs from real-world clinical records of 1,011 oncology patients (796 patients with breast cancer and 215 patients with lung cancer). The QA pairs are the results of a manual annotation process carried out by physicians specialized in oncology and radiotherapy RoMedQA includes 76,416 QA pairs about breast cancer patients and 26,230 about lung cancer patients, with questions grounded in medical case summaries (epicrises).
Romanian legal MCQA dataset, comprising 10,836 questions from three examinations. Each entry essentially consists of a body in which a theoretical question is posed regarding a legal aspect, along with three possible answer choices labeled A, B, and C, out of which at mosttwo answers are correct.
A domain-specific multiple-choice QA dataset with ~14k biology questions in Romanian, designed for evaluating and fine-tuning LLMs in educational contexts aligned with the national curriculum.
A multimodal Romanian benchmark for driving-license exam reasoning, combining text, images, and legal references. Includes over 1,100 questions across tasks like QA, visual QA, and retrieval.
A longitudinal dataset of Romanian math exams spanning 1895โ2025, with over 10k problems and 600+ exam sets. Supports educational AI, retrieval, and reasoning tasks with rare historical depth.
Romanian question-answer data extracted from "Who Wants to Be a Millionaire?" videos, with original and translated JSON files available for download. Published in 2025; the linked public Zenodo deposit was created in July 2026.
Synthetic dataset with ~1.9M records. Altered and correct statement as columns
Romanian Archaisms Regionalisms Lexicon containing ~ 1940 Word definitions
Romanian Rules for Dialects - 1940 regionalisms, meanings and the region of provenience
The dataset was developed mainly for speech processing applications, yet its applicability extends beyond this domain. RoLEX includes over 330,000 curated entries with information regarding lemma, morphosyntactic description, syllabification, lexical stress and phonemic transcription.
The first Romanian legal-domain grammatical error correction dataset with ~350k annotated sentence pairs. Enables research in correction, normalization, and document processing for Romanian legal text.
A Romanian lexical resource containing 9,836 words grouped into four school-grade vocabulary lists, useful for readability analysis and age-adapted lexical simplification.
Underrepresented Speech Dataset from Open Data. Duration of the dataset is 4h 18m 55s. Distribution according to platforms: 83% of the content comes from YouTube, 12% from SoundCloud, and 5% from Vimeo Dataset covers primarily under-represented speech groups (outside the 19โ29 male category)
A Romanian multi-speaker audio-and-transcript dataset with approximately 247,000 records for speech recognition and speech synthesis research. Research and educational use only; consult the dataset card for restrictions relating to the source audio.