-
CorAl Dataset: Aljamiado Qur’an. Digital Corpus of Multiple Versions
CorAl (Aljamiado Qur'an) provides access to the complete corpus of digital editions of Aljamiado translations of the Qur'an. It enables users to compare the different versions dynamically, with the texts aligned to facilitate collation and analysis....
Instituto: Instituto de Lenguas y Culturas del Mediterráneo y Oriente Próximo (ILC), CSIC
-
SimpMedLexSp (Simple Medical Lexicon for Spanish)
A medical lexicon of 14013 pairs of technical word forms and the corresponding simplified synonym or definition. It is aimed at automatic text simplification in Spanish. A subset of the lexicon (4642 term entries) was also normalized to Unified Medical...
Instituto: Instituto de Lengua, Literatura y Antropología (ILLA), CSIC
-
Corpus for Complex Word Identification in Medical Spanish Texts (CWI-Med-Sp)
[Description of methods used for collection/generation of data] The corpus statistics and methods are explained in the following article: Federico Ortega-Riba, Leonardo Campillos-Llanos, Doaa Samy (2025) "Lexical Simplification in Spanish Texts For...
Instituto: Instituto de Lengua, Literatura y Antropología (ILLA), CSIC
-
CT-EBM-SP - Corpus of Clinical Trials for Evidence-Based-Medicine in Spanish
A collection of 1200 texts (292 173 tokens) about clinical trials studies and clinical trials announcements in Spanish: - 500 abstracts from journals published under a Creative Commons license, e.g. available in PubMed or the Scientific Electronic...
Instituto: Instituto de Lengua, Literatura y Antropología (ILLA), CSIC
-
CT-EBM-SP - Corpus of Clinical Trials for Evidence-Based-Medicine in Spanish ...
A collection of 1200 texts (292173 tokens) about clinical trials studies and clinical trials announcements in Spanish: - 500 abstracts from journals published under a Creative Commons license, e.g. available in PubMed or the Scientific Electronic...
Instituto: Instituto de Lengua, Literatura y Antropología (ILLA), CSIC
-
CLARA-MeD corpus
A collection of 24.298 pairs of professional and simplified texts (>96 million tokens): 1) Drug leaflets and summaries of product characteristics (10 211 pairs of texts, >82M words); 2) Cancer-related information summaries (201 pairs of texts,...
Instituto: Instituto de Lengua, Literatura y Antropología (ILLA), CSIC
