Qualitative and quantitative data from contexts of use for the analysis of si...
En este registro se incluyen dos documentos, un Excel (menciones_contextos_UT_THC_covid-19.xlsx) con el número de menciones y porcentajes de aparición de los términos objeto de estudio en el corpus y subcorpus de Ciencias Sociales, Ciencias y...
Organización: Centro de Ciencias Humanas y Sociales (CCHS), CSIC
CLARA-MeD simplified sentences
This dataset contains 1200 manually simplified sentences (144 019 tokens) from clinical trials in Spanish. A total of 1040 announcements from the European Clinical Trials Register (EudraCT) were analyzed to select sentences with ambiguities or...
Organización: Instituto de Lengua, Literatura y Antropología (ILLA), CSIC
SimpMedLexSp (Simple Medical Lexicon for Spanish)
A medical lexicon of 14013 pairs of technical word forms and the corresponding simpli-fied synonym or definition. It is aimed at automatic text simplification in Spanish. A subset of the lexicon (4642 term entries) was also normalized to Unified...
Organización: Instituto de Lengua, Literatura y Antropología (ILLA), CSIC