Diagnostic Interview Corpus - Translation

The Diagnostic Interview Corpus is a multilingual dataset of 12,754 French medical consultation sentences (questions and instructions) with translations into 12 languages and associated UMLS-based semantic glosses. It supports research on low-resource medical machine translation, semantic representation, and pictograph generation. Languages - Source: French - Targets (in translations.csv): Albanian, Modern Standard Arabic, Tunisian Arabic, Moroccan Arabic, Algerian Arabic, Dari (Afghan Persian), Farsi (Iranian Persian), Russian, English, Spanish, Tigrinya, Ukrainian - Semantic gloss (in translations.csv): French sentences aligned with UMLS glosses (concept sequences + functional tokens). - Paraphrases (in paraphrases.csv): French paraphrases aligned with the corresponding French source sentences, generated through a grammar-based approach to ensure controlled syntactic variation Domains and registers - Medical consultations - Questions and instructions (e.g., symptom checks, treatment directives) - Categories by body region (e.g., head, chest, abdomen) Features - Parallel multilingual translations created and adapted with clinical experts - Semantic gloss layer (UMLS CUIs + functional tokens) for pictograph generation - Patient-centered simplifications and cultural adaptations to improve comprehension Example French: Avez-vous des nausées ou des vomissements ? English: Do you have nausea or vomiting? UMLS gloss: You | Nausea | or – article | Vomiting | Question Intended Use - Low-resource multilingual MT research - Semantic representation learning (UMLS-based) - Pictograph translation systems for patients with limited health literacy - Evaluation of medical-domain MT beyond surface-level accuracy Acknowledgements This corpus was developed in the context of the BabelDr and PictoDr projects at the University of Geneva in collaboration with Geneva University Hospitals.This work is part of the PROPICTO project, funded by the Swiss National Science Foundation (N°197864) and the French National Research Agency (ANR-20-CE93-0005). This project also received funding by the ”Fondation Privée des Hôpitaux Universitaires de Genève”.

    Organizational unit
    Propicto Project
    Type
    Dataset
    DOI
    License
    Creative Commons Attribution 4.0 International
    Keywords
    low resource machine translation, medical dialogues, medical domain, medical questionnaires, semantic gloss, UMLS, Standard Modern Arabic, Albanian, Moroccan Arabic, Tunisian Arabic, Algerian Arabic, Dari, Farsi, Russian, English, Spanish, Tigrinya, Ukrainian, French
Publication date25/09/2025
Retention date23/09/2035
accessLevelPublicAccess levelPublic
SensitivityBlue
licenseContract on the use of data
License
37
8
  • Quality (0 Reviews)
  • Usefulness (0 Reviews)

Datacite metadata

    
      <?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<datacite:resource xmlns:premis="http://www.loc.gov/premis/v3" xmlns:mets="http://www.loc.gov/METS/" xmlns:datacite="http://datacite.org/schema/kernel-4" xmlns:dlcm="http://www.dlcm.ch/dlcm/v3" xmlns:fits="http://hul.harvard.edu/ois/xml/ns/fits/fits_output" xmlns:xlink="http://www.w3.org/1999/xlink">
    <datacite:identifier identifierType="DOI">10.26037/yareta:47hsvsq6hngg7gc4eashhqtwhi</datacite:identifier>
    <datacite:creators>
        <datacite:creator>
            <datacite:creatorName nameType="Personal">Bouillon, Pierrette</datacite:creatorName>
            <datacite:givenName>Pierrette</datacite:givenName>
            <datacite:familyName>Bouillon</datacite:familyName>
            <datacite:nameIdentifier xsi:type="datacite:nameIdentifier" nameIdentifierScheme="ORCID" schemeURI="https://orcid.org/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">https://orcid.org/0000-0002-8854-6360</datacite:nameIdentifier>
        </datacite:creator>
        <datacite:creator>
            <datacite:creatorName nameType="Personal">Gerlach, Johanna</datacite:creatorName>
            <datacite:givenName>Johanna</datacite:givenName>
            <datacite:familyName>Gerlach</datacite:familyName>
            <datacite:nameIdentifier xsi:type="datacite:nameIdentifier" nameIdentifierScheme="ORCID" schemeURI="https://orcid.org/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">https://orcid.org/0000-0002-3371-4021</datacite:nameIdentifier>
        </datacite:creator>
        <datacite:creator>
            <datacite:creatorName nameType="Personal">Mutal, Jonathan David</datacite:creatorName>
            <datacite:givenName>Jonathan David</datacite:givenName>
            <datacite:familyName>Mutal</datacite:familyName>
            <datacite:nameIdentifier xsi:type="datacite:nameIdentifier" nameIdentifierScheme="ORCID" schemeURI="https://orcid.org/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">https://orcid.org/0000-0001-7729-8456</datacite:nameIdentifier>
        </datacite:creator>
        <datacite:creator>
            <datacite:creatorName nameType="Personal">Spechbach, Hervé</datacite:creatorName>
            <datacite:givenName>Hervé</datacite:givenName>
            <datacite:familyName>Spechbach</datacite:familyName>
        </datacite:creator>
    </datacite:creators>
    <datacite:titles>
        <datacite:title>Diagnostic Interview Corpus - Translation</datacite:title>
    </datacite:titles>
    <datacite:publisher publisherIdentifier="https://ror.org/01swzsf04" publisherIdentifierScheme="ROR" schemeURI="https://ror.org/">Université de Genève, Yareta</datacite:publisher>
    <datacite:publicationYear>2025</datacite:publicationYear>
    <datacite:resourceType resourceTypeGeneral="Dataset">Dataset</datacite:resourceType>
    <datacite:subjects>
        <datacite:subject subjectScheme="keywords" xml:lang="en">low resource machine translation;medical dialogues;medical domain;medical questionnaires;semantic gloss;UMLS;Standard Modern Arabic;Albanian;Moroccan Arabic;Tunisian Arabic;Algerian Arabic;Dari;Farsi;Russian;English;Spanish;Tigrinya;Ukrainian;French</datacite:subject>
    </datacite:subjects>
    <datacite:contributors>
        <datacite:contributor contributorType="ResearchGroup">
            <datacite:contributorName nameType="Organizational">[3ac8bd75-1155-4a0c-acec-35c441c5e3ef] Propicto Project</datacite:contributorName>
            <datacite:affiliation xsi:type="datacite:affiliation" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">[UNIGE-FTI] Université de Genève - Faculté de Traduction et d'Interprétation</datacite:affiliation>
            <datacite:affiliation xsi:type="datacite:affiliation" affiliationIdentifier="https://ror.org/01swzsf04" affiliationIdentifierScheme="ROR" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">[UNIGE] Université de Genève</datacite:affiliation>
        </datacite:contributor>
        <datacite:contributor contributorType="ProjectManager">
            <datacite:contributorName nameType="Personal">Bouillon, Pierrette</datacite:contributorName>
        </datacite:contributor>
        <datacite:contributor contributorType="ProjectManager">
            <datacite:contributorName nameType="Personal">Ormaechea-Grijalba, Lucia</datacite:contributorName>
        </datacite:contributor>
        <datacite:contributor contributorType="ProjectManager">
            <datacite:contributorName nameType="Personal">Mutal, Jonathan David</datacite:contributorName>
        </datacite:contributor>
        <datacite:contributor contributorType="ProjectManager">
            <datacite:contributorName nameType="Personal">Gerlach, Johanna</datacite:contributorName>
        </datacite:contributor>
        <datacite:contributor contributorType="ProjectManager">
            <datacite:contributorName nameType="Personal">de Graaf, Céline</datacite:contributorName>
        </datacite:contributor>
    </datacite:contributors>
    <datacite:dates>
        <datacite:date dateType="Issued">2025-09-25T00:00:00Z</datacite:date>
        <datacite:date dateType="Collected">2022-01-01T00:00:00Z/2025-01-01T00:00:00Z</datacite:date>
        <datacite:date dateType="Created">2025-09-25T10:15:42.449996Z</datacite:date>
        <datacite:date dateType="Updated">2025-10-16T14:16:24.615341Z</datacite:date>
        <datacite:date dateType="Accepted">2025-10-16T14:16:24.28811305Z</datacite:date>
    </datacite:dates>
    <datacite:alternateIdentifiers>
        <datacite:alternateIdentifier alternateIdentifierType="ARK">ark:64642/rt947hsvsq6hngg7gc4eashhqtwhi</datacite:alternateIdentifier>
    </datacite:alternateIdentifiers>
    <datacite:relatedIdentifiers/>
    <datacite:formats>
        <datacite:format>text/csv</datacite:format>
        <datacite:format>text/markdown</datacite:format>
        <datacite:format>text/xml</datacite:format>
    </datacite:formats>
    <datacite:rightsList>
        <datacite:rights rightsURI="https://www.dlcm.ch/access-level/final">PUBLIC</datacite:rights>
        <datacite:rights rightsURI="https://www.dlcm.ch/data-tag">BLUE</datacite:rights>
        <datacite:rights rightsURI="https://www.dlcm.ch/data-use-policy">LICENSE</datacite:rights>
        <datacite:rights rightsURI="https://creativecommons.org/licenses/by/4.0/" rightsIdentifier="CC-BY-4.0" rightsIdentifierScheme="SPDX">Creative Commons Attribution 4.0 International</datacite:rights>
    </datacite:rightsList>
    <datacite:descriptions>
        <datacite:description descriptionType="Abstract" xml:lang="en">The Diagnostic Interview Corpus is a multilingual dataset of 12,754 French medical consultation sentences (questions and instructions) with translations into 12 languages and associated UMLS-based semantic glosses. It supports research on low-resource medical machine translation, semantic representation, and pictograph generation.

Languages
- Source: French
- Targets (in translations.csv): Albanian, Modern Standard Arabic, Tunisian Arabic, Moroccan Arabic, Algerian Arabic, Dari (Afghan Persian), Farsi (Iranian Persian), Russian, English, Spanish, Tigrinya, Ukrainian
- Semantic gloss  (in translations.csv): French sentences aligned with UMLS glosses (concept sequences + functional tokens).
- Paraphrases (in paraphrases.csv): French paraphrases aligned with the corresponding French source sentences, generated through a grammar-based approach to ensure controlled syntactic variation

Domains and registers
- Medical consultations
- Questions and instructions (e.g., symptom checks, treatment directives)
- Categories by body region (e.g., head, chest, abdomen)

Features
- Parallel multilingual translations created and adapted with clinical experts
- Semantic gloss layer (UMLS CUIs + functional tokens) for pictograph generation
- Patient-centered simplifications and cultural adaptations to improve comprehension

Example
French: Avez-vous des nausées ou des vomissements ?
English: Do you have nausea or vomiting?
UMLS gloss: You | Nausea | or – article | Vomiting | Question

Intended Use
- Low-resource multilingual MT research
- Semantic representation learning (UMLS-based)
- Pictograph translation systems for patients with limited health literacy
- Evaluation of medical-domain MT beyond surface-level accuracy

Acknowledgements
This corpus was developed in the context of the BabelDr and PictoDr projects at the University of Geneva in collaboration with Geneva University Hospitals.This work is part of the PROPICTO project, funded by the Swiss National Science Foundation (N°197864) and the French National Research Agency (ANR-20-CE93-0005). This project also received funding by the ”Fondation Privée des Hôpitaux Universitaires de Genève”.</datacite:description>
    </datacite:descriptions>
    <datacite:fundingReferences/>
</datacite:resource>

    

Packages information

    Metadata version
    4.0
    Disposal approval
    Yes
    Retention expiration
    23/09/2035 00:00:00
    Created
    16/10/2025 14:16:24
    Deposit
    Diagnostic Interview Corpus - Translation
    Deposit creation
    Deposit approval
    16/10/2025 14:16:24
    SIP
    Diagnostic Interview Corpus - Translation
    AIP
    Diagnostic Interview Corpus - Translation
    Agent 1
    Linux | amd64 | 4.18.0-553.47.1.el8_10.x86_64
    Agent 2
    Eclipse Adoptium | Software | 21.0.8
    Agent 3
    DLCM | Software | 3.1.0
    Agent 4
    FITS | Software | 1.6.0
    Agent 5
    ClamAV | Software | 1.4.2

Similar archives

All rights reserved by DLCM and the University of GenevaunigeBlack