Dear Leiden University-only students, the following thesis topics have now been archived, and are no longer available for thesis projects at Leiden University. See the Open Topics page for available projects.

Archived thesis projects @TDS Lab

  1. [ML] Fixing the most popular mixed data type distance measure

    Many AI, ML and data science methods depend on the notion of a distance, which often acts as a dissimilarity measure between observations in the data set. In real-world data sets, variables have various types, e.g. continuous, ordinal, nominal/categorical and binary, contained within one data set. In such cases, dissimilarity is almost always measured using Gower's distance. It min-max-scales numeric variables, and assigns distances to non-numeric variables as 1 if the values are unequal, and 0 if they are. Dimensions are just added directly, like in the Manhattan distance measure. The implication is that distances are dominated by categorical dimensions, as the distance (if non-zero) corresponds to the largest possible distance in the numeric dimensions, which will typically have smaller values. Also, average distances per dimension are not equalized (not even if the dimensions themselves are normalized or standardized first), and are dominated by imbalanced columns. This project will develop a balanced version of Gower's distance that makes the contribution of every feature on average equal, and leaves the possibility to re-weigh the contribution of features. The resulting distance measure will be used for risk stratification of people with metabolic syndrome on a large scale data warehouse with health, demographic and socio-economic data, but is expected to find wide-spread use in distance-based machine learning tasks on heterogeneous data.

    Daily supervisors: Marcel Haas (LUMC), Marco Spruit
  2. [NLP/LLM] From incident to insight: LLMs for classification of medical incidents in healthcare

    In the LUMC, every colleague can file a critical incident in our incident reporting system. So far, there are over 80,000 incidents reported in the last decade. This system therefore offers a wealth of information and opportunities for improvement in patient care . However, these incidents have highly unstructured categories making them very difficult to use and deduct any useful information. Luckily there are always unstructured textual notes that elaborate on the incident giving detailed information.

    That is why we think LLMs can provide a solution as an effective method to handle unstructured textual data.

    We suspect there is especially much to learn from and gain in transfer-related incidents e.g. when a patient is transferred from one department to another. We need your help in anonymising our data and subsequently classifying all transfer-related incidents. We already have our dataset completely available for use and another student already worked on it. Furthermore, the study is already approved by our local ethical committee so the work is publishable so everything is set perfectly for a good thesis! ??

    Daily supervisors: Marco Spruit, Sesmu Arbous, Vincent Ribbens (LUMC)

    (DUTCH understanding required)
  3. [NLP] External validation of NLP model for lung cancer prediction (DUTCH understanding required)

    In a research collaboration with Amsterdam UMC and Erasmus MC, you will externally validate a successful NLP model to predict early detection of lung cancer in GPs' clinical notes, as developed at Amsterdam UMC, at both/either LUMC and/or ErasmusMC using real-world GP clinical texts. Finetuning procedures based on error analyses and ablation studies to further optimise the original model will likely constitute part of your scientific contribution.

    Daily supervisors: Marco Spruit and others

  4. [LLM] Leveraging Large Language Models to Improve Prediction in Geriatric Care (DUTCH understanding required)

    In prospective studies of older patients, such as the TENT study (cancer patients, n=2000) and the APOP study (Emergency Department patients, n=750), we have performed comprehensive baseline geriatric assessments. These included daily functioning, comorbidities, living situation, cognition, nutrition, and frailty, with follow-up for one year on mortality, quality of life, and functional decline. Using these structured data, we developed prediction models, but their performance was modest, with AUCs typically below 0.75. We hypothesize that this limitation reflects the fact that many aspects of frailty, multimorbidity, and patient context are recorded not in structured variables but in free-text data such as referral letters, discharge summaries, and clinician notes. Recent advances in large language models (LLMs) allow scalable extraction of clinically relevant features from such text, potentially yielding more accurate and clinically useful predictions.

    Aim: To validate whether privacy-friendly large language models applied to unstructured electronic health record (EHR) text can improve prediction of mortality, functional decline, and quality of life in older patients, compared with models based only on structured geriatric assessments.

    Design & Setting: Retrospective analysis of existing prospective cohorts, enriched with EHR text data available via CTCue at Leiden University Medical Center.

    Expected Impact: This project directly tests whether LLMs can unlock hidden predictive value from routinely collected clinical text. If successful, it will provide the first validated evidence that LLM-enhanced models outperform standard geriatric assessments in predicting outcomes across two distinct high-risk populations. This could improve prognostication, support shared decision-making, and help allocate geriatric resources more effectively - contributing to more person-centered, equitable care for older patients.

    Daily supervisors: Simon Mooijaart, Bram van Dijk (LUMC), Marco Spruit
  5. [NLP] Identification of psoriatic arthritis flares in patient records

    Psoriatic arthritis (PsA) is a chronic condition impacting about 20% of psoriasis patients, with unpredictable flares affecting nearly a quarter of patients each year. Our understanding of flare triggers remains limited, making improved identification and analysis crucial for better patient care. In collaboration with Maxima Medisch Centrum and the Medical Informatics and Rheumatology departments at Erasmus MC, this master thesis project aims to identify PsA flares and their risk factors within free-text clinical notes by applying advanced natural language processing (NLP) techniques, including large language models. Using real-world data from the DEPAR (Dutch south west Early Psoriatic ARthritis) prospective cohort, you will extract and prepare relevant lifestyle and flare information from medical records. The project includes hands-on experience with NLP-driven data analysis on real world data and offers on-site supervision and the opportunity to co-author a research paper for publication. Good knowledge of Dutch is preferred due to the need to work with Dutch clinical free text.

    Daily supervisors: Tom Seinen (EMC) and Marco Spruit

  6. [FML] Ontwerp en Ontwikkeling van een Gezamenlijk Regionaal Informatie Platform (GRIP) met Decentrale Architectuur (DUTCH understanding highly prefered)

    De zorg in Nederland staat onder druk door oplopende arbeidstekorten en een toename van zorgvraag bij ouderen en er komen meer inwoners met chronische aandoeningen. Om in een regionale samenwerking in de zorg samen te werken vanuit verschillende databronnen worden Privacy Enhancing Technologieen ingezet zoals Multi Party Computation. In de regionale samenwerking RIGA (www.riga.nl) wordt de data vervolgens samengebracht op geaggregeerd niveau, hier worden rapportages van gemaakt waar regionaal op gestuurd wordt:

    In dit project komen verschillende aspecten aan bod:

    Daily supervisors: Els Roorda (RIGA), Marco Spruit

  7. [NLP] Narrative reporting with LLMs in HealthyChronos

    HealthyChronos is an app that helps young people after cancer treatment to manage their energy levels and regain control over their lives. Based on structured data such as sleep and activity patterns, and unstructured data such as free text inputs concerning the user's personal history and set goals, this application provides a report and advice on how the user is doing in the light of its goals. In this internship, the state-of-the-art regarding retrieval-augmented generation with Large Language Models (LLMs) will be explored, for extraction of information from external databases, and for exploring what narrative formats for presenting the report with LLMs are most appealing to the user, through live testing with a small user panel. One main goal is to have the report as accurate and reliable as possible w.r.t. the information in the database, and narrative formats could be useful in achieving this. There is also room for secondary research interests, such as how LLMs could help produce reports in simple language, other languages, or what else sparks your interest. Part of the internship will be carried out at HealthyChronos in Leiden so that use of source code and data is possible. If you want to make positive impact with your NLP skills in other humans' lives, this is the internship for you!

    NB: This project requires mastery of Dutch.

    Supervisors: Bram van Dijk (LUMC), Max van Duijn (LIACS), Marco Spruit
  8. [NLP] News mining for Homicide characteristics in Indonesia

    In the context of the LUGF project Understanding Homicide in Indonesia: Harnessing Traditional and New Media Data for Insight, a collaboration with ISGA (prof Marieke Liem, dr Olga Bogolyubova, student), you:

    Results may build upon SNPcurator for real-time literature information extraction and this thesis for topic modelling.

    Supervisor: Marco Spruit, Olga Bogolyubova (FGGA)
  9. [NLP] Extracting Adverse Drug Reactions from SmPC Using Large Language Models

    Background
    Previous research has demonstrated the effectiveness of natural language processing techniques in extracting adverse drug reactions (ADRs) from Summary of Product Characteristics (SmPC) documents. However, the potential of large language models (LLMs) for this task remains unexplored.

    Objective
    To develop and evaluate a method using large language models to automatically extract adverse drug reactions from SmPC documents, comparing its performance to previous NLP approaches.

    Methods

    1. Data Collection:
      • Scrape SmPC documents from the Electronic Medicines Compendium (EMC), focusing on section 4.8 (Undesirable effects).
      • Use the same dataset of 647 medicines as in the previous study for comparability.
    2. LLM-based Extraction:
      • Fine-tune a pre-trained LLM (e.g., BERT, RoBERTa, or GPT) on a subset of manually annotated SmPC documents.
      • Develop prompts to guide the LLM in identifying and extracting ADRs and their frequencies.
      • Implement post-processing steps to clean and standardize extracted ADRs.
    3. Evaluation:
      • Use the same subset of 37 commonly prescribed medicines for manual review.
      • Calculate precision, recall, and F1-score to assess performance.
      • Compare results with the previous rule-based NLP approach.
    4. Error Analysis:
      • Analyze false positives and false negatives to identify areas for improvement.

    Expected Outcomes

    Significance
    This study will explore the potential of LLMs in improving the accuracy and efficiency of ADR extraction from SmPC documents, potentially enhancing pharmacovigilance and drug safety monitoring processes.

    Advisors: Ian Shen (ext.), Marco Spruit