Student project catalogue

Master in AI for Drug Discovery

Available projects for students interested in molecular simulation, cheminformatics, computational toxicology, biomedical AI, spectroscopy, regulatory networks and molecular optimization.

12available projects
2–3typical duration in months
AIapplied to chemistry, biology and simulation
Notebooktechnical report and final presentation

Selection rules

Students must contact the organizers before starting the project selection process.

What students must send

  1. Contact one of the organizers by email: piotto@unisa.it or lucsessa@unisa.it.
  2. Indicate the first-choice project and the second-choice project.
  3. List the two selected projects on separate lines, one above the other.
  4. Indicate the two months during which you are available to work on the project.
  5. Wait for confirmation before considering the project assigned.
First choice Project number and project title
Second choice Project number and project title
Availability Two available months

Available projects

Projects are listed one below the other. Each card contains a short description and a download link for the full project file.

01MD + Symbolic AI

PolySTRIPES: Symbolic AI Representation of Molecular Dynamics Trajectories

Prototype a workflow that converts molecular dynamics trajectories of polymers, membranes or membrane-related systems into symbolic sequences of local interactions, compact descriptors and AI-ready graph or sequence representations.

MD trajectoriesTokenizationEntropy descriptorsGraphs
Download description
Duration: 2–3 months
Biological difficulty: ★★★☆☆
Computational difficulty: ★★★☆☆
02Biomarkers + Safety

Cardiac Monitoring and AI

Design an AI-oriented workflow integrating ECG-derived features and cardiac imaging phenotypes to support the interpretation of drug action, cardiotoxicity and translational safety signals.

ECGCardiac imagingBiomarkersMultimodal AI
Download description
Duration: 2–3 months
Biological difficulty: ★★★★☆
Computational difficulty: ★★★☆☆
03Adaptive Simulation

Agent-Driven Adaptive Multiscale Simulation

Implement a simplified 2D adaptive multiscale simulation in which coarse molecular agents can keep, merge or split their representation according to local order, uncertainty and computational need.

Agent systemsUncertaintyMerge/Split rulesMultiscale modelling
Download description
Duration: 2–3 months
Biological difficulty: ★★☆☆☆
Computational difficulty: ★★★★★
04DFT + Bayesian AI

Bayesian Correction of DFT-Based NMR Predictions

Develop a Bayesian calibration layer that corrects DFT-based NMR chemical-shift predictions using reference molecules, while quantifying the uncertainty of the corrected spectra.

DFTNMRBayesian modellingUncertainty
Download description
Duration: 2–3 months
Biological difficulty: ★★★☆☆
Computational difficulty: ★★★★☆
05Omics + Networks

GENOA: Categorical Clustering and Regulatory Network Inference

Build a proof-of-concept pipeline that discretizes tumor gene expression into categorical states, clusters patients in this discrete space and reconstructs cluster-specific regulatory networks.

TranscriptomicsCategorical clusteringNetwork inferencePrecision oncology
Download description
Duration: 2–3 months
Biological difficulty: ★★★☆☆
Computational difficulty: ★★★★★
06Molecular Optimization

MolParetoBO: Surrogate-Assisted Multi-Objective Molecular Optimization

Implement a simplified multi-objective molecular optimization workflow combining molecular mutations, validity filters, surrogate models, active learning and Pareto-based candidate selection.

RDKitSurrogate modelsPareto frontADMET
Download description
Duration: 2–3 months
Biological difficulty: ★★★☆☆
Computational difficulty: ★★★★☆
07Ligandability Mapping

ProbeMap-AI: Mixed-Solvent Molecular Dynamics and AI-Ready FragMaps

Prototype a simplified mixed-solvent molecular dynamics workflow using molecular probes to identify ligandable regions, cryptic pockets and chemically interpretable hotspots on protein surfaces.

OpenMMXenon probesFragMapsHotspot ranking
Download description
Duration: 2–3 months
Biological difficulty: ★★★☆☆
Computational difficulty: ★★☆☆☆
08QSTR + Drug Safety

OTO-QSTR: OECD-Compliant Quantitative Structure–Toxicity Relationship Modelling for Ototoxicity

Develop an interpretable QSTR workflow to predict drug-induced ototoxicity from molecular structure, including dataset curation, descriptor and fingerprint calculation, machine-learning models, applicability-domain analysis and OECD-compliant validation.

QSAR/QSTROtotoxicityRDKitSHAPOECD validation
Download description
Duration: 2–3 months
Biological difficulty: ★★★☆☆
Computational difficulty: ★★★☆☆
09Active Learning + Experimental Design

ActiveLearn-ED: Active Learning for Experiment Selection and Model Improvement in Drug Discovery

Prototype a Python active learning workflow that recommends the next experimental measurements in a drug discovery campaign, first by selecting molecules expected to improve a predictive model and then by prioritizing the most informative molecule-assay pairs in sparse experimental matrices.

Active learningExperimental designUncertaintyAssay selectionApplicability domain
Download description
Duration: 2–3 months
Biological difficulty: ★★☆☆☆
Computational difficulty: ★★★★☆
10 PPI + AI Lead Design

PEARL: From Dimers to Drugs

Prototype an end-to-end computational pipeline that starts from a kinase dimer structure, identifies protein-protein interface hotspots, extracts a seed peptide and applies molecular dynamics, protein language models and generative docking to propose peptide, peptidomimetic and small-molecule inhibitor candidates.

PPI hotspots Molecular dynamics ESM-2 / ProteinMPNN Pharmacophore design DiffDock / REINVENT
Download description
Duration: 6 - 8 weeks
Biological difficulty: ★★★★☆
Computational difficulty: ★★★★★
11 AMPs + ML Profiling

AMPACT: Machine Learning for Antimicrobial Peptide Activity Prediction and Profiling on YADAMP

Build a complete machine learning workflow on the YADAMP antimicrobial peptide dataset, moving from descriptor cleaning and exploratory analysis to feature engineering, MIC regression, activity classification, multi-target prediction, clustering and anomaly detection to profile peptide activity across microbial species.

YADAMP Antimicrobial peptides MIC prediction Regression / classification Clustering / anomaly detection
Download description
Duration: 2–3 months
Biological difficulty: ★★★☆☆
Computational difficulty: ★★★☆☆
12 Statistics + Inference

DAISY: Data Analysis, Inference and Statistical studY

Build a complete and reproducible statistical data analysis workflow on a real-world dataset, moving from data examination, cleaning and descriptive statistics to distribution analysis, correlation analysis, parameter estimation, confidence intervals, hypothesis testing and decision analysis to draw statistically sound and interpretable conclusions.

Statistical inference Descriptive statistics Confidence intervals Hypothesis testing Decision analysis Python / SciPy
Download description
Duration: 1–2 months
Biological difficulty: ★☆☆☆☆
Computational difficulty: ★★★☆☆


Confidentiality and intellectual property notice


The project descriptions, ideas, methodologies, documents, datasets, and related materials presented in this catalogue are confidential and intended exclusively for internal educational use within the Master in AI for Drug Discovery. All intellectual property rights, know-how, concepts, and original materials remain the property of their respective authors, tutors, institutions, or rights holders. Any unauthorized copying, reproduction, disclosure, distribution, publication, adaptation, external use, commercial exploitation, or development of derivative projects based on these materials is strictly prohibited without prior written authorization.