Data Scientist · Clinical Research · Bergamo, Italy
Rasoul Samei
I work on intensive-care data at Istituto Mario Negri, mostly on making fragmented clinical databases usable enough to answer a question. I'm looking for a doctoral position or a research role in health data next.
About
I'm a data scientist in clinical research. Since November 2024 I've been at the Istituto di Ricerche Farmacologiche Mario Negri in Ranica. Most of my work is intensive-care data. I harmonise fragmented clinical databases, build the analytical tables other people's studies run on, and do the statistics that go with them.
I built a dashboard with an embedded language-model assistant so researchers can query study data in plain language, and a no-code extraction workbook a physician used to analyse ICU data herself for her residency thesis. She had no coding background and did not need me in the loop.
Much of the rest is judging what the data can actually support, which I learned working next to clinicians and statisticians.
Before Mario Negri I spent six months in the UK at The Openwork Partnership, on anomaly detection for regulated financial advisory data, and before that I was a teaching assistant in machine learning at the University of Bergamo. I'm Persian and live in Bergamo.
- 3Peer-reviewed publications
- 900+Lines of BigQuery SQL in one production pipeline
- 28+Business KPIs defined and validated
- #1Stat-Hackathon, Bergamo 2022
Research
I work on severity scores and model calibration. But first someone has to make multi-centre clinical databases comparable enough that a question can be asked of them at all, and that is where most of my time goes. So far that has meant GiViTI and MIMIC data.
-
2025
Development and Validation of the Sequential Organ Failure Assessment (SOFA)-2 Score
JAMA, 334(23): 2090–2103
Contributing author. Extracted and processed data from the MargheritaTre (GiViTI) clinical database, statistical analysis, manuscript review.
-
2025
Anomaly Detection in Financial Advisory Services: A Machine Learning Approach for Mortgage Conduct of Business Advisers
ICSEM 2025
Unsupervised anomaly detection on regulated mortgage advisory conduct data, built in Azure Machine Learning for compliance review. Developed from my master's thesis.
-
2025
Assessing the Efficacy of a Sequential General Variational Mode Decomposition-Based Combination Model for US Wind Power Forecasting
Modeling Earth Systems and Environment
Decomposition-based hybrid time series model for wind power forecasting.
Work
-
A single validated table from a fragmented ICU database
Roughly 900 lines of SQL on Google BigQuery that turn a highly fragmented multi-table intensive-care database into one validated analytical table.
Now used as a standard data asset by other teams at the institute and by external researchers. Automating the recurring processing and reporting around it cut manual handling across several projects.
BigQuery · SQL · Python · R
-
SOFA-2 score, published in JAMA
Data extraction and processing from the MargheritaTre (GiViTI) database, statistical analysis, and manuscript review for the updated Sequential Organ Failure Assessment score.
Contributing author.
R · SQL · Severity scores · Model calibration
-
Tools so researchers don't have to ask me
A dashboard with an embedded language-model assistant that lets researchers query study data in plain language, and a no-code extraction workbook for ICU data.
A physician used the workbook to analyse ICU data herself for her residency thesis, without a data scientist in the loop.
Python · LangChain · SQL
-
Anomaly detection for regulated financial advice
An unsupervised anomaly detection system in Azure Machine Learning that surfaces irregular patterns in mortgage advisory conduct data for compliance review. Alongside it, 28+ business KPIs defined, computed, validated, and delivered through Power BI dashboards to senior supervisors.
Method published at ICSEM 2025. Also built Azure Python SDK and Spark SQL pipelines and converted legacy COBOL applications into web-based services.
Azure ML · Python · Spark SQL · Power BI
-
Independent projects
HL7 FHIR R4 records mapped into OMOP-style staging tables on Databricks. An NLP pipeline in spaCy feeding XGBoost classifiers for outcome prediction. A retrieval-augmented question answering assistant over a private document corpus, built with LangChain and pgvector, that keeps an audit log so every answer traces back to its source.
Databricks · spaCy · XGBoost · LangChain · pgvector · FastAPI · Docker · MLflow
Path
- 2024 –Data Scientist, Istituto di Ricerche Farmacologiche Mario Negri IRCCS, Ranica
- 2023 – 2024Data Scientist (six months), The Openwork Partnership, Swindon, UK
- 2021 – 2024MA Economics and Data Analysis, Università degli Studi di Bergamo. Teaching assistant in Machine Learning.
- 2022First place, Stat-Hackathon, Bergamo (Python, R, SAS)
- 2015BA Business Management, Islamic Azad University of Neyshabur
Toolkit
- Languages
- Python, R, SQL / Spark SQL
- Statistics
- Hypothesis testing, regression modelling, time series forecasting, multivariate analysis
- Machine learning
- Supervised and unsupervised methods, anomaly detection, clustering, feature engineering, NLP
- Cloud & data
- Google BigQuery, Azure ML, Azure Synapse, Databricks, Power BI
- Shipping models
- MLflow, FastAPI, Docker, Git, LangChain, Claude Code
- Clinical data
- GiViTI / MargheritaTre, MIMIC, HL7 FHIR R4, OMOP
- Spoken
- Persian (native), English (professional, IELTS 7.5), Italian and German (basic)
Get in touch.
Open to research and industry roles across Europe, and to doctoral positions in health data. Salaried research posts rather than stipend-only studentships, if I get the choice.
rasoul@rasoulsamei.com