Module Specifications
Academic Year 2026 - 2027
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Description This module introduces the fundamental techniques in data analysis with a focus on its application to financial and actuarial problems using R. Topics include the bias/variance trade-off and model complexity, cross-validation techniques to evaluate models and estimate hyper-parameters, and the use of regularisation to mitigate overfitting in highly parameterised models. In addition, students will gain hands-on experience in applying supervised learning techniques for regression and classification tasks using R, evaluating binary classifiers with metrics such as precision, recall, F1 score, ROC curves, and confusion matrices. Unsupervised learning methods, including principal component analysis (PCA) and K-means clustering, will also be covered to reduce data dimensionality, identify latent substructures, and detect anomalies. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Learning Outcomes 1. Explain the bias/variance trade-off and its relationship with model complexity. 2. Implement cross-validation techniques in R to evaluate models on unseen data and estimate hyper-parameters. 3. Apply regularisation methods (e.g., LASSO, ridge regression) to reduce overfitting in highly parameterised models. 4. Utilize R software to implement supervised learning techniques to solve regression and classification problems. 5. Evaluate the performance of binary classifiers using metrics such as precision, recall, F1 score, ROC curves, and confusion matrices. 6. Apply unsupervised learning techniques (e.g., PCA, K-means clustering) to reduce data dimensionality, identify latent substructures, and detect anomalies. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
All module information is indicative and subject to change. For further information,students are advised to refer to the University's Marks and Standards and Programme Specific Regulations at: http://www.dcu.ie/registry/examinations/index.shtml |
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Indicative Content and Learning Activities
Introduction to ML and Actuarial Data Science Overview of ML vs. traditional actuarial models (GLMs). Relevance of ML to insurance pricing, reserving, and risk management; Data Preprocessing and Feature Engineering in R Data cleaning, transformation, and exploratory data analysis. Assets’ Returns and their distribution. Tests for normality. Feature selection methods. Supervised Learning Techniques Linear and logistic regression fundamentals and limitations. Decision trees, random forests, and gradient boosting. Model tuning, cross-validation, and performance evaluation. Unsupervised Learning and Dimensionality Reduction Principal Component Analysis (PCA). Clustering methods (e.g., k-means, hierarchical clustering). Application examples in risk segmentation. Interpretability and Model Diagnostics Techniques to interpret “black-box” models (e.g., variable importance, partial dependence, SHAP). Model validation, stress testing, and sensitivity analysis. Practical Applications in Actuarial Science Pricing, reserving, and forecasting. Discussion of case studies and research. Ethical and regulatory implications in ML deployment performing model diagnostics. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Indicative Reading List Books: None Articles: None | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Other Resources None | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||