This proposal aglomerates three subprojects that represent the next steps in the responsible use and access to big data, specifically focusing on taking the usage of big data for AI applications a step further. Subproject 1 aims at building upon the existing work on digital twins by exploring a method that enables the coupling of digital twins while avoiding the extremely intensive computations that would be required when the models were coupled in a ‘naïve’ way. Probabilistic programming is used to achieve this goal. Subproject 2 extends the capabilities of NLP (Natural Language Processing) for EHR (Electronic Health Records). The subproject leverages the abundant information hidden in the free texts of EHR to improve three medical prediction research areas: cohort identification, predictor extraction and outcome extraction. Subproject 3 takes the FAIR principles (Findable, Accessible, Interoperable, Reusable) for data one step further and applies them to AI models. Tools will be created to support developers publishing their FAIR models. The approach will be demonstrated using two medical prediction models.