2026
BA-FedSHAP: A reproducible toolkit for auditing background-induced attribution drift
SSRN
Preprint on how the choice of background distribution alters federated attribution results, and what that does to an audit.
PublicationsAuthor
Research area
Claims about an AI system should be checkable by someone who was not involved in building it.
What evidence about an AI system is adequate for a particular decision, held by a particular role, in a particular deployment?
Most assurance work asks whether a model is good. That is the wrong unit. A model is not deployed, a system is, into an institution, for a decision, under a regulation, with someone accountable for the outcome. The same model can be adequately evidenced for one of those situations and badly evidenced for the next.
My work here builds methods and software that make that distinction operational: analysing risk as a function of deployment context, reasoning over regulatory obligations as executable statements rather than prose, and testing whether a body of audit evidence actually supports the decision it is offered for, including the cases where plausible-looking evidence does not.
Outputs
2026
SSRN
Preprint on how the choice of background distribution alters federated attribution results, and what that does to an audit.
PublicationsAuthor
2026
Digital Skills and Jobs Platform, European Commission
A deep-dive for the European Commission's skills platform on where clinical AI is genuinely useful, and what assurance regulated care requires of it.
PolicyAuthor
2026
A reproducible public-records method for analysing AI risk as a function of deployment context rather than of the model alone.
Research softwareAuthor
2026
A benchmark for reasoning over regulatory obligations as executable statements rather than prose.
Research softwareAuthor
2026
Asks whether audit evidence is adequate for the specific decision, and the specific role, relying on it.
Research softwareAuthor
2026
Stress-tests the ways apparently adequate assurance evidence can be assembled to mislead.
Research softwareAuthor
2026
Measures how much a federated explanation depends on an arbitrary modelling choice.
Research softwareAuthor
Projects
A connected set of reproducible tools for deployment-conditioned risk analysis, executable regulatory reasoning, evidence adequacy scoring and adversarial stress testing.