Definitions, methods, and applications in interpretable machine learning
Abstract
Significance The recent surge in interpretability research has led to confusion on numerous fronts. In particular, it is unclear what it means to be interpretable and how to select, evaluate, or even discuss methods for producing interpretations of machine-learning models. We aim to clarify these concerns by defining interpretable machine learning and constructing a unifying framework for existing methods which highlights the underappreciated role played by human audiences. Within this framework, methods are organized into 2 classes: model based and post hoc. To provide guidance in selecting and evaluating interpretation methods, we introduce 3 desiderata: predictive accuracy, descriptive accuracy, and relevancy. Using our framework, we review existing work, grounded in real-world studies which exemplify our desiderata, and suggest directions for future work.
Journal: Proceedings of the National Academy of Sciences
Publisher: National Academy of Sciences
Citations are the number of DOI-registered works in Crossref that cite this paper; references are how many works it cites. Full text is on the publisher site via the DOI link.