Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.

Regularizing Black-box Models for Improved Interpretability

2019·37 Zitationen·arXiv (Cornell University)Open Access

Volltext beim Verlag öffnen

Zitationen

Autoren

2019

Jahr

Abstract

Most of the work on interpretable machine learning has focused on designing either inherently interpretable models, which typically trade-off accuracy for interpretability, or post-hoc explanation systems, whose explanation quality can be unpredictable. Our method, ExpO, is a hybridization of these approaches that regularizes a model for explanation quality at training time. Importantly, these regularizers are differentiable, model agnostic, and require no domain knowledge to define. We demonstrate that post-hoc explanations for ExpO-regularized models have better explanation quality, as measured by the common fidelity and stability metrics. We verify that improving these metrics leads to significantly more useful explanations with a user study on a realistic task.

Autoren

Institutionen

Carnegie Mellon University(US)

Themen

Explainable Artificial Intelligence (XAI)Adversarial Robustness in Machine LearningMachine Learning in Healthcare

Volltext beim Verlag öffnen

Regularizing Black-box Models for Improved Interpretability

Abstract

Ähnliche Arbeiten

Autoren

Institutionen

Themen