OpenAlex · Aktualisierung stündlich · Letzte Aktualisierung: 15.03.2026, 08:53

Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.

Explanations can be manipulated and geometry is to blame

2019·145 Zitationen·arXiv (Cornell University)Open Access
Volltext beim Verlag öffnen

145

Zitationen

6

Autoren

2019

Jahr

Abstract

Explanation methods aim to make neural networks more trustworthy and interpretable. In this paper, we demonstrate a property of explanation methods which is disconcerting for both of these purposes. Namely, we show that explanations can be manipulated arbitrarily by applying visually hardly perceptible perturbations to the input that keep the network's output approximately constant. We establish theoretically that this phenomenon can be related to certain geometrical properties of neural networks. This allows us to derive an upper bound on the susceptibility of explanations to manipulations. Based on this result, we propose effective mechanisms to enhance the robustness of explanations.

Ähnliche Arbeiten

Autoren

Institutionen

Themen

Explainable Artificial Intelligence (XAI)Scientific Computing and Data ManagementArtificial Intelligence in Healthcare and Education
Volltext beim Verlag öffnen