"That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks
2204.04636

Authors

Edoardo Mosca,Shreyash Agarwal,Javier Rando-Ramirez,Georg Groh,Javier Rando

Abstract

Adversarial attacks are a major challenge faced by current machine learning research. These purposely crafted inputs fool even the most advanced models, precluding their deployment in safety-critical applications.

Extensive research in computer vision has been carried to develop reliable defense strategies. However, the same issue remains less explored in natural language processing.

Our work presents a model-agnostic detector of adversarial text examples. The approach identifies patterns in the logits of the target classifier when perturbing the input text.

The proposed detector improves the current state-of-the-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks.

Resources

Ray graphicRay graphicRay graphicRay graphic

Stay in the loop

Every AI paper that matters, free in your inbox daily.

Details

  • takara.ai
  • Custom AI and machine learning from the Frontier Research Team.
  • © 2026 takara.ai Ltd
  • Content is sourced from third-party publications.
Ray graphicRay graphicRay graphic