Evading Toxicity Detection with ASCII-art: A Benchmark of Spatial Attacks on Moderation Systems

September 27, 2024 ยท Declared Dead ยท ๐Ÿ› https://aclanthology.org/2025.woah-1.13/

๐Ÿ‘ป CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Sergey Berezin, Reza Farahbakhsh, Noel Crespi arXiv ID 2409.18708 Category cs.CL: Computation & Language Cross-listed cs.AI, cs.CR Citations 3 Venue https://aclanthology.org/2025.woah-1.13/ Last Checked 5 months ago
Abstract
We introduce a novel class of adversarial attacks on toxicity detection models that exploit language models' failure to interpret spatially structured text in the form of ASCII art. To evaluate the effectiveness of these attacks, we propose ToxASCII, a benchmark designed to assess the robustness of toxicity detection systems against visually obfuscated inputs. Our attacks achieve a perfect Attack Success Rate (ASR) across a diverse set of state-of-the-art large language models and dedicated moderation tools, revealing a significant vulnerability in current text-only moderation systems.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Computation & Language

๐ŸŒ… ๐ŸŒ… Old Age

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, ... (+6 more)

cs.CL ๐Ÿ› NeurIPS ๐Ÿ“š 166.0K cites 9 years ago

Died the same way โ€” ๐Ÿ‘ป Ghosted