The Architectural Bottleneck Principle

November 11, 2022 · Declared Dead · 🏛 Conference on Empirical Methods in Natural Language Processing

Repo contents: .gitignore, LICENSE, README.md

Authors Tiago Pimentel, Josef Valvoda, Niklas Stoehr, Ryan Cotterell arXiv ID 2211.06420 Category cs.CL: Computation & Language Cross-listed cs.LG Citations 5 Venue Conference on Empirical Methods in Natural Language Processing Repository https://github.com/rycolab/attentional-probe ⭐ 3 Last Checked 1 month ago

Abstract

In this paper, we seek to measure how much information a component in a neural network could extract from the representations fed into it. Our work stands in contrast to prior probing work, most of which investigates how much information a model's representations contain. This shift in perspective leads us to propose a new principle for probing, the architectural bottleneck principle: In order to estimate how much information a given component could extract, a probe should look exactly like the component. Relying on this principle, we estimate how much syntactic information is available to transformers through our attentional probe, a probe that exactly resembles a transformer's self-attention head. Experimentally, we find that, in three models (BERT, ALBERT, and RoBERTa), a sentence's syntax tree is mostly extractable by our probe, suggesting these models have access to syntactic information while composing their contextual representations. Whether this information is actually used by these models, however, remains an open question.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 💻 Repository 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Computation & Language

🌅 🌅 Old Age

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, ... (+6 more)

cs.CL 🏛 NeurIPS 📚 166.0K cites 8 years ago

🌅 🌅 Old Age

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin, Ming-Wei Chang, ... (+2 more)

cs.CL 🏛 NAACL 📚 110.2K cites 7 years ago

R.I.P. 👻 Ghosted

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann, ... (+29 more)

cs.CL 🏛 NeurIPS 📚 54.2K cites 5 years ago

R.I.P. 👻 Ghosted

RoBERTa: A Robustly Optimized BERT Pretraining Approach

Yinhan Liu, Myle Ott, ... (+8 more)

cs.CL 🏛 arXiv 📚 28.4K cites 6 years ago

R.I.P. 👻 Ghosted

BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Mike Lewis, Yinhan Liu, ... (+6 more)

cs.CL 🏛 ACL 📚 12.3K cites 6 years ago

R.I.P. 👻 Ghosted

Deep contextualized word representations

Matthew E. Peters, Mark Neumann, ... (+5 more)

cs.CL 🏛 NAACL 📚 12.0K cites 8 years ago

Died the same way — 📜 Death by README

R.I.P. 📜 Death by README

Momentum Contrast for Unsupervised Visual Representation Learning

Kaiming He, Haoqi Fan, ... (+3 more)

cs.CV 🏛 CVPR 📚 14.3K cites 6 years ago

R.I.P. 📜 Death by README

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Peng Gao, Jiaming Han, ... (+10 more)

cs.CV 🏛 arXiv 📚 716 cites 2 years ago

R.I.P. 📜 Death by README

Revisiting Graph based Collaborative Filtering: A Linear Residual Graph Convolutional Network Approach

Lei Chen, Le Wu, ... (+3 more)

cs.IR 🏛 AAAI 📚 609 cites 6 years ago

R.I.P. 📜 Death by README

Diffusion Models for Medical Image Analysis: A Comprehensive Survey

Amirhossein Kazerouni, Ehsan Khodapanah Aghdam, ... (+5 more)

eess.IV 🏛 MedIA 📚 599 cites 3 years ago