R.I.P.
π»
Ghosted
PUFFIN: Protein Unit Discovery with Functional Supervision
April 16, 2026 Β· Grace Period Β· π ISMB 2026 proceedings
Authors
GΓΆkΓ§e UludoΔan, Buse Giledereli, Elif Ozkirimli, Arzucan ΓzgΓΌr
arXiv ID
2604.14796
Category
q-bio.BM
Cross-listed
cs.LG
Citations
0
Venue
ISMB 2026 proceedings
Abstract
Proteins carry out biological functions through the coordinated action of groups of residues organized into structural arrangements. These arrangements, which we refer to as protein units, exist at an intermediate scale, being larger than individual residues yet smaller than entire proteins. A deeper understanding of protein function can be achieved by identifying these units and their associations with function. However, existing approaches either focus on residue-level signals, rely on curated annotations, or segment protein structures without incorporating functional information, thereby limiting interpretable analysis of structure-function relationships. We introduce PUFFIN, a data-driven framework for discovering protein units by jointly learning structural partitioning and functional supervision. PUFFIN represents proteins as residue-level structure graphs and applies a graph neural network with a structure-aware pooling mechanism that partitions each protein into multi-residue units, with functional supervision that shapes the partition. We show that the learned units are structurally coherent, exhibit organized associations with molecular function, and show meaningful correspondence with curated InterPro annotations. Together, these results demonstrate that PUFFIN provides an interpretable framework for analyzing structure-function relationships using learned protein units and their statistical function associations. We made our source code available at https://github.com/boun-tabi-lifelu/puffin.
Community Contributions
Found the code? Know the venue? Think something is wrong? Let us know!
π Similar Papers
In the same crypt β q-bio.BM
R.I.P.
π»
Ghosted
Protein secondary structure prediction using deep convolutional neural fields
R.I.P.
π
404 Not Found
LinearFold: linear-time approximate RNA folding by 5'-to-3' dynamic programming and beam search
R.I.P.
π»
Ghosted
What is a meaningful representation of protein sequences?
R.I.P.
π»
Ghosted
Protein Secondary Structure Prediction Using Cascaded Convolutional and Recurrent Neural Networks
R.I.P.
π»
Ghosted