Billion-scale Similarity Search Using a Hybrid Indexing Approach with Advanced Filtering

January 23, 2025 Β· Declared Dead Β· πŸ› Cybernetics and Information Technologies

πŸ‘» CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Simeon Emanuilov, Aleksandar Dimov arXiv ID 2501.13442 Category cs.IR: Information Retrieval Cross-listed cs.DB, cs.DC, cs.LG Citations 5 Venue Cybernetics and Information Technologies Last Checked 4 months ago
Abstract
This paper presents a novel approach for similarity search with complex filtering capabilities on billion-scale datasets, optimized for CPU inference. Our method extends the classical IVF-Flat index structure to integrate multi-dimensional filters. The proposed algorithm combines dense embeddings with discrete filtering attributes, enabling fast retrieval in high-dimensional spaces. Designed specifically for CPU-based systems, our disk-based approach offers a cost-effective solution for large-scale similarity search. We demonstrate the effectiveness of our method through a case study, showcasing its potential for various practical uses.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Information Retrieval

Died the same way β€” πŸ‘» Ghosted