Corpus Conversion Service: A machine learning platform to ingest documents at scale [Poster abstract]

May 15, 2018 · Declared Dead · 🏛 SysML 2018

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Peter W J Staar, Michele Dolfi, Christoph Auer, Costas Bekas arXiv ID 1805.09687 Category cs.DL: Digital Libraries Cross-listed cs.CL, cs.CV, cs.DC, cs.IR Citations 0 Venue SysML 2018 Last Checked 3 months ago

Abstract

Over the past few decades, the amount of scientific articles and technical literature has increased exponentially in size. Consequently, there is a great need for systems that can ingest these documents at scale and make their content discoverable. Unfortunately, both the format of these documents (e.g. the PDF format or bitmap images) as well as the presentation of the data (e.g. complex tables) make the extraction of qualitative and quantitive data extremely challenging. We present a platform to ingest documents at scale which is powered by Machine Learning techniques and allows the user to train custom models on document collections. We show precision/recall results greater than 97% with regard to conversion to structured formats, as well as scaling evidence for each of the microservices constituting the platform.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Digital Libraries

R.I.P. 👻 Ghosted

Constructing bibliometric networks: A comparison between full and fractional counting

Antonio Perianes-Rodriguez, Ludo Waltman, Nees Jan van Eck

cs.DL 🏛 J. Informetrics 📚 1.1K cites 9 years ago

R.I.P. 👻 Ghosted

Measuring academic influence: Not all citations are equal

Xiaodan Zhu, Peter Turney, ... (+2 more)

cs.DL 🏛 J. Assoc. Inf. Sci. Technol. 📚 262 cites 11 years ago

R.I.P. 👻 Ghosted

The Open Access Advantage Considering Citation, Article Usage and Social Media Attention

Xianwen Wang, Chen Liu, ... (+2 more)

cs.DL 🏛 Scientometrics 📚 224 cites 11 years ago

R.I.P. 👻 Ghosted

A Bibliometric Review of Large Language Models Research from 2017 to 2023

Lizhou Fan, Lingyao Li, ... (+4 more)

cs.DL 🏛 ACM TIST 📚 208 cites 3 years ago

R.I.P. 👻 Ghosted

On the Performance of Hybrid Search Strategies for Systematic Literature Reviews in Software Engineering

Erica Mourão, João Felipe Pimentel, ... (+4 more)

cs.DL 🏛 IST 📚 157 cites 6 years ago

R.I.P. 👻 Ghosted

A Systematic Identification and Analysis of Scientists on Twitter

Qing Ke, Yong-Yeol Ahn, Cassidy R. Sugimoto

cs.DL 🏛 PLoS ONE 📚 147 cites 9 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Federated Learning: Strategies for Improving Communication Efficiency

Jakub Konečný, H. Brendan McMahan, ... (+4 more)

cs.LG 🏛 arXiv 📚 5.2K cites 9 years ago

R.I.P. 👻 Ghosted

In-Datacenter Performance Analysis of a Tensor Processing Unit

Norman P. Jouppi, Cliff Young, ... (+73 more)

cs.AR 🏛 ISCA 📚 5.1K cites 9 years ago

R.I.P. 👻 Ghosted

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Hoo-Chang Shin, Holger R. Roth, ... (+7 more)

cs.CV 🏛 IEEE TMI 📚 4.9K cites 10 years ago

R.I.P. 👻 Ghosted

Explanation in Artificial Intelligence: Insights from the Social Sciences

Tim Miller

cs.AI 🏛 AI 📚 4.9K cites 8 years ago