Corpus Conversion Service: A machine learning platform to ingest documents at scale [Poster abstract]

May 15, 2018 Β· Declared Dead Β· πŸ› SysML 2018

πŸ‘» CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Peter W J Staar, Michele Dolfi, Christoph Auer, Costas Bekas arXiv ID 1805.09687 Category cs.DL: Digital Libraries Cross-listed cs.CL, cs.CV, cs.DC, cs.IR Citations 0 Venue SysML 2018 Last Checked 3 months ago
Abstract
Over the past few decades, the amount of scientific articles and technical literature has increased exponentially in size. Consequently, there is a great need for systems that can ingest these documents at scale and make their content discoverable. Unfortunately, both the format of these documents (e.g. the PDF format or bitmap images) as well as the presentation of the data (e.g. complex tables) make the extraction of qualitative and quantitive data extremely challenging. We present a platform to ingest documents at scale which is powered by Machine Learning techniques and allows the user to train custom models on document collections. We show precision/recall results greater than 97% with regard to conversion to structured formats, as well as scaling evidence for each of the microservices constituting the platform.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Digital Libraries

Died the same way β€” πŸ‘» Ghosted