Open source repositories tagged with #content, ranked by health score.
Mirror of Apache PDFBox
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).