History of OCR processing on Linux
<p>My current search only processing of pdf's hinges around <a href="https://github.com/ocrmypdf/OCRmyPDF">OCRmyPDF</a> and another manual I need to read is <a href="https://ocrmypdf.readthedocs.io/en/latest/index.html">the one for this</a>. It is essentially just a wrapper around <a href="https://github.com/tesseract-ocr/tesseract">Tesseract OCR</a> engine which recognizes more than <a href="https://github.com/tesseract-ocr/tessdata">100 languages</a>, but I'm just happy with English.</p>
<p>While there are several wrappers listed for tesseract, I'm currently playing with <a href="https://github.com/manisandro/gImageReader">gImageReader</a> but I think it's the underlying code I need to understand first. Another offering which is showing some promise is <a href="https://scribeocr.com/">ScribeOCR</a> and while that is an on-line service, the <a href="https://github.com/scribeocr/scribeocr">code base</a> will allow me to run a local copy if that turns out to be a better option. Documentation for that <a href="https://docs.scribeocr.com/">is here</a>.</p>
<p>There are several other packages listed on the <a href="https://tesseract-ocr.github.io/tessdoc/User-Projects-%E2%80%93-3rdParty.html">tesseract site</a>, many of which can be ignored, but there may be something else useful.</p>
<p>In addition to the OCR engine, I've downloaded a shed load of other applications that were supposed to help pre-process the scanned PDF's. I'm sure a lot of them could be wipped, but I need to go through and work out just what is useful in a Swiss army knife of tools. For instance OCRmyPDF (I think) provided an extension to the Dolphin file manager menu where I can select it directly and start it running on a number of pdf's. Tailoring that function and adding my own preferences to the list would be good to do.</p>
<p>While there are several wrappers listed for tesseract, I'm currently playing with <a href="https://github.com/manisandro/gImageReader">gImageReader</a> but I think it's the underlying code I need to understand first. Another offering which is showing some promise is <a href="https://scribeocr.com/">ScribeOCR</a> and while that is an on-line service, the <a href="https://github.com/scribeocr/scribeocr">code base</a> will allow me to run a local copy if that turns out to be a better option. Documentation for that <a href="https://docs.scribeocr.com/">is here</a>.</p>
<p>There are several other packages listed on the <a href="https://tesseract-ocr.github.io/tessdoc/User-Projects-%E2%80%93-3rdParty.html">tesseract site</a>, many of which can be ignored, but there may be something else useful.</p>
<p>In addition to the OCR engine, I've downloaded a shed load of other applications that were supposed to help pre-process the scanned PDF's. I'm sure a lot of them could be wipped, but I need to go through and work out just what is useful in a Swiss army knife of tools. For instance OCRmyPDF (I think) provided an extension to the Dolphin file manager menu where I can select it directly and start it running on a number of pdf's. Tailoring that function and adding my own preferences to the list would be good to do.</p>
