<p>My current search only processing of pdf&#39;s hinges around <a href="https://github.com/ocrmypdf/OCRmyPDF">OCRmyPDF</a> and another manual I need to read is <a href="https://ocrmypdf.readthedocs.io/en/latest/index.html">the one for this</a>. It is essentially just a wrapper around <a href="https://github.com/tesseract-ocr/tesseract">Tesseract OCR</a> engine which recognizes more than <a href="https://github.com/tesseract-ocr/tessdata">100 languages</a>, but I&#39;m just happy with English.</p>

<p>While there are several wrappers listed for tesseract, I&#39;m currently playing with <a href="https://github.com/manisandro/gImageReader">gImageReader</a> but I think it&#39;s the underlying code I need to understand first. Another offering which is showing some promise is <a href="https://scribeocr.com/">ScribeOCR</a> and while that is an on-line service, the <a href="https://github.com/scribeocr/scribeocr">code base</a> will allow me to run a local copy if that turns out to be a better option. Documentation for that <a href="https://docs.scribeocr.com/">is here</a>.</p>

<p>There are several other packages listed on the <a href="https://tesseract-ocr.github.io/tessdoc/User-Projects-%E2%80%93-3rdParty.html">tesseract site</a>, many of which can be ignored, but there may be something else useful.</p>

<p>In addition to the OCR engine, I&#39;ve downloaded a shed load of other applications that were supposed to help pre-process the scanned PDF&#39;s. I'm sure a lot of them could be wipped, but I need to go through and work out just what is useful in a Swiss army knife of tools. For instance OCRmyPDF (I think) provided an extension to the Dolphin file manager menu where I can select it directly and start it running on a number of pdf's. Tailoring that function and adding my own preferences to the list would be good to do.</p>
Page History
Date/CommentUserIPVersion
26 Feb 2025 (19:34 UTC)
—
Lester Caine192.168.1.2541
Current • Source
No records found